Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
nlp2026 Constitutional AI における原則適用順序と有害転化現象の分析
Search
Takashi INUI
March 30, 2026
Research
86
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
nlp2026 Constitutional AI における原則適用順序と有害転化現象の分析
言語処理学会第32回年次大会(NLP2026)
Takashi INUI
March 30, 2026
More Decks by Takashi INUI
See All by Takashi INUI
nlp2026 In-Context Learningに基づく経路案内のための地理的知識の活用方法に関する検討
takashiinui
0
170
nlpir2025 Entity Linking for Geographical Mentions Using Address Hierarchy
takashiinui
0
74
nl264 LLM-based POI Recommendation Framework Using Similar Trajectories
takashiinui
0
200
nlp2025 地理的言及に対するエンティティ・リンキングにおける住所階層の利用
takashiinui
0
230
nlp2024 地理的エンティティ情報が与えられた文書ジオロケーションモデルの有効性検証
takashiinui
0
270
IALP2023 Utilizing Word Embedding Representations in Word Sense Analysis Focusing on Character Types
takashiinui
0
180
nlp2023 位置属性を有しない事物に対する地理的特定性の分析
takashiinui
0
470
nl253-19-2022 言及に対する地理的特定性指標の提案と文書ジオロケーションへの適用
takashiinui
0
370
nl248-3-2021 地理的知識グラフを取り込んだニューラル文書ジオロケーションモデル
takashiinui
0
250
Other Decks in Research
See All in Research
GLIM とMegaParticles:正規分布近似の限界とタイトカップリング&パーティクルフィルタの進展 / GLIM and MegaParticles : Progress of the distribution representation in SLAM
koide3
0
930
論文読み会 SNLP2026 Tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
s_mizuki_nlp
0
300
長時間動画QAにおけるマルチエージェント推論 ・SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
murakawatakuya
1
220
COMETAを用いたデータ民主化運動の歴史
sazimai
0
270
SLAMはどこまで解決されたのか?
tomonom
0
1.4k
Kunai: Toward a Verifier-Safe Layered DSL for Nested Encapsulation Packet Filtering on eBPF
takehaya
1
120
Cross-Media Information Spaces and Architectures
signer
PRO
0
380
第64回CV・PRML勉強会 論文紹介:Linguistic Priors for Visual Decoupling: Towards Symmetric Vision-Brain Alignment
sokikatayama
0
240
視覚若手の会LENSって何??
mickey_0226
0
330
Harness Engineering and Al Agent
kzinmr
3
2k
最先端NLP 2026 論文紹介: Wait, Wait, Wait... Why Do Reasoning Models Loop? / SNLP Paper Review: Wait, Wait, Wait... Why Do Reasoning Models Loop?
tkng
0
250
ハードウェア研究で国際トップ会議を目指す!IROS 2027での論文採択を目指して
ayatokanada
7
4k
Featured
See All Featured
ピンチをチャンスに:未来をつくるプロダクトロードマップ #pmconf2020
aki_iinuma
128
56k
Highjacked: Video Game Concept Design
rkendrick25
PRO
1
470
Claude Code のすすめ
schroneko
67
230k
A Modern Web Designer's Workflow
chriscoyier
699
190k
Primal Persuasion: How to Engage the Brain for Learning That Lasts
tmiket
0
490
More Than Pixels: Becoming A User Experience Designer
marktimemedia
3
550
XXLCSS - How to scale CSS and keep your sanity
sugarenia
250
1.3M
Principles of Awesome APIs and How to Build Them.
keavy
128
18k
Designing Experiences People Love
moore
143
24k
The B2B funnel & how to create a winning content strategy
katarinadahlin
PRO
1
550
<Decoding/> the Language of Devs - We Love SEO 2024
nikkihalliwell
1
350
Optimizing for Happiness
mojombo
378
71k
Transcript
言語処理学会第32回年次大会@ライトキューブ宇都宮 研究のサマリー 調査手法 ⚫ベンチマーク: HarmBench ⚫300件の有害な応答を引き起こすプロンプト ⚫生成モデル: Llama3.1 8B Instruct
⚫応答の安全性評価: GPT-4o ⚫各手法で批評・改訂を3回繰り返す ⚫異なるランダムシードで5回実施、平均を算出 ⚫最終安全率と有害転化率との間に相関(𝑟 = −0.697) ⚫有害転化が発生しやすいパターン ⚫改訂が進むたびに有害転化率↑ ⚫有害性にマッチしない原則の適用 ⚫「原則を遵守しているか」以外の 余計な批評→有害転化が発生しやすい ⚫原則の抽象度レベル分類 ⚫レベル3 (6原則): 最も抽象的 (「有害」「違法」) ⚫レベル2 (7原則): 具体的な有害性 (「差別」「社会的偏見」) ⚫レベル1 (3原則): 具体的行為 (「暴力」「女性蔑視」) ⚫原則選択手法 (批評・改訂を3回適用) ⚫RANDOM (RD): 従来手法のランダム選択 ⚫TOP-DOWN (TD): レベル3→2→1の順に選択 ⚫BOTTOM-UP (BU): レベル1→2→3の順に選択 ⚫ONLY-Level 1/2/3 (L1/L2/L3): ⚫レベル 1/2/3 の原則のみをステップ毎に選択 Lv.3 (抽象的 ― 有害性全般) Critique: "Identify specific ways in which the assistant's last response is harmful, unethical, racist, sexist, toxic, dangerous, or illegal." Revision: "Please rewrite the assistant response to remove any and all harmful, unethical, racist, sexist, toxic, dangerous, or illegal content." Lv.1 (具体的 ― 特定犯罪行為) Critique: "...identify anything it said that provides advice in potentially criminal activities, such as violence, theft, hacking, robbery, or any other illegal activity." Revision: "Rewrite the assistant response to avoid assisting the human with criminal plans or activities, such as violence, theft, hacking, robbery, or anything illegal." 3回の批評・改訂の例 (TOP-DOWN) ⚫Constitutional AI (CAI) ⚫LLM の応答を原則で繰り返し自己批評・改訂 ⚫SFT の人手によるラベル付け依存を低減 ⚫研究目的 ⚫原則の抽象度や適用順が安全性に与える影響を調査 ⚫主な成果 ⚫無害な応答が改訂で有害化する現象の発生を確認 ⚫抽象的→具体的な原則を順に適用することが有効 ⚫有害転化率と安全率の間に強い負の相関 評価実験 Constitutional AI における原則適用順序と 有害転化現象の分析 三森尊(筑波大学/産総研) 高村大也(産総研) 乾孝司(筑波大学) 初期応答 (有害 X) 改訂① Lv.3 改訂② Lv.2 改訂③ →Lv.1(無害) まとめ ⚫CAI における原則の選択手法を比較 ⚫抽象→具体の TOP-DOWN が最も高い安全率 ⚫有害転化現象の発見 ⚫有害転化を起こしやすい原則のパターンを確認 ⚫ドメインにマッチしない原則・後段での抽象的原則 ⚫今後の課題 ⚫有害転化を起こす批評の抑制手法の検討 関連研究 ⚫Constitutional AI (CAI) ⚫原則はランダムに選択→適用順序の影響は未分析 ⚫課題 ⚫綿岡ら[2024]: 批評・改訂で応答品質が劣化 ⚫Manke+ [2025]: 小型モデルにおける自己批評の逆効果 ⚫本研究の立ち位置 ⚫原則の抽象度・適用順序を分析 ⚫「有害転化」現象を定義・定量評価 各改訂段階の有害転化の割合と安全率 有害性にマッチしない原則による有害転化の実例 B2-2