Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
AIは公平な評価, 決断を行えるか? 〜 LLM-as-a-Judgeの限界と意思決定バイア...
Search
Sponsored
·
SiteGround - Reliable hosting with speed, security, and support you can count on.
→
Neurogica
May 04, 2026
Technology
63
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
AIは公平な評価, 決断を行えるか? 〜 LLM-as-a-Judgeの限界と意思決定バイアス 〜 Can AI Make Fair Evaluations and Decisions?
Neurogica
May 04, 2026
More Decks by Neurogica
See All by Neurogica
Sparse Autoencoder 〜 大規模言語モデルの内部表現を読む 〜 Sparse Autoencoder: Understanding LLM's Latent Representations
neurogica
0
18
LLMは“⼼”をモデル化しているのか? Do LLMs Model the Mind?
neurogica
0
1
LLM-MAS×時系列予測 LLM-MAS × Time-Series Forecasting
neurogica
0
35
推薦タスクの整理と 時系列基盤モデル活用を考える How Recommender Systems Use Time-Series Data and Where Foundation Models Fit
neurogica
0
14
画像生成AIの生成速度の遅さと 画風のばらつきを解決する手法 Methods for Solving Image Generation AI's Slowness and Style Inconsistency
neurogica
0
19
DBコネクションプール Database Connection Pooling
neurogica
0
50
時系列基盤モデルは作れるのか? Can We Build a Foundation Model for Time Series?
neurogica
1
110
双曲空間と機械学習 〜 階層性を活かした学習〜 Hyperbolic Space and Machine Learning ~ Learning that Leverages Hierarchical Structure ~
neurogica
0
46
PENGUIN: General Vital Sign Reconstruction from PPG with Flow Matching State Space Models | ICASSP 2026
neurogica
0
71
Other Decks in Technology
See All in Technology
Amazon Quick on DesktopがIAM Identity Centerで動かない理由
yukiogawa
0
190
ユーザー価値を届け続けるためにウォンテッドリーが大切にしている文化
kotaminato
0
160
AI de Idea
kawaguti
PRO
2
110
研究開発部の紹介 / Sansan R&D Profile
sansan33
PRO
5
25k
30座EKS, 180次升級淬煉的EKS Upgrade Skill 的歷程
eric8230
0
160
The Agent Builder Loop from Daily Work to OSS
minorun365
PRO
3
200
AIによるクリエイティブ生成を行う上での試行錯誤
plaidtech
PRO
0
150
AI に書かせたその API、 “信頼” できますか?
nagix
0
110
事業課題から技術的負債に向き合う
sansantech
PRO
2
1.8k
映像変換サーバーなしで端末内でHLSを生成してライブ配信
hikarusato
0
110
Spring BootからQuarkusへの移行
tatsuya1bm
2
120
幾何アルゴリズムで なめらかなピン操作を / iOSDC Japan 2026 / smoothpin
kazumanagano
0
350
Featured
See All Featured
The Hidden Cost of Media on the Web [PixelPalooza 2025]
tammyeverts
2
500
Connecting the Dots Between Site Speed, User Experience & Your Business [WebExpo 2025]
tammyeverts
11
1k
Sam Torres - BigQuery for SEOs
techseoconnect
PRO
0
540
For a Future-Friendly Web
brad_frost
183
10k
Rails Girls Zürich Keynote
gr2m
96
14k
Redefining SEO in the New Era of Traffic Generation
szymonslowik
1
420
Joys of Absence: A Defence of Solitary Play
codingconduct
1
520
Agile Leadership in an Agile Organization
kimpetersen
PRO
0
230
A Soul's Torment
seathinner
8
3.6k
Navigating Team Friction
lara
192
16k
Thoughts on Productivity
jonyablonski
76
5.4k
Beyond borders and beyond the search box: How to win the global "messy middle" with AI-driven SEO
davidcarrasco
3
240
Transcript
AIは公平な評価, 決断を⾏えるか? 〜 LLM-as-a-Judgeの限界と意思決定バイアス 〜 株式会社ニューロジカ 開発部 三ツ井 智哉 AIは公平な評価,
決断を⾏えるか? 〜 LLM-as-a-Judgeの限界と意思決定バイアス 〜
⾃然⾔語⽣成(NLG)の評価 • ⼈間による評価:コストが⾼く時間がかかる • Large Language Models (LLMs) による評価:低コスト、時短 G-Eval
(EMNLP 2023) [1] • GPT-4などのLLMをNLGの評価者として利⽤ • Chain-of-Thought(CoT)などを活⽤し, ⼈間の評価との⾼い相関を実現 1,000件のモデル応答 LLM as a Judge とは © Neurogica Inc. はじめに [1] Liu, Yang, et al. "G-eval: NLG evaluation using gpt-4 with better human alignment." Proceedings of the 2023 conference on empirical methods in natural language processing. 2023. 問題点 解決案 エージェント まとめ 数⽇〜数週間 数⼗万円 ⼈によりばらつき 数分 数万円 同⼀の基準 ⼈間 LLM
Large Language Models are not Fair Evaluators (ACL 2024) [2]
• LLMの位置バイアスの実証 • “AとB, どちらの出⼒が良いか?” という⽐較において, 選択肢の提⽰順(AとB)を⼊れ替えるだけで勝敗判定が変化 • 順番を⼊れ替えた結果を平均化する⼯夫が必要 LLM as a Judge の問題点 ① はじめに [2] Wang, Peiyi, et al. "Large language models are not fair evaluators." Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. 問題点 解決案 エージェント まとめ Which response is better? Response 1: … Response 2: … Which response is better? Response 1: … Response 2: … Response 1: … Response 2: … LLM Judge © Neurogica Inc.
Self-Preference Bias in LLM-as-a-Judge (NeurIPS 2024 Workshop) [3] • LLMは⾃⾝の⽣成物を不当に⾼く評価してしまう傾向
• モデルにとって馴染みやすい, 予測しやすい⽂章を⾼く評価 • ⽣成側と評価側のスタイルの相性が混⼊ LLM as a Judge の問題点 ② はじめに [3] Wataoka, Koki, Tsubasa Takahashi, and Ryokan Ri. "Self-Preference Bias in LLM-as-a-Judge." Neurips Safe Generative AI Workshop 2024. 問題点 解決案 エージェント まとめ © Neurogica Inc.
Pairwise or Pointwise? (COLM 2025) [4] • 評価プロトコル⾃体の違いによるバイアスの受けやすさを⽐較 • Pairwise(A/B⽐較):順番の⼊れ替えなどで評価が逆転する割合が約35%
と⾼い • Pointwise(絶対評価・スコアリング):評価のブレが約9%に留まり, より ノイズに対して頑健 問題の解決案 はじめに [4] Tripathi, Tuhina, et al. "Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation." Second Conference on Language Modeling., 2025 問題点 解決案 エージェント まとめ Which response is better? Response 1: … Response 2: … Scoring this response Response 1: … Pairwise Pointwise © Neurogica Inc..
• LLM-as-a-Judge:出⼒を評価 • LLM Agent:検索・ブラウザ操作・外部ツール実⾏・購買・予約・調査 評価から意思決定(エージェント)へ はじめに 問題点 解決案 エージェント
まとめ • Claudeは画⾯操作・クリック・⼊⼒を⾏う Computer Useを提供 • OpenAI Operatorはブラウザを使ってタスクを 実⾏するAgentとして公開 © Neurogica Inc.
Actions Speak Louder than Words [5] • 差別的な回答をしないよう調整されたLLMでも, エージェントとしての意思決 定には潜在的な社会的バイアス
エージェントにおけるバイアス ① はじめに [5] Li, Yuxuan, Hirokazu Shirado, and Sauvik Das. "Actions speak louder than words: Agent decisions reveal implicit biases in language models." Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 2025. 問題点 解決案 エージェント まとめ • LLMにペルソナを与え, 避難・融資・採⽤などの 意思決定シナリオで⽐較 • ほぼすべてのシミュレー ションで有意な意思決定 格差が観測 © Neurogica Inc.
What Is Your AI Agent Buying? [6] • LLM購買エージェントについて, モデルごとの選好バ
イアスを実証 • 商品の位置, 価格, レビュー数, 広告タグの有無などに 対する感応度がモデルごと, 世代ごとに異なる. エージェントにおけるバイアス ② はじめに [6] Allouah, Amine, et al. "What Is Your AI Agent Buying? Evaluation, Biases, Model Dependence, & Emerging Implications of Agentic E-Commerce." Proceedings of the ACM Web Conference 2026. 問題点 解決案 エージェント まとめ ←位置バイアスを⽰す タグや評価による→ 影響を⽰す © Neurogica Inc.
まとめ はじめに 問題点 解決案 エージェント まとめ LLM評価の問題 • 選択肢の順番にバイア スが存在
• 評価するモデルと同じ モデルの⽣成物を⾼く 評価 エージェントへの進化 • テキスト評価だけでな く操作, 意思決定も • LLMの意思決定にもバ イアスが存在 © Neurogica Inc.. LLM as a Judge ⼈間と近い評価を⼈間よりは るかに短時間低コストで実⾏ 可能