Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Japanese_MT-Bench_を使った_LLM_モデルの評価.pdf
Search
Keisuke Kamata
January 24, 2024
1.6k
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Japanese_MT-Bench_を使った_LLM_モデルの評価.pdf
Keisuke Kamata
January 24, 2024
More Decks by Keisuke Kamata
See All by Keisuke Kamata
AI Agent評価体系構築セミナー ハンズオンパート
olachinkei
0
250
AIAgentOps_Weave
olachinkei
0
250
AI Agent評価構築 webinar introduction
olachinkei
0
230
Physical AIを支えるWeights & Biases
olachinkei
1
610
W_Bハッカソン説明会202602.pdf
olachinkei
0
540
MCPサーバー連携をLLMに学ばせる強化学習フレームワークARTを使ってみる (CyberAgent 三橋 亮太)
olachinkei
1
570
W&Bが新しくリリースしたServerless RLの紹介 (W&B 鎌田啓輔)
olachinkei
0
520
WeaveでMCPを記録する & W&BのMCP
olachinkei
1
360
LLMアプリケーションの品質担保に向けた プラクティスと LLMオブザーバビリティツール
olachinkei
1
350
Featured
See All Featured
Principles of Awesome APIs and How to Build Them.
keavy
128
18k
Building AI with AI
inesmontani
PRO
1
1.2k
Measuring Dark Social's Impact On Conversion and Attribution
stephenakadiri
2
280
Redefining SEO in the New Era of Traffic Generation
szymonslowik
1
420
Accessibility Awareness
sabderemane
1
210
Optimising Largest Contentful Paint
csswizardry
37
4k
Future Trends and Review - Lecture 12 - Web Technologies (1019888BNR)
signer
PRO
0
3.7k
Refactoring Trust on Your Teams (GOTO; Chicago 2020)
rmw
35
3.8k
The SEO identity crisis: Don't let AI make you average
varn
0
560
Hiding What from Whom? A Critical Review of the History of Programming languages for Music
tomoyanonymous
3
1.2k
Utilizing Notion as your number one productivity tool
mfonobong
4
590
エンジニアに許された特別な時間の終わり
watany
109
250k
Transcript
Japanese MT-Bench を使った LLM モデルの評価 Meng Lee, Stability AI @
W&B Webinar 2024/01/24
Agenda • 自己紹介 • Japanese Stable LM シリーズ • Japanese
MT-Bench
Meng Lee (メン・リ) 自己紹介 • Stability AI で機械学習エンジニア。日本語大規 模言語モデル(LLM)の研究開発を主導
• SmartNews 時期は会社初の BERT・DistilBert に基づく大規模ニュース分類システムを構築 • 台湾大学情報管理科で情報検索と自然言語処 理を専攻 • 台湾育ての多言語モデル。日本語、英語と中国 語。コードもそこそこ書けます
🦜 Japanese Stable LM シリーズ • モデルサイズ 3B から 70B
の日本語 LLM を公開 • ゼロから学習か、英語のベースモデルから継続学習 • 基盤言語モデルとチャットモデル • 日本語特化の lm-evaluation-harness を開発し、 JGLUE ベースで LLM の言語理解を評価
⚖ Japanese MT-Bench での日本語 LLM の言語生成評価 • Chatbot Arena で有名な
LLM-as-a-judge 論 文の手法に沿って作られた会話形式の日本 語特化の LLM 言語生成評価データセット (MT は Multi-Turn の省略) • 8つの応用領域の問題を含む。各問題に正確 に答えるために、LLMは以下の要求を同時に 満たす必要があります: • 流暢な日本語を生成する • 世界の知識を理解する • 日本文化、社会を理解する • 推論や数学の能力を持つ • 文脈を理解し、利用者と対話すること
⚖ Japanese MT-Bench での日本語 LLM の言語生成評価
⚖ Weights & Biases で Japanese MT-Bench を利用 • Japanese
MT-Bench は、GPT-4 のような強 力な LLM を使用して自動評価を行い、企業 や研究所のための迅速な LLM 開発を可能 にします。 • lm-evaluation-harness・Jaster と一緒に使 用することをお勧めします。これにより、これ らの日本語 LLM のパフォーマンスをより深く 理解することができます。 • Nejumiリーダーボードは日本語特化の LLM 評価を簡単にしてくれる
Stability AI 採用情報:https://ja.stability.ai/careers Japanese Stable LM: https://huggingface.co/stabilityai