Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
論文読み会 SNLP2026 Tau2-Bench: Evaluating Conversat...
Search
S
August 19, 2026
Research
2
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
論文読み会 SNLP2026 Tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
S
August 19, 2026
More Decks by S
See All by S
論文読み会 SNLP2025 Learning Dynamics of LLM Finetuning. In: ICLR 2025
s_mizuki_nlp
0
490
論文読み会 SNLP2024 Instruction-tuned Language Models are Better Knowledge Learners. In: ACL 2024
s_mizuki_nlp
1
630
埋め込み表現の意味適応による知識ベース語義曖昧性解消
s_mizuki_nlp
2
610
論文読み会 SNLP2018 Sequence to Action: End to End Semantic Graph Generation for Semantic Parsing
s_mizuki_nlp
0
140
論文読み会 SNLP2019 Ordered neurons: Integrating tree structures into recurrent neural networks
s_mizuki_nlp
0
150
論文読み会 SNLP2020 ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
s_mizuki_nlp
0
210
論文読み会 SNLP2021 A Distributional Approach to Controlled Text Generation
s_mizuki_nlp
0
170
Other Decks in Research
See All in Research
[IR Reading 2026春 論文紹介] LLM-based Listwise Reranking under the Effect of Positional Bias (ECIR 2026) /IR-Reading-2026-Spring
koheishinden
PRO
0
330
Language and AI
ayaniwa
0
210
2025年度秋葉原ウォーカブルプロジェクト調査報告 「アキバらしいウォーカブル」とは何か
izumiyama_lab
1
180
AIで最適化を解けるか?
mickey_kubo
0
150
[CV勉強会@関東 CVPR2026] PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow / kantocv 67th CVPR 2026
shunk031
0
250
多様なデータを許容し学習し続ける模倣学習 / Advanced Imitation Learning for VLA
prinlab
0
280
マーケットストリート 社会実験2024 in 秋葉原ジャンク通り 調査報告書
izumiyama_lab
1
110
260624_NLP-colloquium: Hubness
de9uch1
1
180
Anthropic が提案する LLM の内部状態を自然言語で説明可能にした Natural Language Autoencoders / Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
shunk031
0
170
AIエージェント時代のLLM-jpモデルのあるべき姿
k141303
0
570
Fukui Shibiten 39 - AI Art
butchi
0
180
EIRによる不正端末のブロッキング 5G時代におけるデバイス識別と不正対策の進化
stellarcraft
0
120
Featured
See All Featured
The Psychology of Web Performance [Beyond Tellerrand 2023]
tammyeverts
49
3.5k
Google's AI Overviews - The New Search
badams
0
1.1k
The agentic SEO stack - context over prompts
schlessera
0
870
End of SEO as We Know It (SMX Advanced Version)
ipullrank
3
4.4k
Understanding Cognitive Biases in Performance Measurement
bluesmoon
32
3k
[SF Ruby Conf 2025] Rails X
palkan
2
1.3k
Creating an realtime collaboration tool: Agile Flush - .NET Oxford
marcduiker
35
2.5k
The Cost Of JavaScript in 2023
addyosmani
55
10k
The #1 spot is gone: here's how to win anyway
tamaranovitovic
3
1.1k
First, design no harm
axbom
PRO
2
1.3k
GraphQLの誤解/rethinking-graphql
sonatard
75
12k
個人開発の失敗を避けるイケてる考え方 / tips for indie hackers
panda_program
123
22k
Transcript
ICML2026 Spotlight Tau2-bench: Evaluating Conversational Agents in a Dual-Control Environment.
In: ICML 2026 Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik R Narasimhan 第18回 最先端NLP勉強会 Hottolink/SciTokyo Okazaki Lab/AIST: Sakae Mizuki 2026-08-30
概要 2
まとめと考察 3