Environment. In: ICML 2026. [Yao+, ICLR25] tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. In: ICLR 2025. [Ray+, ICML26] tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains. In: ICML 2026. [Shi+, ICML26] tau-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge. In: ICML 2026. [Cuadron+, arXiv25] SABER: Small Actions, Big Errors — Safeguarding Mutating Steps in LLM Agents. arXiv:2512.07850 [Naous+, ICLR26] Flipping the Dialogue: Training and Evaluating User Language Models. In: ICLR 2026. [Seshadri+, ACL26] Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. In: ACL 2026. [Kapoor+, ICML26WS] Open-World Evaluations for Measuring Frontier AI Capabilities. In: ICML 2026 AIWILD Workshop. [Wang+, ICML26] Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning. In: ICML 2026. [Lyu+, ACL26] Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards. In: ACL 2026. [NVidia, arXiv26] Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning. arXiv:2604.12374 [Raghavendra+, arXiv26] SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions. arXiv:2606.30573 [Budzianowski+, EMNLP18] MultiWOZ — A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling. In: EMNLP 2018. [Chen+, NAACL21] Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems. In: NAACL 2021. [矢野, SBInt26] 日本語エージェントベンチマーク「J-tau telecom」の公開. SB Intuitions Tech Blog, 2026-06-19. [Anthropic, 2025] Project Vend: Can Claude run a small shop? Anthropic Research, 2025-06-27. 22