自前モデル化=レジリエンス 05 シリコン共設計まで垂直統合 MAI の勝ち筋は「モデル×ハーネス×RLE」の共最適化。Suleyman は 7/29 に「Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus」と宣言した Build 2026(6/2)で7モデル同時発表、その後 Image-2.5-Pro / Voice-2-Flash / Image-2.6 / Thinking-1 / Cyber-1-Flash / Code-1.1-Flash を連続投入。四半期で「a dozen 超」(Satya Q4 決算) PowerPoint で最大 84%、Dynamics 365 Contact Center で最大 89%、MDASH で 50% のコスト削減を実測値として公表。品質は維持または向上 「Every model in a product or agentic system should be substitutable」。単一モデル依存のリスク(セキュリティ・商流・地政学)を構造的に排除する経営判断 Maia 200 上で MAI モデルを動かすと 40% 高い性能/W(Build時点では 1.4x)。モデル・ハーネス・RLE・シリコンの全層を自社で握る Sources: microsoft.ai/news 2026-06-02 / 07-23 / 07-29 / 08-10 / 08-11 / 08-12 / 08-13 2
Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus across the industry. 「トークン最大化」がここ数ヶ月の物語だった。だが次の焦点は業界全体でトークン効率だ。 OPTIMIZING THE FRONTIER PERFORMANCE CURVE · 2026.07.29 “ In most cases, frontier generalist models aren't necessary for every task. By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically. 多くの場合、すべてのタスクにフロンティア汎用モデルは要らない。プロダクト特化のチューニングで、フロンティア性能を維持あるいは超えながらトークンコストを劇的に下げられる。 OPTIMIZING THE FRONTIER PERFORMANCE CURVE · 2026.07.29 “ Every business now must assume that any one model it depends on could disappear, through a security incident, a business or policy misalignment, or a geopolitical shift. いまやどの企業も、依存する単一モデルが消える前提を置かねばならない。セキュリティ事故、商流・方針の不一致、あるいは地政学的な変化によって。 OPTIMIZING THE FRONTIER PERFORMANCE CURVE · 2026.07.29 Source: microsoft.ai/news — Optimizing the frontier performance curve, Mustafa Suleyman (2026-07-29) 7
2025年11月6日に「MAI Superintelligence Team」を組成。Mustafa Suleyman が Microsoft AI の一部として直接リード。 Humanist Superintelligence(人間中心の超知能)。「人と組織に 奉仕し、置き換えないAI」。 SCALE 「lean, talent-dense team of explorers, researchers, and fullstack engineers」。少人数・高密度・短サイクル。 COMPUTE 次世代 GB200 クラスタが稼働中。MAI-1-preview は約15,000基の NVIDIA H100 で事前学習・事後学習。 SILICON 自社アクセラレータ Maia 200 とモデルを共設計。Build 時点で 1.4x、7 月時点で 40% 高い性能/W。 01 No distillation 他社モデルからの蒸留を一切行わない。「capabilities should be learned, not inherited」。模倣は教師の設計判断に縛られ、操縦性を失う。 02 Clean data クリーンで追跡可能・商用ライセンス済みのエンタープライズ級データのみ。「何がモデルを形 作ったか説明できなければ、挙動も理解できず改善もできない」。 03 Own the stack アーキテクチャ・学習パイプライン・事後学習・RL フレームワークまで内製。 04 Own the silicon Maia 200 とのモデル共設計。長期的な自給自足(self-sufficiency)を目的とする。 Source: microsoft.ai/news — Towards Humanist Superintelligence, Mustafa Suleyman (2025-11-06) / Source: microsoft.ai/news — Two in-house models in support of our mission (2025-08-28) / Source: microsoft.ai/news — Introducing MAIThinking-1 (2026-08-12) / Source: microsoft.ai/news — Optimizing the frontier performance curve, Mustafa Suleyman (2026-07-29) 10
capabilities that always work for, in service of, people and humanity more generally. We think of it as systems that are problem-oriented and tend towards the domain specific. Not an unbounded and unlimited entity with high degrees of autonomy – but AI that is carefully calibrated, contextualized, within limits. 常に人と人類のために働く、極めて高度な AI 能力。問題志向でドメイン特化に傾く systems であり、高い自律 性を持つ無制限の存在ではない。慎重に較正され、文脈化され、限界の内側にある AI。 — We reject narratives about a race to AGI. AGI レースという物語を拒否する — HSI is a vision to ensure humanity remains at the top of the food chain. 人類が食物連鎖の頂点に留まることを保証するビジョン — At Microsoft AI, we believe humans matter more than AI. AI より人間が重要だと信じている — Creating superintelligence is one thing; but creating provable, robust containment and alignment alongside it is the urgent challenge of the 21st century. 超知能を作ることと、封じ込め・アラインメントを同時に作ることは別物であり、後者こそ21世紀の急務 Source: microsoft.ai/news — Towards Humanist Superintelligence (2025-11-06) 11
inherited intelligence lacks the steerability essential for real world usage: an imitator is fundamentally tied to the design choices of its teacher and struggles to adapt to new situations. MAI-Thinking-1 was trained without distillation from third party models, forcing our model to truly learn the tasks at hand. 獲得は速いが、実運用に不可欠な操縦性を欠く。模倣者は教師の設計判断に本質的に縛られ、新し い状況への適応に苦しむ。 MAI-Thinking-1 は第三者モデルからの蒸留なしで学習。モデルに「タスクそのものを本当に学ぶ」ことを 強制する。 この選択がもたらす3つの帰結 QUALITY PROVENANCE CONTROL 品質 来歴 制御 ベンチマークを狙い撃ちせず「ゼロから登った」ため、汎化する伸 びしろが残る。AIME 2025 で 97.0%、SWE-Bench Pro で 52.8%。 商用ライセンス済み・追跡可能なデータ系譜。エンタープライズ が法務・監査で説明できる。 安全性を能力と同じ RL インフラで学習。過剰拒否と不適切 応諾を同一の報酬構成で「欠陥」として扱う。 Source: microsoft.ai/news — Introducing MAI-Thinking-1 (2026-08-12) / Source: microsoft.ai/news — Microsoft Build 2026: MAI keynote transcript (2026-06-02) 14
2026 の7モデル MAI-Thinking-1 推論 MAI-Code-1-Flash コーディング 2026年6月2日、Microsoft AI は画像・音声・文字起こし・コーデ MAI-Image-2.5 ィング・推論にまたがる7モデルを一度に発表した。同時に 画像生成・編集 35B active / 約1T total の sparse MoE、256k コンテキスト。AIME 2025 97.0%、SWE-Bench Pro 52.8%。Sonnet 4.6 に対しブラインド人手評価で選好。 5B active。SWE-Bench Pro 51.2%(Haiku 4.5 は 35.2%)。GitHub Copilot ハーネスで直接 学習。 Arena Image Edit で2位(1403±9)、Nano Banana 2 を上回る。$5/$8/$47。 Frontier Tuning と Mayo Clinic との医療フロンティアモデル共 同開発を公表。 MAI-Image-2.5-Flash 大規模本番向けの超効率版。$1.75/$1.75/$19.50。 画像(効率) “The compute used to train frontier models has increased by a factor of one trillion. Now we expect another thousand-fold increase over the next three years.” MAI-Transcribe-1.5 文字起こし MAI-Voice-2 43言語に拡張(25→43)、1時間の音声を15秒未満で処理。キーワードバイアスで WER 最大 30%改善。 15言語、感情タグ制御、5〜60秒のゼロショット音声プロンプト。MAI-Voice-1 比 72% 選好。 音声生成 — Mustafa Suleyman MAI-Voice-2-Flash 超低レイテンシのボイスエージェント向け。7月23日にパブリックプレビュー。 音声(効率) Source: microsoft.ai/news — Building a hill-climbing machine: Launching seven new MAI models (2026-06-02) / Source: microsoft.ai/news — Microsoft Build 2026: MAI keynote transcript (2026-06-02) 19
Our models must not refuse legitimate requests under the guise of safety and compliance as then they are not truly serving humans. 安全性やコンプライアンスを口実に正当な要求を拒否してはならない。それ では人間に本当に奉仕していない。 Source: microsoft.ai/news — Introducing MAI-Thinking-1 (2026-08-12) – 「不適切な応諾」と「不要な拒否」を、同一の報酬構成における欠陥として扱う – 集約は潜在的な危害の重大度に基づく – 安全性は能力と同じ強化学習インフラで学習 — 安全性が能力と常に整合し、付随 的にならない – 結果として、機微な不安全要求には基準を守りつつ、非機微な内容では有用性を保 てている 27
– より大きなモデル、より長い思考、より多いトークンが正義 – 投じたトークンあたりの性能、投じた1ドルあたりの顧客成果 – フロンティア汎用モデルを全タスクに充てる – プロダクトごとにチューニングされた小型モデル群 + ルーティング – コストは「性能の対価」として受け入れられていた – 節約分は顧客に還元する(MAI の明示的な方針) Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus across the industry. How do we get the best possible performance per token invested, and the best real customer outcome per dollar invested? トークン最大化が主役だったが、次の焦点は業界全体でトークン効率だ。1トークン・1ドルあたりの成果をどう最大化するか。 MUSTAFA SULEYMAN · MICROSOFT.AI/NEWS · 2026.07.29 There is definitely a vibe that is shifting from token maxing — how can I burn as many tokens — to how can I actually get to token efficiency, where I can deploy models that are sustainable for the products I'm building, that are sustainable for the business. 「どれだけトークンを燃やせるか」から「トークン効率」へ空気は確実に変わりつつある。プロダクトにもビジネスにも持続可能なモデルを展開する ために。 SOPHIE, MAI PRODUCT TEAM · MICROSOFT REACTOR · 2026.08.11 Why are you using a Ferrari for a job that doesn't necessarily need it? その仕事に本当にフェラーリが必要なのか? NITYA / AMY BOYD · MICROSOFT REACTOR · 2026.08.11 Source: microsoft.ai/news — Optimizing the frontier performance curve, Mustafa Suleyman (2026-07-29) / Source: Microsoft Reactor — Model Mondays S4E1 “Microsoft AI Models Spotlight” (2026-08-11) 36
IMPROVES IT. ヒルクライミング・マシン Data and workflows RLEs ループを所有することの意味 01 Data and workflows 実業務のトレース — 手順、判断、実行されたアクション。「組織の中で仕事が実際にどう進む か」の記録。MAI 曰く「最も価値あるデータはあなたのもの」。 02 RLEs(強化学習環境) AI のための訓練ジム。あなただけがアクセスできる。静的データではなく、ツール呼び出しやスキル 実行を伴うエージェント的挙動を学ばせる。 A hill-climbing machine 03 Tuned model Frontier Tuning で生まれる自社モデル。学習した組織知はモデルの中に残り、所有権もあな たに残る。 04 Usage Tuned model Usage 本番での利用が新たなデータとシグナルを生み、次の周回の入力になる。MAI は「ship, learn, improve, repeat」と表現。 Source: Microsoft Reactor — Model Mondays S4E1 “Microsoft AI Models Spotlight” (2026-08-11) / Source: microsoft.ai/news — Building a hill-climbing machine: Launching seven new MAI models (2026-06-02) / Source: microsoft.ai/news — MAICode-1.1-Flash: Better, faster, at a quarter of the cost (2026-08-11) 42
“ With MAI you don't rent intelligence from a shared model that learns from everyone. Only you keep the benefits of your hard-earned workflows, know-how, data and institutional knowledge. Only you control the resulting model. RLE とは何か Frontier Tuning の効果 所有権の設計 – – – 「AI のための訓練ジム」。決定的・実行可能で、実テス トスイートに採点される環境 – 静的データの学習ではなく、ツール呼び出し・多段作業 同等で最大 10倍効率 – ・失敗からの復旧を練習させる – 自社の環境は自社だけがアクセスできる Excel 向けチューニング済み MAI モデルは GPT-5.4 と McKinsey のタスク向けにチューニングした際は、GPT- 残る – 5.5 を品質で上回りつつコストは 1/10 – 「custom models are both better and more efficient」 学習した組織知はモデルに取り込まれ、そのまま自社に Mayo Clinic との医療フロンティアモデルは、モデルその ものを Mayo Clinic が所有する – Foundry の重み調整(tune the weights yourself )も開発者に開放 Source: microsoft.ai/news — Microsoft Build 2026: MAI keynote transcript (2026-06-02) / Source: microsoft.ai/news — Building a hill-climbing machine: Launching seven new MAI models (2026-06-02) / Source: microsoft.ai/news — Our values in operation: Health (2026-06-02) 43
HEALTH 医療プロダクトでの実測値 Dragon Copilot −50% 文字起こしと言語識別のエラー率を相対的に半減(多言語録音の内部評価、大半の言語で)。 170,000名の医療従事者が使用し、前四半期に2,800万件の患者診療を処理。58言語の多言語ワーク フローを MAI-Transcribe-1.5 が担う。 MAI にとって医療は Humanist Superintelligence の3領域の ひとつであり、同時に最も具体的な ROI が出ている領域でもあ る。 MAI-DxO 85.5% Copilot Health “Once validated, organizations worldwide will be able to access the frontier AI health model. The health model itself will be owned by Mayo Clinic.” 50M+/日 Mayo Clinic — Our values in operation: Health, 2026.06.02 共同開発 NEJM の304症例から作った SD Bench での正答率(o3 併用時)。5〜20年の臨床経験を持つ医師 21名の平均は20%。検査コストも医師・単体モデルより低い。※研究デモであり非公開。 Microsoft の消費者向けプロダクト全体で扱う健康関連の質問数。24カ国230名超の医師による外部パ ネル、ISO/IEC 42001 認証、5万超の米国医療機関と50超のウェアラブルに接続。 医療特化のフロンティアモデルを共同開発。まず Mayo Clinic 自身の環境に展開し、検証後に Foundry 経由で他組織へ。モデルの所有権は Mayo Clinic に帰属する。 Sources: microsoft.ai/news — MAI-Image-2.5-Pro and MAI-Voice-2-Flash (07-23) / Our values in operation: Health (06-02) / The Path to Medical Superintelligence (2025-06-30) / Introducing Copilot Health (03-12) 51
尺 Model Mondays(Microsoft Reactor) シーズン4 第1回 2026年8月11日 Microsoft AI Models Spotlight 約56分 A S Y S Amy Boyd ホスト(UK) Sophie MAI プロダクトチーム Yanan(Jana) MAI PM/コーディングモデル Sharmila ニュースハイライト Microsoft Developer Relations。後半で MAI 画像モデル3種を Foundry 上で実践検証したデモを披露。 Microsoft AI 内のフロンティアモデル研究チームのプロダクト側。MAI の思想、トークン効率、モデルファミリー、ヒルクライミング、RLE を担当。 MAI-Code-1-Flash の学習・評価・A/B 運用を解説し、4本のライブデモを実施。1.1-Flash の画像理解を先出し。 Foundry 全体のモデル動向(AT&T×AMD、GPT-Transcribe、Kimi K3、Claude Opus 5、Microsoft AI Labs)。 位置づけ:本セッションは 8月11日時点の一次情報であり、ブログ未掲載の実運用データ(GitHub Copilot の A/B フライト履歴、MAI-Code-1.1-Flash のベンチ表、1.1-Flash の画像理解)を含む。以降のスライドではこれらを 中心に扱う。 Source: Microsoft Reactor — Model Mondays S4E1 “Microsoft AI Models Spotlight” (2026-08-11) 55
models are built from scratch, not distilled 2 Silicon MAI モデルはゼロから構築され、蒸留されていない 3 Reasoning 4 Coding 5 Tool use 6 Helpfulness 7 Safety 8 Efficiency 9 Harnesses 10 RLEs Trained on a clean lineage of commercially safe training data 商用上安全な、クリーンな系譜の学習データで訓練されている Designed to ensure more efficient Return on Token spend トークン支出に対するリターンをより効率的にするよう設計されている Continuously improving through our hill climbing machine ヒルクライミング・マシンを通じて継続的に改善し続ける 注:スライド上の表記は “RLEEs” だが、文脈・他資料から RLEs(Reinforcement Learning Environments)を指す。 Source: Microsoft Reactor — Model Mondays S4E1 “Microsoft AI Models Spotlight” (2026-08-11) 56
Every gain proven on live product signals, not public benchmarks. Jun 2 Jun 18 Jun 28 Jul 20 Build launch More clients: CLI · Apps · VS GA: Business + Enterprise Experiments for 1.1-Flash RC May Jun 8 Jun 24 Jul 6 Jul 13 Jul 16 Next A/B experiments to decide release candidate Flight: big wins Didn't beat Flight: big wins Didn't beat Flight: big wins 1.1-Flash Jun 8 flight — big wins Jul 6 flight — new engine Jul 16 flight — big wins −25% P50 tokens / user 2× Throughput +2% Survival rate +9% 3-day engagement −10% Time to first token −13% Errors +4% User-initiated turns −25% Time between tokens −10% Turn cancellations −14% Turn cancellations Source: Microsoft Reactor — Model Mondays S4E1 “Microsoft AI Models Spotlight” (2026-08-11) 57