Upgrade to Pro — share decks privately, control downloads, hide ads and more …

逆最適化と機械学習

Avatar for Shinsaku Sakaue Shinsaku Sakaue
September 04, 2026
530

 逆最適化と機械学習

日本オペレーションズ・リサーチ学会 2026 年秋季シンポジウム(第 94 回)
https://orsj.org/2026f/symposium/

Avatar for Shinsaku Sakaue

Shinsaku Sakaue

September 04, 2026

More Decks by Shinsaku Sakaue

Transcript

  1. 自己紹介 2016/4 2 / 46 2018/10 2020/4 修士@東大 (武田先生) 2024/8

    2025/4 NTT CS 研 社会人博士@京大 (湊先生) 2025/7 CyberAgent AI Lab 出向:特任助教@東大(数理情報第 7 研) RIKEN AIP 客員 NII (JST BOOST) RIKEN AIP 客員 研究の興味 最適化(離散・連続) 学習理論(リグレット・汎化誤差解析) データ構造(BDD・ZDD) … コミュニティ NeurIPS ICML COLT AISTATS OR 学会 IBIS …
  2. 逆最適化 4 / 46 • x は決定変数,X は実行可能領域 • c

    ∈ Rd は目的関数パラメータ,f (·, c) は目的関数 • xobs ∈ X は観測された解(観測解) 順問題 Given: c, X, f (·, ·) 逆問題 obs Given: x , X, f (·, ·) Find: x⋆ (c) ∈ arg min f (x, c) x∈X Find: ĉ s.t. xobs ∈ arg min f (x, ĉ) x∈X 逆最適化は観測解 xobs の最適性を説明する目的関数パラメータ ĉ を探す問題
  3. 文脈付き目的関数パラメータ予測 5 / 46 z は意思決定時に利用できる文脈特徴量 予測モデル:文脈 z から目的関数パラメータ c

    を予測 z ĉ = Mθ (z) xpred ∈ arg min f (x, ĉ) x∈X • Mθ は目的関数パラメータの予測モデル,θ は Mθ のパラメータ • ĉ は文脈からモデルが予測した目的関数パラメータ • xpred は予測された ĉ による順問題の求解結果 観測データから文脈と目的関数パラメータの関係 Mθ の学習を目指す
  4. 例:購買選択からの選好推定 z :購買文脈 6 / 46 c:選好の重み 予測 x:購買選択 求解

    価格・容量・ブランド 価格・品質・健康を 選択される商品・数量 顧客属性・利用場面 どれだけ重視するか 商品棚で定まる選択肢 X 順問題:商品属性の加重効用を最大化する購買選択 学習:過去の購買選択と整合する選好の重み c の推定 利用:購買選択の予測,需要予測・品揃え設計への活用
  5. 例:発注量からの利益モデル推定 c:利益モデルの パラメータ z :発注文脈 予測 天気・曜日・過去の販売 現在庫・調達リードタイム 7 /

    46 x:発注量 求解 商品ごとの発注量 需要の見積もり 予算・保管容量で 売れ残り・欠品への評価 定まる実行可能領域 X 順問題:期待利益を最大化する発注量選択 学習:熟練者の発注と整合する利益モデルパラメータ c の推定 利用:新しい発注文脈における熟練者に近い発注判断の自動化
  6. 学習データの二つの設定 8 / 46 目的関数パラメータの観測 {(zi , Xi , ci

    )}N i=1 ci 観測可能.観測解 xobs を含む場合もある(xobs が ci に対し最適なら不要) i i 選ばれた解の観測 N {(zi , Xi , xobs i )}i=1 ci 観測不可.xobs は i 番目の観測解. i 共通の観測:文脈 zi と実行可能領域 Xi 異なる観測:目的関数パラメータ ci or 観測解 xobs i
  7. 二つの評価指標と学習手法 10 / 46 MSE:目的関数パラメータの予測誤差 1 X MSE = ∥ĉi

    − ci ∥2 N i=1 N Regret:予測が誘導した意思決定の品質  1 X Regret = f (x⋆ (ĉi ), ci ) − f (x⋆ (ci ), ci ) N i=1 N 注.x⋆ (ĉi ) は任意に tie-break PFL:MSE 最小化による ĉ の学習(二段階法) DFL:Regret の最小化を目指した ĉ の学習(end-to-end) PFL = prediction-focused learning,DFL = decision-focused learning [Mandi et al. 2024] Mandi et al. (2024). Decision-focused learning: Foundations, state of the art, benchmark and future opportunities. Journal of Artificial Intelligence Research.
  8. 二段階法:MSE に基づく通常の回帰 11 / 46 損失関数と更新式 1 X LMSE (θ)

    = ∥ĉi − ci ∥2 , N i=1 N ĉi = Mθ (zi ), θ ← θ − η∇θ LMSE (θ) • 損失関数:目的関数パラメータの MSE • 学習時計算:標準的な回帰 • 予測時計算:順問題求解 xpred (ĉ) = x⋆ (ĉ) ∈ arg min f (x, ĉ) x∈X
  9. DiffOpt:微分可能最適化層による DFL 12 / 46 [Amos and Kolter 2017; Wilder

    et al. 2019] z ĉ = Mθ (z) 微分可能求解 xpred (ĉ) f (xpred (ĉ), c) 誤差逆伝播 ĉi = Mθ (zi ), LReg (θ) = N  1 X f (xpred (ĉi ), ci ) − f (x⋆ (ci ), ci ) , θ ← θ − η∇θ LReg (θ) N i=1 • 損失関数:Regret • 学習時計算:微分可能求解,目的関数値評価,誤差逆伝播 • 予測時計算:順問題求解 注.DiffOpt という呼び方は非標準的 Amos and Kolter (2017). OptNet: Differentiable optimization as a layer in neural networks. International Conference on Machine Learning. Wilder et al. (2019). Melding the data-decisions pipeline: Decision-focused learning for combinatorial optimization. AAAI Conference on Artificial Intelligence.
  10. DiffOpt の実装イメージ 13 / 46 JAX による実装イメージ def batch_loss(params, z,

    X, c): c_pred = model(params, z) x_pred = vmap(differentiable_solve)(c_pred, X) # vmap は solve の自動ベクトル化 return mean(f(x_pred, c)) loss, grads = value_and_grad(batch_loss)(params, z, X, c) JAX 実装例:JAXopt(QP)[Blondel et al. 2022],MPAX(LP・QP)[Lu et al. 2025] 汎用の凸最適化層:CVXPYlayers(CVXPY で記述した問題)[Agrawal et al. 2019] Blondel et al. (2022). Efficient and modular implicit differentiation. Advances in Neural Information Processing Systems. Lu et al. (2025). MPAX: Mathematical programming in JAX. Conference on Neural Information Processing Systems, Workshop on GPU-Accelerated and Scalable Optimization. Agrawal et al. (2019). Differentiable convex optimization layers. Advances in Neural Information Processing Systems.
  11. 各手法の計算要件 14 / 46 二段階法:MSE を用いる高速な回帰ベースライン SPO+:代理損失を用いる中間的手法 [Elmachtoub and Grigas

    2022] (末尾補足参照) DiffOpt:微分可能最適化層を用いる DFL 手法 必要な観測 学習時の求解 ソルバー微分 二段階法 SPO+ DiffOpt c c c なし 修正係数で通常求解 ĉ で求解 不要 不要 必要 Elmachtoub and Grigas (2022). Smart “Predict, then Optimize”. Management Science.
  12. 実験例の紹介:制約付き QP 15 / 46 [高橋 2025, Case 2] 文脈

    z に基づくパラメータ c の予測と,線形不等式制約付き二次計画の求解 z ∈ R5 , c, x ∈ R2 , ĉ = Mθ (z)   1 ⊤ pred ⊤ x ∈ arg min x Qx + ĉ x : Gx ≤ h . 2 • Q, G, h は固定 • Mθ は多層パーセプトロン(二段階法と DiffOpt で共通) • zi から ci を非線形な関係で生成した 1000 標本:訓練 800,テスト 200 高橋 優輝 (2025). JAXopt で実践する Decision-focused Learning ~機械学習×数理最適化の新たなアプローチ~ . NTT ドコモ R&D ENGINEERING BLOG.
  13. 比較結果 16 / 46 [高橋 2025, Case 2] 学習方法 二段階法

    DiffOpt(JAXopt) MSE Regret 49.75 4130.17 15.70 3.07 注.表記対応:記事中の PFL / DFL,本講演の二段階法/ DiffOpt. • DiffOpt は二段階法より MSE は大きく Regret は小さい • 極端な予測が安定した低 Regret な解の積極的選択に寄与 Top-k などの単純な制約では二段階法が Regret で有利な場合もある [Geng et al. 2024] 高橋 優輝 (2025). JAXopt で実践する Decision-focused Learning ~機械学習×数理最適化の新たなアプローチ~ . NTT ドコモ R&D ENGINEERING BLOG. Geng et al. (2024). Benchmarking PtO and PnO methods in the predictive combinatorial optimization regime. Advances in Neural Information Processing Systems.
  14. 観測解に基づく学習 18 / 46 N 訓練データは {(zi , Xi ,

    xobs i )}i=1 ,目的関数パラメータ ci の観測は無し. 学習プロトコル:損失関数を L(後述)として, " # N 1 X ĉi = Mθ (zi ), θ ← θ − η∇θ L(ĉi ; xobs i , Xi ) N i=1 予測時の処理: z ĉ = Mθ (z) xpred ∈ arg min f (x, ĉi ) x∈Xi
  15. 逆最適化で観測欠損を補う手法 19 / 46 [Mishra et al. 2024; Hikima and

    Kamiyama 2025; Hikima et al. 2026] 仮定:目的関数 f (x, c) = c⊤ x は線形 入力:実行可能領域 Xi ,観測解 xobs i ,基準係数 ri (≈ Mθ (zi ))    obs ⊤ , qi ∈ arg min∥q − ri ∥ Ci := q xi ∈ arg min q x = −Nconv(Xi ) xobs i x∈Xi q∈Ci qi − r i Ci ri :基準係数 qi :逆最適化の出力 目的関数パラメータ空間 現在のモデルの出力 Mθ (zi ) を観測解 xobs を最適にする方向に修正 i Mishra et al. (2024). From inverse optimization to feasibility to ERM. International Conference on Machine Learning. Hikima and Kamiyama (2025). An inverse optimization approach to contextual inverse optimization. International Joint Conference on Artificial Intelligence. Hikima et al. (2026). A sampling-based relaxation approach to contextual inverse optimization. International Joint Conference on Artificial Intelligence.
  16. 逆最適化で観測欠損を補う手法 20 / 46 [Mishra et al. 2024; Hikima and

    Kamiyama 2025; Hikima et al. 2026] 損失関数:逆最適化で得た目標係数 qi への二乗距離 LIO (ĉi ; xobs i , Xi ) := 1 ∥ĉi − qi ∥22 , 2 ∇θ LIO = [Jθ Mθ (zi )]⊤ (ĉi − qi ) 注.Jθ Mθ (zi ):θ に関する Jacobian;qi は固定ターゲット扱い. 各反復で逆最適化とモデルの更新を交互に実行: 1 観測解 xobs を最適解として持つ係数 qi を逆最適化手法で計算 i 2 欠損している観測 ci を qi で置き換え二段階法と同様の回帰で学習 Mishra et al. (2024). From inverse optimization to feasibility to ERM. International Conference on Machine Learning. Hikima and Kamiyama (2025). An inverse optimization approach to contextual inverse optimization. International Joint Conference on Artificial Intelligence. Hikima et al. (2026). A sampling-based relaxation approach to contextual inverse optimization. International Joint Conference on Artificial Intelligence.
  17. Suboptimality loss を用いる手法 21 / 46 [Troutt et al. 2006;

    Mohajerin Esfahani et al. 2018; Kitaoka 2024] 仮定:目的関数パラメータへの線形依存 f (x, c) = q(x) + c⊤ x, q : X → R は既知の c 非依存関数 例:固定行列 Q に対する f (x, c) = 12 x⊤ Qx + c⊤ x Suboptimality loss:ĉ の下での観測解 xobs の非最適度 Lsub (ĉ; xobs , X) = f (xobs , ĉ) − min f (x, ĉ) x∈X | {z } | {z } 観測解での値 最適値 凸性:Lsub は ĉ について凸 ∵ 第一項は ĉ のアフィン関数,第二項はアフィン関数の点ごとの最小値の負 Troutt et al. (2006). Behavioral estimation of mathematical programming objective function coefficients. Management Science. Mohajerin Esfahani et al. (2018). Data-driven inverse optimization with imperfect information. Mathematical Programming. Kitaoka (2024; rev. 2026). Exact solution to data-driven inverse optimization of MILPs in finite time via gradient-based methods. arXiv:2405.14273.
  18. Suboptimality loss の劣勾配 22 / 46 pred obs − x

    定義:ĉ の下での観測解 xobs の非最適度 xpred x Lsub (ĉ; xobs , X) = f (xobs , ĉ) − min f (x, ĉ) −ĉ x∈X 劣勾配:観測解と予測解の差 予測解を xpred ∈ arg min f (x, ĉ) とすると,Danskin の定理より x∈X xobs − xpred ∈ ∂ĉ Lsub (ĉ; xobs , X) ⊤ 予測モデルパラメータ θ についての劣勾配は [Jθ Mθ (z)] (xobs − xpred ) 実装上の要点:ソルバー微分不要;解差分を予測モデルへ誤差逆伝播 より一般の Fenchel–Young 損失から継承される性質 [Blondel et al. 2020] Blondel et al. (2020). Learning with Fenchel–Young losses. Journal of Machine Learning Research. X xobs
  19. 注意と実装イメージ 23 / 46 注意:零ベクトルへの退化 f (x, c) = c⊤

    x の場合は Lsub (0; xobs , X) = 0 正則化や制約がなければ,ĉ = 0 が自明な最小解. JAX による実装イメージ:正則化付き学習 def batch_loss(params, z, X, x_obs): c_pred = model(params, z) x_pred = stop_gradient(vmap(solve)(c_pred, X)) subopt = f(x_obs, c_pred) - f(x_pred, c_pred) return mean(subopt) + lam * R(c_pred) loss, grads = value_and_grad(batch_loss)(params, z, X, x_obs) 注.後述の実験では正則化なし (lam = 0) で実装 # ソルバー微分なし # 正則化
  20. 実験 24 / 46 [坂上 2025] 高橋 (2025, Case 2)

    の設定で二段階法・DiffOpt との比較 Suboptimality loss は学習時には c 不使用(MSE, Regret の評価にのみ使用) 学習方法 二段階法 DiffOpt(JAXopt) Suboptimality loss 学習時間(秒) MSE Regret 0.33 59.11 53.98 113.93 3087.79 2682.62 18.28 6.32 2.76 • この実験では suboptimality loss による学習が最良 Regret を達成 • 計算時間は二段階法が最短,suboptimality loss は DiffOpt よりやや短い 高橋 優輝 (2025). JAXopt で実践する Decision-focused Learning ~機械学習×数理最適化の新たなアプローチ~ . NTT ドコモ R&D ENGINEERING BLOG. 坂上 晋作 (2025). パラメータ観測が不要な「予測→最適化」 ~機械学習×数理最適化のもう一つのアプローチ~ . Zenn.
  21. オンライン学習設定 26 / 46 オンライン学習 各時刻 t で Xt の観測

    xalg の出力 xobs の観測 ĉt+1 への更新 t t 逐次的により良い ĉt を学習し,より良い xalg を出力することを目指す t 仮定 目的関数は線形 f (x, c) = ⟨c, x⟩,c は全時刻 t = 1, . . . , T で共通 注.Mθ が線形ならば文脈 zt は Xt に含めて扱える: c = Mθ を zt から係数ベクトル Mθ zt への変換行列,Xtorig を zt 非依存の実行可能領域, Xt = {xorig zt⊤ : xorig ∈ Xtorig } とすると,⟨Mθ zt , xorig ⟩ = ⟨c, xorig zt⊤ ⟩ (∀xorig ∈ Xtorig ).
  22. プロトコルと応用 27 / 46 [Bärmann et al. 2017] オンライン学習プロトコル For

    t = 1, . . . , T : 1 実行可能領域 Xt を観測 2 xalg t ∈ arg min{⟨ĉt , x⟩ : x ∈ Xt } を出力 # ĉ1 = 0 とする 3 xobs ∈ arg min{⟨c, x⟩ : x ∈ Xt } を観測 t # 最適フィードバック,c は未知 4 観測に応じて ĉt+1 へと更新 応用例:商品推薦・経路選択などにおける xalg の逐次的推薦 t Bärmann et al. (2017). Emulating the expert: Inverse optimization through online learning. International Conference on Machine Learning.
  23. 評価指標:Regret 28 / 46 Regret:真の係数ベクトル c に対する累積目的関数値ギャップ RT (c) =

    T D X t=1 obs c, xalg t − xt E = T X t=1 | f (xalg t , c) {z } 学習器の累積目的関数値 − T X t=1 | f (xobs t , c) {z } 観測解の累積目的関数値 • 最適フィードバックの場合 RT (c) ≥ 0 • RT (c) が小さいほど学習器は未知の c に対して良い解 xalg を出力 t
  24. Regret 上界と計算量の改善 29 / 46 d は目的関数パラメータ c ∈ Rd

    の空間の次元,T はラウンド数. Bärmann et al. (2017, 2018) Besbes et al. (2021, 2025) Gollapudi et al. (2021) Sakaue et al. (2025a) Sakaue (2026) Regret 上界 √ O( T ) O(d4 log T ) O(d log T ) O(d log T ) O(d log T ) 各時刻の計算量 T 非依存 d, T の多項式時間 d, T の多項式時間 典型的には O(d3 )(T 非依存) 典型的には O(d2 )(T 非依存) • Regret の下界は Ω(d) [Sakaue et al. 2025a] • 実行可能領域が M 凸集合の場合は O(d log d) Regret [Oki and Sakaue 2026] • 真の c が決定境界から ∆ 離れている場合は O(1/∆2 ) Regret [Sakaue et al. 2025b]
  25. √ Perceptron 型更新と O( T ) Regret 各時刻の手続き ζ0 =

    0 で初期化し,各時刻 t で • Return: xalg t ∈ arg min{⟨ĉt , x⟩ : x ∈ Xt } for ĉt = −ζt−1 obs • Update: Observe xobs − xalg and ζt = ζt−1 + gt t , set gt = xt t √ 定理:O( T ) Regret 上界 √ xobs ∈ Xt ,diam(Xt ) ≤ D,∥c∥ ≤ 1 ならば RT (c) ≤ D T t 30 / 46
  26. √ O( T ) Regret の証明 31 / 46 学習器の出力の最適性:

    alg ⟨ζt−1 , gt ⟩ = −⟨ĉt , xobs t − xt ⟩ ≤ 0 二乗ノルムポテンシャルの上界: ∥ζt ∥2 − ∥ζt−1 ∥2 = 2⟨ζt−1 , gt ⟩ + ∥gt ∥2 ≤ D2 から √ ∥ζT ∥ ≤ D T Cauchy–Schwarz: RT (c) = T D X t=1 obs c, xalg t − xt | {z =−gt } E √ = −⟨c, ζT ⟩ ≤ ∥c∥ ∥ζT ∥ ≤ D T
  27. Second-order Perceptron に基づく手法 32 / 46 [Sakaue 2026] 各時刻の手続き A0

    = λI (λ > 0),ζ0 = 0 で初期化し,各時刻 t で −1 • Return: xalg t ∈ arg min{⟨ĉt , x⟩ : x ∈ Xt } for ĉt = −At−1 ζt−1 obs − xalg • Update: Observe xobs t , set gt = xt t , At = At−1 + gt gt⊤ and ζt = ζt−1 + gt 定理:O(d log T ) Regret 上界 xobs ∈ arg min{⟨c, x⟩ : x ∈ Xt },diam(Xt ) ≤ D = O(1),∥c∥ ≤ 1 ならば, t λ = D2 d とすることで RT (c) = O(d log T ). Sakaue (2026). Simple projection-free algorithm for contextual recommendation with logarithmic regret and robustness. arXiv:2603.20826.
  28. スケーリングの直感とポテンシャル Perceptron −gt 33 / 46 Second-order Perceptron −A−1 t−1

    gt −gt ĉt At = 過去の 解残差 ĉt Pt ⊤ s=1 gs gs + λI の逆行列スケーリングは過去の解残差累積が少ない成分を重視 証明で比較する二つのポテンシャル 累積ポテンシャル ζt , A−1 t ζt 楕円ポテンシャル t X s=1 gs , A−1 s gs
  29. ポテンシャル比較不等式 34 / 46 補題:ポテンシャル比較不等式 ζT , A−1 T ζT

    ≤ T X gt , A−1 t gt t=1 証明の要点:学習期の出力の最適性と半正定値性から D E alg obs νt := gt , A−1 ζ = ĉ , x − x ≤ 0, µt := gt , A−1 t t t t−1 t−1 t−1 gt ≥ 0. Sherman–Morrison 公式と符号条件から µt + 2νt − νt2 µt ≤ = gt , A−1 t gt . 1 + µt 1 + µt 各時刻について総和を取り,ζ0 = 0 から補題の不等式を得る. −1 ζt , A−1 t ζt − ζt−1 , At−1 ζt−1 =
  30. 楕円ポテンシャル補題 35 / 46 楕円ポテンシャル補題 (Elliptical Potential Lemma) 定理の仮定(および d,

    T ≥ 2)のもとで T X gt , A−1 t gt ≤ d log T t=1 スカラーの場合の直感:λ = 1, gt = 1 ならば,各時刻で t X At = 1 + gs2 = t + 1, gt , A−1 t gt = s=1 この場合楕円ポテンシャルは T T X X −1 gt , A t gt = t=1 1 . t+1 1 = O(log T ). t+1 t=1
  31. O(d log T ) Regret の証明 36 / 46 ∥v∥2A

    := ⟨v, Av⟩ とする.Cauchy–Schwarz と二つの補題から v u T q uX p −1 RT (c) = −⟨c, ζT ⟩ ≤ ∥c∥AT ζT , AT ζT ≤ ∥c∥AT t gt , A−1 d log T . t gt ≤ ∥c∥AT t=1 obs xalg の最適性,∥c∥ ≤ 1, diam(Xt ) ≤ D から,⟨c, xalg t t − xt ⟩ ∈ [0, D].よって ! T T X X 2 ⊤ ⊤ 2 obs ∥c∥AT = c λI + gt gt c ≤ λ∥c∥ + D ⟨c, xalg t − xt ⟩ = λ + DRT (c). t=1 t=1 RT (c) ≥ 0 についてのニ次不等式を解くと (RT (c))2 ≤ (λ + DRT (c))d log T =⇒ RT (c) ≤ Dd log T + p λd log T .
  32. 非最適フィードバック下での Regret 上界 37 / 46 [Sakaue 2026] 各時刻の観測 xobs

    の非最適度とその累積を以下で定義: t δt (c) = ⟨c, xobs t ⟩ − min ⟨c, x⟩, x∈Xt ∆T (c) = T X δt (c). t=1 定理:非最適フィードバック下での Regret 上界 xobs ∈ Xt ,diam(Xt ) ≤ D = O(1),∥c∥ ≤ 1 ならば,同一のアルゴリズムで t   p RT (c) = O d log T + ∆T (c) d log T . 特に ∆T (c) = 0 なら RT (c) = O(d log T ). Sakaue (2026). Simple projection-free algorithm for contextual recommendation with logarithmic regret and robustness. arXiv:2603.20826.
  33. まとめ 39 / 46 1.目的関数パラメータ観測からのオフライン学習 二段階法・DiffOpt の比較 MSE は悪化しても Regret

    が改善するケースがある 2.観測解からのオフライン学習 逆最適化で観測欠損を補う手法:パラメータ観測の欠損を補完して回帰 Suboptimality loss を用いる手法:解残差に基づく更新 3.逆最適化とオンライン学習 √ 解残差に基づく Perceptron 型更新による O( T ) Regret 上界 Second-order Perceptron に基づく手法による O(d log T ) Regret 上界
  34. 関連文献 40 / 46 サーベイ・ベンチマーク DFL, 文脈付き最適化のサーベイ [Mandi et al.

    2024; Sadana et al. 2025] 逆最適化のモデル分類・方法論・応用 [Chan et al. 2025] 大規模なベンチマーク比較 [Geng et al. 2024] 制約も考慮した拡張 最適解から LP の目的係数・制約係数を共同学習 [Tan et al. 2020] 既知目的と複数観測による実行可能領域の推定 [Ghobadi and Mahmoudzadeh 2021] MILP の制約を先に学習し,その後に目的関数を学習 [Kitaoka 2025] 制約パラメータ予測における実行可能性と意思決定品質の両立 [Mandi et al. 2025] 最近の発展 線形最適化に対する訓練時ソルバー不要の DFL [Berden et al. 2025; Wan and Liu 2026] Attention 機構を用いて期待目的関数を直接学習する DFL [Kong et al. 2025] Transformer と制約推論による実行可能解の直接予測 [Navarro et al. 2026]
  35. 参考文献 (1/5) 41 / 46 1 Agrawal, A., Amos, B.,

    Barratt, S., Boyd, S., Diamond, S., and Kolter, J. Z. (2019). Differentiable convex optimization layers. Advances in Neural Information Processing Systems, 32, 9558–9570. 2 Amos, B., and Kolter, J. Z. (2017). OptNet: Differentiable optimization as a layer in neural networks. Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 136–145. 3 Bärmann, A., Pokutta, S., and Schneider, O. (2017). Emulating the expert: Inverse optimization through online learning. Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 400–410. 4 Bärmann, A., Martin, A., Pokutta, S., and Schneider, O. (2018; rev. 2020). An online-learning approach to inverse optimization. arXiv:1810.12997. 5 Berden, S., Mahmutoğulları, A. İ., Tsouros, D., and Guns, T. (2025). Solver-free decision-focused learning for linear optimization problems. Advances in Neural Information Processing Systems, 38, 158127–158145. 6 Besbes, O., Fonseca, Y., and Lobel, I. (2021). Online learning from optimal actions. Proceedings of the 34th Conference on Learning Theory, PMLR 134, 586–586. 7 Besbes, O., Fonseca, Y., and Lobel, I. (2025). Contextual inverse optimization: Offline and online learning. Operations Research, 73(1), 424–443. DOI: 10.1287/opre.2021.0369.
  36. 参考文献 (2/5) 42 / 46 8 Blondel, M., Martins, A.

    F. T., and Niculae, V. (2020). Learning with Fenchel–Young losses. Journal of Machine Learning Research, 21(35), 1–69. 9 Blondel, M., Berthet, Q., Cuturi, M., Frostig, R., Hoyer, S., Llinares-López, F., Pedregosa, F., and Vert, J.-P. (2022). Efficient and modular implicit differentiation. Advances in Neural Information Processing Systems, 35, 5230–5242. DOI: 10.52202/068431-0378. 10 Chan, T. C. Y., Mahmood, R., and Zhu, I. Y. (2025). Inverse optimization: Theory and applications. Operations Research, 73(2), 1046–1074. DOI: 10.1287/opre.2022.0382. 11 Elmachtoub, A. N., and Grigas, P. (2022). Smart “Predict, then Optimize”. Management Science, 68(1), 9–26. DOI: 10.1287/mnsc.2020.3922. 12 Geng, H., Ruan, H., Wang, R., Li, Y., Wang, Y., Chen, L., and Yan, J. (2024). Benchmarking PtO and PnO methods in the predictive combinatorial optimization regime. Advances in Neural Information Processing Systems, 37, 65944–65971, Datasets and Benchmarks Track. DOI: 10.52202/079017-2108. 13 Ghobadi, K., and Mahmoudzadeh, H. (2021). Inferring linear feasible regions using inverse optimization. European Journal of Operational Research, 290(3), 829–843. DOI: 10.1016/j.ejor.2020.08.048. 14 Gollapudi, S., Guruganesh, G., Kollias, K., Manurangsi, P., Paes Leme, R., and Schneider, J. (2021). Contextual recommendations and low-regret cutting-plane algorithms. Advances in Neural Information Processing Systems, 34, 22498–22508. 15 Hikima, Y., and Kamiyama, N. (2025). An inverse optimization approach to contextual inverse optimization. Proceedings of the 34th International Joint Conference on Artificial Intelligence, Main Track, 5354–5362. DOI: 10.24963/ijcai.2025/596.
  37. 参考文献 (3/5) 43 / 46 16 Hikima, Y., Kamiyama, N.,

    Sakaue, S., and Tsuchiya, T. (2026). A sampling-based relaxation approach to contextual inverse optimization. 35th International Joint Conference on Artificial Intelligence. 17 Kitaoka, A. (2024; rev. 2026). Exact solution to data-driven inverse optimization of MILPs in finite time via gradient-based methods. arXiv:2405.14273. 18 Kitaoka, A. (2025; rev. 2026). Inverse mixed-integer programming: Learning constraints then objective functions. arXiv:2510.04455. 19 Kong, L., Mu, W., Cui, J., Zhuang, Y., Prakash, B. A., Dai, B., and Zhang, C. (2025). DF 2 : Distribution-free decision-focused learning. Proceedings of the 41st Conference on Uncertainty in Artificial Intelligence, PMLR 286, 2269–2290. 20 Lu, H., Peng, Z., and Yang, J. (2025). MPAX: Mathematical programming in JAX. 39th Conference on Neural Information Processing Systems, Workshop on GPU-Accelerated and Scalable Optimization. arXiv:2412.09734. 21 Mandi, J., Kotary, J., Berden, S., Mulamba, M., Bucarey, V., Guns, T., and Fioretto, F. (2024). Decision-focused learning: Foundations, state of the art, benchmark and future opportunities. Journal of Artificial Intelligence Research, 80, 1623–1701. DOI: 10.1613/jair.1.15320. 22 Mandi, J., Defresne, M., Berden, S., and Guns, T. (2025). Feasibility-aware decision-focused learning for predicting parameters in the constraints. Advances in Neural Information Processing Systems, 38, 163285–163309.
  38. 参考文献 (4/5) 44 / 46 23 Mishra, S. K., Raj,

    A., and Vaswani, S. (2024). From inverse optimization to feasibility to ERM. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, 35805–35828. 24 Mohajerin Esfahani, P., Shafieezadeh-Abadeh, S., Hanasusanto, G. A., and Kuhn, D. (2018). Data-driven inverse optimization with imperfect information. Mathematical Programming, 167(1), 191–234. DOI: 10.1007/s10107-017-1216-6. 25 Navarro, M., van Hoeve, W.-J., and Singh, K. (2026). Inverse optimization without inverse optimization: Direct solution prediction with Transformer models. arXiv:2602.05306. 26 Oki, T., and Sakaue, S. (2026). Finite and corruption-robust regret bounds in online inverse linear optimization under M-convex action sets. Proceedings of the 43rd International Conference on Machine Learning, to appear. arXiv:2602.01682. 27 Sadana, U., Chenreddy, A., Delage, E., Forel, A., Frejinger, E., and Vidal, T. (2025). A survey of contextual optimization methods for decision-making under uncertainty. European Journal of Operational Research, 320(2), 271–289. DOI: 10.1016/j.ejor.2024.03.020. 28 Sakaue, S., Tsuchiya, T., Bao, H., and Oki, T. (2025a). Online inverse linear optimization: Efficient logarithmic-regret algorithm, robustness to suboptimality, and lower bound. Advances in Neural Information Processing Systems, 38, 85189–85217. 29 Sakaue, S., Bao, H., and Tsuchiya, T. (2025b). Revisiting online learning approach to inverse linear optimization: A Fenchel–Young loss perspective and gap-dependent regret analysis. Proceedings of the 28th International Conference on Artificial Intelligence and Statistics, PMLR 258, 46–54.
  39. 参考文献 (5/5) 45 / 46 30 Sakaue, S. (2026). Simple

    projection-free algorithm for contextual recommendation with logarithmic regret and robustness. arXiv:2603.20826. 31 坂上 晋作 (2025). パラメータ観測が不要な「予測→最適化」 ~機械学習×数理最適化のもう一つの アプローチ~. Zenn. URL: zenn.dev. 32 高橋 優輝 (2025). JAXopt で実践する Decision-focused Learning ~機械学習×数理最適化の新たな アプローチ~. NTT ドコモ R&D ENGINEERING BLOG. URL: nttdocomo-developers.jp. 33 Tan, Y., Terekhov, D., and Delong, A. (2020). Learning linear programs from optimal decisions. Advances in Neural Information Processing Systems, 33, 19738–19749. 34 Troutt, M. D., Pang, W.-K., and Hou, S.-H. (2006). Behavioral estimation of mathematical programming objective function coefficients. Management Science, 52(3), 422–434. DOI: 10.1287/mnsc.1050.0445. 35 Wan, B., and Liu, M. (2026). A solver-free training method for predict-then-optimize. 43rd International Conference on Machine Learning. arXiv:2606.19587. OpenReview: Tz69nG4l87. 36 Wilder, B., Dilkina, B., and Tambe, M. (2019). Melding the data-decisions pipeline: Decision-focused learning for combinatorial optimization. Proceedings of the AAAI Conference on Artificial Intelligence, 33(1), 1658–1665. DOI: 10.1609/aaai.v33i01.33011658.
  40. 補足:SPO+:意思決定損失の凸上界 SPO 損失: ℓSPO (ĉ; c) = c⊤ x⋆ (ĉ)

    − c⊤ x⋆ (c) SPO+:SPO 損失に対する凸代理損失 最小化問題に対する SPO+損失 ℓSPO+ (ĉ; c) = max(c − 2ĉ)⊤ x + 2ĉ⊤ x⋆ (c) − c⊤ x⋆ (c) x∈X • ℓSPO+ は ℓSPO の凸上界 • 劣勾配:2(x⋆ (c) − x⋆ (2ĉ − c)) • ただし,真の係数ベクトル c の観測が必要 46 / 46