Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
論文読み会 ICLR2019 | Attention, Learn to Solve Rout...
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
cocomoff
May 21, 2020
Research
950
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
論文読み会 ICLR2019 | Attention, Learn to Solve Routing Problems!
cocomoff
May 21, 2020
More Decks by cocomoff
See All by cocomoff
論文読み会 NeurIPS2024 | UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction
cocomoff
1
140
論文読み会 AMAI | Personalized choice prediction with less user information
cocomoff
0
100
論文読み会 KDD2024 | Relevance meets Diversity: A User-Centric Framework for Knowledge Exploration through Recommendations
cocomoff
0
290
論文読み会 KDD2022 | Multi-Behavior Hypergraph-Enhanced Transformer for Sequential Recommendation
cocomoff
0
190
論文読み会 AISTATS2024 | Deep Learning-Based Alternative Route Computation
cocomoff
0
84
論文読み会 AAAI2021 | Knowledge-Enhanced Top-K Recommendation in Poincaré Ball
cocomoff
0
160
論文読み会 WWW2022 | Learning Probabilistic Box Embeddings for Effective and Efficient Ranking
cocomoff
0
370
ClimaX: A foundation model for weather and climate
cocomoff
0
680
論文読み会 AAAI2022 | MIP-GNN: A Data-Driven Framework for Guiding Combinatorial Solvers
cocomoff
0
300
Other Decks in Research
See All in Research
XDPerf: A High-Performance Traffic Generator Built with WASM and eBPF
takehaya
1
290
[CV勉強会@関東 CVPR2026] PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow / kantocv 67th CVPR 2026
shunk031
0
290
Sleuthcon Keynote - How Cybercriminals (ab)use AI
fr0gger
0
320
[BlackHatAsia2026] Hidden Telemetry: Uncovering TraceLogging ETW Providers You're Not Using (Yet)
asuna_jp
1
700
CVPR2026論文紹介_VLMにとって良いvision encoderとは何か?Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
kobayashi31
1
210
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
satai
3
260
某助成金プロジェクト採択に向けて企業研究所のアウトリーチ専任者がやったこと
afroscript
0
180
J-STAGEの現況と全文XML登載必須化について
xspa2012
0
200
Harness Engineering and Al Agent
kzinmr
3
1.9k
SAM3を用いたコマ・吹き出しの 領域検出と分割構造からの読み順推定
kzmssk
0
110
Anthropic が提案する LLM の内部状態を自然言語で説明可能にした Natural Language Autoencoders / Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
shunk031
0
210
PHTalks Bengaluru - SSRF When All Else Fails
dk999
0
1.1k
Featured
See All Featured
WENDY [Excerpt]
tessaabrams
12
39k
Bioeconomy Workshop: Dr. Julius Ecuru, Opportunities for a Bioeconomy in West Africa
akademiya2063
PRO
1
340
How to Grow Your eCommerce with AI & Automation
katarinadahlin
PRO
1
260
Building a A Zero-Code AI SEO Workflow
portentint
PRO
0
710
The Power of CSS Pseudo Elements
geoffreycrofte
82
6.5k
Future Trends and Review - Lecture 12 - Web Technologies (1019888BNR)
signer
PRO
0
3.7k
The AI Revolution Will Not Be Monopolized: How open-source beats economies of scale, even for LLMs
inesmontani
PRO
3
3.7k
The MySQL Ecosystem @ GitHub 2015
samlambert
251
13k
Code Reviewing Like a Champion
maltzj
528
40k
Being A Developer After 40
akosma
91
590k
jQuery: Nuts, Bolts and Bling
dougneiner
66
8.6k
brightonSEO & MeasureFest 2025 - Christian Goodrich - Winning strategies for Black Friday CRO & PPC
cargoodrich
3
810
Transcript
May 21, 2020 @cocomoff
概要 これまで専⽤アルゴリズムで解いていた最適化問題 (TSP/VRP/OP/PCTSP) をPointer-Network を拡張したNN で解く PyTorch 版 https://github.com/wouterkool/attention-learn-to-route モデルの基本的なアイデア
⼊⼒はTSP の地点座標 ( 地点数 のtorch.tensor) Encoder-decoder 型のモデルで次に訪問すべき地点を決定 BN とmulti-head attention を使って隠れ層に埋め込む コンテキスト ( 今何地点訪れたかを考慮して を拡張) を計算 適当に出⼒をclip してsoftmax に⼊れて地点を決定する 学習の基本的なアイデア 強化学習界隈で基本的な⼿法であるベースラインつきREINFORCE ベースラインとしてgreedy な経路を使う 1/12 n × 2 h
実験 | 環境設定 地点数 について各問題を解いた 学習時にたくさんインスタンスをつくる (100,000 とか?) Every epoch
we process 2500 batches of 512 instances e.g., で5:30/epoch, で16:20/epoch の訓練時間 100 epochs 学習してから10000 個のテストインスタンスで評価 学習率は定数 ( ) だけど,適当にdecay した⽅が安定した Encoder は3 層 Decoder greedy 毎回最良の⾏動を洗濯して解をつくる sampling 1280 解をサンプルして,最良を選ぶ 既存⼿法は3 パターン ( 専⽤exact/heuristics ,既存のNN ⼿法) できるだけ環境を揃えて実⾏して⽐較した ( とのこと) 2/12 n = 20, 50, 100 n = 20 n = 50 η = 10−4
実験 | TSP 3/12
実験 | CVRP/SDVRP CVRP はTSP の容量付き+ 復数⾞両版 SDVRP はルートを分割しても良いタイプのTSP 3/12
実験 | OP/PCTSP OP はmin. cost ではなくmax. profit 型の問題 PCTSP/SPCTSP
は訪問しなくてもいいタイプのTSP ( ペナルティ付き) 4/12
背景 ( 既存の研究アプローチ) (1/3) NN で最適化するアプローチはHopfield&Tank (1985) ぐらいからある Pointer Network
(Vinyals et al. NIPS2015) 左: seq2seq で直接頂点番号を出⼒するアプローチ 右: PN .Decoder 側にも⼊⼒データの特徴 ( 座標) を⼊れ,凸包の頂 点を指し⽰すようなAttention を作る 5/12
背景 ( 既存の研究アプローチ) (2/3) Bello et al. 2016 (ICLR2017 WS)
PN + Actor-Critic Actor-Critic: ⾏動する側(Actor) と⾏動を評価する側(Critic) を同時学 習するタイプの⼿法.実際にはREINFORCE( 評価は状態⾏動価値を 学習ではなく,報酬の平均で⾏う) Nazari et al. NeurIPS2018 VRP を解くようにPN (LSTM を変更) 6/12
背景 ( 既存の研究アプローチ) (3/3) Daiet al. NIPS2017, Nowak et al.
2017, Kaempfer et al. 2018 Decoder-Encoder 型ではなく,1 つのモデル (GNN とか) で解く 他にもTransformer-inspired なモデルとかもある Deudon et al. CPAIOR2018 2OPT local search という探索⼿法をアテンションで学習して再現 これはBello et al. の追試っぽい感じ ( 報酬が2OPT-based) これぐらいのが解ける 7/12
提案⼿法 (1/4) | Encoder 次の確率をモデル化したい ( その後サンプリングして解を作成): は問題イ ンスタンス (e.g.,
2 次元座標) , は出⼒の順列: Encoder 8/12 s π
は座標 を線形で埋め込み ( ) は 個のMulti-Head Attention を適⽤して作成 h(0) x
h = (0) W x + x b h(i) M
Attention (Dot-product attention (?)) Multi-head attention (MHA) 別のパラメータを作って 本アテンションを作り,線形結合 FF
(Feed-forward sublayer) 線形変換してReLU (1 層⽬以外) BN (Batch Normalization) 例のあれ M
提案⼿法 (2/4) | Decoder 埋め込んだ を使い,順列の⽣成を⾏う ⼤まかな構造は普通のDecoder と同じ Decoder への⼊⼒として,埋め込んだ
だけじゃなく,「最初の地点」と 「1 つ前の地点」を使う ( これをcontext embedding と呼ぶ.図の◦3 つ) 事前に訪問した地点は訪問しないので,mask で-∞ にする 3124 を⽣成する図 9/12 h h
提案⼿法 (3/4) | 学習 作ったNN はインスタンス から を⽣成するモデル 期待コスト =
loss をgrad. descent する REINFORCE . はNN でモデル化.ベースライン は このサンプル評価値からのgrad. descent を安定させる. いろいろな⼿法で を⼊れてもよいが,「インスタンスにアルゴリズ ムを適⽤してみたら,難しさは評価できるだろう」という期待 ある地点のパラメータ を固定して,greedy rollout を作成 greedy rollout より良い順回路が⾒つかれば,報酬が伝わって もしパラメータに優位な差ができたら,更新する 10/12 s π p (π ∣ θ s) L(θ ∣ s) = E [L(π)] p (π∣s) θ ∇ log p (π ∣ θ s) b(s) b(s) θ
提案⼿法 (4/4) | 学習 5 ⾏⽬でランダムインスタンスを作成 6 ⾏⽬〜7 ⾏⽬で現在と今設定されているベースラインをやってみる 8
⾏⽬〜9 ⾏⽬で誤差を評価し,学習を⾏う 11 ⾏⽬〜13 ⾏⽬で適当にパラメータを更新する 11/12
ベースラインの⽐較 提案⼿法(AM) とPointer Network(PN) の⽐較 ロス計算に⽤いているベースライン を3 つ変えた場合の挙動 Rollout と書いてあるやつが実験で使っているGreedy
Rollout Critic は をNN で推定して使う (NN 中⾝はEncoder と似ている) Exponential は計算したロスをexp でdecay させるbaseline 12/12 b(s) V (s)