Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
[RSJ26] AnoleVLA: Lightweight Vision-Language-A...
Search
Semantic Machine Intelligence Lab., Keio Univ.
PRO
August 29, 2026
Technology
21
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
[RSJ26] AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
Semantic Machine Intelligence Lab., Keio Univ.
PRO
August 29, 2026
More Decks by Semantic Machine Intelligence Lab., Keio Univ.
See All by Semantic Machine Intelligence Lab., Keio Univ.
[RSJ26] Building a VLA Model Based on Self-Distilled Classification
keio_smilab
PRO
0
28
[RSJ26] NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields
keio_smilab
PRO
0
18
[RSJ26] Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields
keio_smilab
PRO
0
87
[RSJ26] Hierarchy-Aware Multimodal Retrieval-Augmented Generation for Embodied Question Answering
keio_smilab
PRO
0
44
[MIRU26] Open-Vocabulary Intention-Guided Object Detection in Diverse Scenes
keio_smilab
PRO
0
180
[Journal club] Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
keio_smilab
PRO
0
42
[MIRU26] To What Extent Does MLLM-as-a-Judge Exhibit Cross-Model Preference Bias?
keio_smilab
PRO
0
260
[Journal club] FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
keio_smilab
PRO
0
32
[Journal club] DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
keio_smilab
PRO
1
66
Other Decks in Technology
See All in Technology
【5分でわかる】セーフィー エンジニア向け会社紹介
safie_recruit
0
54k
指示待ちから変化に応じるClaude Codeへ!~環境からAgentへの帰り道を作る~
gotalab555
6
970
Data Hubグループ 紹介資料
sansan33
PRO
0
3.2k
パスキーでドライブする アカウント統合(OAuth/OIDC Numa (Immersion) Workshop 2026)
oidfj
PRO
0
330
markdown-poster Introduction
kazamori
0
330
AI時代の「OAuth認証」にどう物申すか?(OAuth/OIDC Numa (Immersion) Workshop 2026)
oidfj
PRO
0
310
Introduction to Bill One Development Engineer
sansan33
PRO
0
470
全社に広がるMCPサーバーを、 どう安全に管理するか MCPass開発の舞台裏
mtpooh
2
140
EMの役割で 変わったこと・変わらなかったこと
sansantech
PRO
0
180
サービス内で複数のOP・ASを連鎖させる(OAuth/OIDC Numa (Immersion) Workshop 2026)
oidfj
PRO
0
320
AIエージェントのためのデータ設計
daiz21
0
460
All About Sansan – for New Global Engineers
sansan33
PRO
1
1.5k
Featured
See All Featured
Navigating the Design Leadership Dip - Product Design Week Design Leaders+ Conference 2024
apolaine
2
400
Measuring Dark Social's Impact On Conversion and Attribution
stephenakadiri
2
260
The Curse of the Amulet
leimatthew05
2
14k
SEOcharity - Dark patterns in SEO and UX: How to avoid them and build a more ethical web
sarafernandez
0
250
Building AI with AI
inesmontani
PRO
1
1.2k
AI: The stuff that nobody shows you
jnunemaker
PRO
9
940
Game over? The fight for quality and originality in the time of robots
wayneb77
1
260
Noah Learner - AI + Me: how we built a GSC Bulk Export data pipeline
techseoconnect
PRO
0
400
Breaking role norms: Why Content Design is so much more than writing copy - Taylor Woolridge
uxyall
1
380
Responsive Adventures: Dirty Tricks From The Dark Corners of Front-End
smashingmag
254
22k
How to Get Subject Matter Experts Bought In and Actively Contributing to SEO & PR Initiatives.
livdayseo
0
180
Have SEOs Ruined the Internet? - User Awareness of SEO in 2025
akashhashmi
0
450
Transcript
深層状態空間モデルに基づく 軽量VLAによる物体操作 慶應義塾大学 髙木裕輔 神原元就 八島大地 妹尾幸樹 戸倉健登 杉浦孔明
Motivation: VLAのメモリ消費量・推論時間を削減したい 背景: VLAを実機で動作させる際、推論時の計算コストが問題 メモリ消費量が大きく高性能な計算機が必要 軌道生成の研究 [Kambara+, RA-L26], [Kaichi+,
IROS26],... 推論時間が長くアームの動作がjerkyに 本研究: 軽量かつ高速なVLAを提案 “Pick up the cube and put it into the basket.” 推論時間 [ms] OpenVLA jerkyな動作 GPUメモリ消費 17GB以上 𝜋0 𝜋0.5 提案手法 𝜋0.5 推論時の動作 [Jiaming+, 26] GPUメモリ消費量 [GB] 2
関連研究: VLAの計算コストを抑制する研究 手法 特徴 SmolVLA 0.45Bの軽量VLA ・action chunkingを活用 Transformerに基づくバックボーン・長系列の処理が困難
深層状態空間モデルに基づくバックボーン 画像上の接触点 / 姿勢予測のみ・軌道を生成しない [Shukor+, 25] RoboMamba [Liu+, NeurIPS24] SmolVLA RoboMamba 3
提案手法: 深層状態空間モデルに基づく軽量なVLA 4
提案手法: 深層状態空間モデルに基づく軽量なVLA 入力を埋め込みトークン系列を生成 ロボット状態 ロボット状態の時間差分 画像 指示文 5
提案手法: 深層状態空間モデルに基づく軽量なVLA バックボーンにて軌道生成に必要な情報を集約 入力系列 バックボーンLLM 出力系列 6
提案手法: 深層状態空間モデルに基づく軽量なVLA 最終トークンを利用しチャンク長の軌道を生成 軌道 最終トークンに 情報が集約 7
高速かつメモリ消費量の少ないバックボーンLLM ▪ Mamba [Gu+, COLM24] による系列処理 隠れ状態 入力 約6Mの学習可能パラメータ 出力
ブロックの多層化 ☺ 𝒪(𝑁)での系列処理 cf. Transformerの計算量は𝒪(𝑁 2 ) ☺ 370Mパラメータの軽量モデル cf. RoboMamba: 2.8B 8
実験設定: シミュレーション・実機ロボットにおける実験 ▪ シミュレーション実験: Meta-World [Yu+, CoRL19] ▪ 実機実験: モバイルマニピュレーションタスク(HSRを使用)
▪ リーダ・フォロワシステムを用い、データを収集 タスク数 エピソード数 試行回数/タスク シミュレーション 50 2,500 10 実機 5 250 10 Meta-Worldのタスク例 実機実験のデータ収集 9
定量的結果: シミュレーションにてベースラインを上回った 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 10
定量的結果: シミュレーションにてベースラインを上回った 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 平均成功率 提案手法 +22pt 11
定性的結果: シミュレーションで様々なタスクに成功 "Grasp a stick and pull a box with
the stick." "Sweep a puck off the table." 提案手法 ベースライン手法 提案手法 ベースライン手法 ☺ 目標位置に物体を移動 不適切な位置で停止 ☺ 物体を把持し移動 物体を把持できず失敗 12
実機実験: モバイルマニピュレーションタスクは困難 ベースの回転・移動に伴う視点変化があり困難 ”Insert the lemon into the cup.
” 提案手法 事前学習データセットに モバイルタスクが少ない 2.4% 𝜋0.5 x2 その他 97.6% 𝜋0.5 の事前学習データセット 13
定量的結果: モバイルマニピュレーションにてベースラインを上回る 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 100 100 Mobile
pick Mobile move Mobile open Mobile push 100 100 Mobile close 14
定量的結果: モバイルマニピュレーションにてベースラインを上回る 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 100 100平均成功率 100
100 +11pt Mobile pick Mobile move Mobile open Mobile push Mobile close 15
定量的結果: モバイルマニピュレーションにてベースラインを上回る 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 100 100平均成功率 100
100 +21pt Mobile pick Mobile move Mobile open Mobile push Mobile close 16
まとめ: 深層状態空間モデルに基づく軽量なVLA ▪ 背景:VLAにおいて推論時のメモリ消費量・推論時間が課題 ▪ 新規性:深層状態空間モデルに基づく軽量なバックボーン ▪ 結果:シミュレーション・実機のタスク成功率でベースライン手法を上回った LLMの事前学習知識を活用するため蒸留を用いた軽量VLA ◼
9/4 13:44~ @大会議室A “Self-Distilled Classificationに基づくVLAの構築” 17
Appendix
定量的結果:高速かつ推論時のメモリ消費量が少ない 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 19
定量的結果:高速かつ推論時のメモリ消費量が少ない 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 1 × 10 1
× 3 20
Ablation Study: 加速度損失・DeepSSMが性能向上に寄与 w/o 加速度損失 DeepSSM → Transformer 提案手法 21
加速度に対する損失関数を導入した2段階訓練 ▪ 軌道の時間差分を効率的に捉えるため2段階訓練を導入 ▪ Stage1: エンドエフェクタの速度に対するL1損失 ▪ Stage2: 加速度に対するL1損失. .
を追加 22
エラー分析: 対象物体の位置の認識が困難 ▪ シミュレーションの失敗例20例に対してエラー分析を実施 ▪ 物体位置の認識エラー: 空間的に誤った位置にアームを移動させる ▪ 把持点推定の失敗: 適切な位置で把持できず物体を落とす
▪ 動作の未完了: 動作途中でアームが停止する エラーカテゴリ エラー数 物体位置の認識エラー 10 把持点推定の失敗 6 動作の未完了 4 合計 20 23
実機実験:提案手法の失敗例 早くグリッパを閉じる アームでボトルを倒す 24