Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
【ECCV2026:Genception】映像生成に秘められた力
Search
小島瑞貴
October 07, 2026
100
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
【ECCV2026:Genception】映像生成に秘められた力
小島瑞貴
October 07, 2026
More Decks by 小島瑞貴
See All by 小島瑞貴
視覚若手の会LENSって何??
mickey_0226
0
330
画像生成モデルは1つの損失関数(FID)を入れるだけで性能が激変するらしい
mickey_0226
0
1.7k
【Zozo Research 技術共有会】三次元領域の現在と展望
mickey_0226
3
670
学術バーQってどんなところ??
mickey_0226
0
200
さわって動かす人工知能
mickey_0226
0
92
動画生成と三次元生成を融合して最強の生成モデルを作ろう
mickey_0226
0
73
CVPR2026_VGGTとその仲間たち
mickey_0226
0
1.2k
Transformerの推論を線形時間にして皆を驚かせましょう
mickey_0226
0
76
Featured
See All Featured
GitHub's CSS Performance
jonrohan
1033
470k
Balancing Empowerment & Direction
lara
6
1.3k
実際に使うSQLの書き方 徹底解説 / pgcon21j-tutorial
soudai
PRO
203
76k
コードの90%をAIが書く世界で何が待っているのか / What awaits us in a world where 90% of the code is written by AI
rkaga
63
46k
Agile Actions for Facilitating Distributed Teams - ADO2019
mkilby
0
300
技術選定の審美眼(2025年版) / Understanding the Spiral of Technologies 2025 edition
twada
PRO
120
120k
Organizational Design Perspectives: An Ontology of Organizational Design Elements
kimpetersen
PRO
1
840
Measuring Dark Social's Impact On Conversion and Attribution
stephenakadiri
2
300
What Being in a Rock Band Can Teach Us About Real World SEO
427marketing
0
1.1k
The Director’s Chair: Orchestrating AI for Truly Effective Learning
tmiket
1
320
Self-Hosted WebAssembly Runtime for Runtime-Neutral Checkpoint/Restore in Edge–Cloud Continuum
chikuwait
0
860
Impact Scores and Hybrid Strategies: The future of link building
tamaranovitovic
0
450
Transcript
名古屋ECCV2026読み会: Genception 東京科学大学 小島 瑞貴
自己紹介 【名前】 小島 瑞貴 【所属】東京科学大学 博士1年 【課外活動】 運営 視覚若手の会LENS 代表
2 幹事
論文情報 Genception: Video Generation Models are General-Purpose Vision Learners Letian
Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu % 3
問題設定 汎用的な視覚モデルの構築 VGGT(三次元復元) SAM3(セグメンテーション) % 高い性能だが,両タスクを「統一的に」扱えない 4
背景: 言語領域での基盤モデル 翻訳 言語基盤モデル 要約 統一 % 種々の言語タスクを基盤モデルが統一! 5
背景: 言語モデルの学習手続き ❶事前学習 ❷Supervised Fine-Tuning 視覚基盤モデルを作りたい…! 6 ❸選好最適化
前提: 視覚基盤モデルに必要な要件 ❶時空間理解 ❷言語との接続 ❸スケール可能性 これって,映像拡散モデルそのものでは? 7
前提: 事前学習としての映像拡散モデル ❶時空間理解 ❷言語との接続 ❸スケール可能性 ❶映像生成の特性 ❷テキスト条件付け ❸テキスト-画像ペアのみ 8
提案: Genception ❶ 映像拡散モデルで,事前学習 ❷ 様々な下流タスクで,事後学習 9
提案手法: 下流タスク一覧 奥行き・法線マップ・カメラ姿勢などなど ポイント: 全部RGB形式に統一された出力(キーポイントを除く;後述) 10
提案手法: 事後学習のやりかた 要点 ❶ ノイズなし(t=0)を入力→1-ステップ出力 ❷ タスク定義は,テキストプロンプトにて(最重要ポイント) 【概要図】 11
提案手法: 事後学習のやりかた 要点 ❶ ノイズなし(t=0)を入力→1-ステップ出力 ❷ タスク定義は,テキストプロンプトにて(最重要ポイント) 【概要図】 拡散モデル 12
提案手法: 事後学習のやりかた 要点 ❶ ノイズなし(t=0)を入力→1-ステップ出力 ❷ タスク定義は,テキストプロンプトにて(最重要ポイント) 【概要図】 拡散モデル タスク定義はここ!(例:デプス)
13
定量評価 デプス推定 カメラ姿勢推定 前景抽出 セグメンテーション キーポイント 法線推定 単一モデルで全部(ほぼ)SOTA 14
定性評価
定性評価(セグメンテーション)
定性評価(キーポイント)
定性評価(言語対応付け)
まとめ ★論文名: Genception: Video Generation Models areGeneral-Purpose Vision Learners ★目的:
言語のような,汎用的な視覚基盤モデルを作りたい ★言語側の学習: ❶事前学習❷ファインチューニング❸選好最適化 ★視覚モデルに必要な要件: ❶時空間理解 ❷ 言語との接続 ❸スケール可能性 ★Genception(本提案手法): ❶映像拡散モデルで事前学習 ❷様々な下流タスクで事後学習 ★下流タスク: 奥行き・法線・キーポイント・カメラ姿勢などなど ★事後学習のやりかた: ❶ノイズなし1ステップ出力❷言語でタスク指示 ★評価: すべてのタスクで(ほとんど)最高性能