Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
TransGAN: Two Transformers Can Make One Strong GAN
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
kiyo
April 18, 2021
Technology
390
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
TransGAN: Two Transformers Can Make One Strong GAN
第六回全日本コンピュータビジョン研究会 Transformer読み会での発表資料です
kiyo
April 18, 2021
More Decks by kiyo
See All by kiyo
Agent Skill Acquisition for Large Language Models via CycleQD
kiyohiro8
0
47
Active Retrieval Augmented Generation
kiyohiro8
3
970
Reinforcement Learning: An Introduction 輪読会 第5回
kiyohiro8
0
480
Reinforcement Learning: An Introduction 輪読会 第3回
kiyohiro8
0
650
CycleGAN and InstaGAN
kiyohiro8
0
1.6k
Bridging_by_Word__Image-Grounded_Vocabulary_Construction_for_Visual_Captioning.pdf
kiyohiro8
0
1k
Attention on Attention for Image Captioning
kiyohiro8
1
570
Progressive Growing of GANs for Improved Quality, Stability, and Variation
kiyohiro8
1
190
Graph-Based Global Reasoning Networks
kiyohiro8
0
1.4k
Other Decks in Technology
See All in Technology
2026-09-26 Platform Engineering Kaigi 2026 インフラとアプリの境界線と委譲の設計 / Drawing the Infra and App Line
masasuzu
0
520
Codex概要
ymiya55
0
110
[Kiro Meetup #7] Kiro Crew Dive Deep
konippi
0
430
TiDBファミリーにDWHが新登場!! TiDB最新情報 / TiDB update 202609
yoshiakiyamasaki
0
160
顧客の成果創出とプロダクトの成長を 両立するためのFDE
sansantech
PRO
0
530
IR Today: Theory, Practice, and Agents
dtunkelang
0
280
DORA_Metrics.pdf
wagnerfusca
1
140
AI Native Platform Engineering 〜PlatformとAgileで“作る速さ”を“価値”へ〜
uya116
0
530
Terraformを用いたJamf Pro構成のIaC, GitOps化への挑戦
yukun
0
120
HolmesGPTで始めるSREエージェント入門!プラットフォームの障害調査はAIにお任せ 〜
sanghyuk
0
250
全社共通データ基盤をつくる。ソニーのDatabricks活用とデータガバナンス設計の裏側
sony
0
270
なぜAI任せのゲームは面白くならないのか?
hirohasuyoutube
0
350
Featured
See All Featured
Heart Work Chapter 1 - Part 1
lfama
PRO
10
36k
Agile Leadership in an Agile Organization
kimpetersen
PRO
0
250
Building AI with AI
inesmontani
PRO
1
1.3k
Rails Girls Zürich Keynote
gr2m
96
14k
Six Lessons from altMBA
skipperchong
29
4.5k
Building Applications with DynamoDB
mza
96
7.2k
Noah Learner - AI + Me: how we built a GSC Bulk Export data pipeline
techseoconnect
PRO
0
440
The World Runs on Bad Software
bkeepers
PRO
72
12k
Rebuilding a faster, lazier Slack
samanthasiow
85
9.7k
コードの90%をAIが書く世界で何が待っているのか / What awaits us in a world where 90% of the code is written by AI
rkaga
63
46k
Data-driven link building: lessons from a $708K investment (BrightonSEO talk)
szymonslowik
1
1.3k
brightonSEO & MeasureFest 2025 - Christian Goodrich - Winning strategies for Black Friday CRO & PPC
cargoodrich
3
850
Transcript
TransGAN: Two Transformers Can Make One Strong GAN 第六回 全日本コンピュータビジョン勉強会
Transformer 読み会 2021/04/18 kiyo (hrs1985)
自己紹介 twitter : @hrs1985 Qiita : https://qiita.com/hrs1985 github : https://github.com/kiyohiro8
株式会社カブクで機械学習エンジニアをしています。 深層生成モデル、画像の変換 ゲームの強化学習 あたりに興味があります。 twitter アイコン
論文の概要 TransGAN: Two Transformers Can Make One Strong GAN (https://arxiv.org/abs/2102.07074)
1. Transformer のみで GAN を構成した (CNN が非必須であることを示した) 2. アーキテクチャと学習方法を工夫することで CIFAR-10 や STL-10 で CNN ベースの GAN に匹敵する性能が出せた。 モデルは https://github.com/VITA-Group/TransGAN に公開されている ただし推論のみ
Generative Adversarial Models Generator はノイズ (z) から fake sample を作る
Discriminator は入力された画像の real / fake を判別する
Attention (Transformer) と GAN CNN + Attention の GAN は
Self-Attention GAN などで使われており、性能向上に寄与している 今回は Convolutional Layer を一切使わずにAttention (Transformer) のみで GAN を構成した Self-Attention Generative Adversarial Networks (https://arxiv.org/abs/1805.08318) より
Transformer Generator / Discriminator Generator / Discriminator ともに Transformer だけで構成されている
Transformer Encoder Block Multi-Head Self Attention → MLP を繋げて 1つのブロックにする
Multi-Head Self Attention と MLP の前に Layer Normalization を挟む
Memory-Friendly Generator 画像サイズは NLP でいう文の長さ (単語数) に相当する。 32x32 の低解像度でも 1024
単語の文となってしまい Attention の計算量がかさむ。 Transformer Encoder を何回か通す → UpScaling (pixel shuffle) →これを繰り返し、目的の画像サイズまで大きくしていく ←各 pixel が NLP でいう word に相当する
Discriminator 画像を 8x8 のパッチに分割 →Transformer Encoder を通す →最終層で特徴を集約して real /
fake 判定
シンプルな TransGAN Transformer の Generator はよい Transformer の Discriminator はダメ
データ拡張 データ拡張 (DiffAug) を導入することで IS も FID も改善
Self-Supervised Auxiliary Task 補助タスクとして、Generator に画像の高解像度化タスクも解かせる 低解像度画像 高解像度化された画像 MSE loss
Locality-Aware Initialization query 位置 (赤) に対して参照できる key の範囲を制限する 学習初期では狭く、後期では広い範囲を参照する
モデルサイズの効果 モデルサイズが大きいほど強い
既存手法との比較 CIFAR-10、STL-10 で SoTA またはそれに匹敵する程度の性能が出た
出力画像例
結論 ・Transformer のみで構成された GAN である TransGAN を提案した ・学習を工夫することで CNN ベースの
GAN に匹敵する性能が出せた ・今後自然言語処理分野のテクニックを取り入れることで性能向上ができるかも?
None
Network Architecture
学習の計算量
Settings