Features B4 Van-Quang Nguyen1 , Masanori Suganuma2,1 , and Takayuki Okatani1,2 1Graduate School of Information Sciences, Tohoku University 2RIKEN Center for AIP ECCV2022 Nguyen, V. Q., Suganuma, M., & Okatani, T. (2022, October). Grit: Faster and beCer image capEoning transformer using dual visual features. In ECCV (pp. 167-184).
a remote GT2: a couple of people that are staring at a tv GT3: two women playing a video game in a living room GRIT: two women playing a video game in a living room GT1:a sewage lid on the ground with a para sail chute in the background GT2:there is a balloon that is flying over the ground parachute over a large valley with a man made structure GRIT: a pair of scissors sitting on the ground with GT1:a collection of artwork leaning against a wooden fence a collection of poster arts lined up on the fence GT2*a collection of paintings against a fence outside several paintings leaning against Polos: 92.4 ☺ GRIT: a stop sign on a sidewalk next to a stop sign Polos: 9.64 Polos: 9.75 ☹ ☹ 改善案: Polosを報酬として利用する強化学習の実施・MLLMの説明能力の利用 Polos [Wada+ CVPR24] にて評価: 67.02