Skip to content
VinVL: Revisiting Visual Representations in Vision-Language Models — Pengchuan Zhang (2021) | RDL Network