1 code implementation • 27 Oct 2022 • Ge-Peng Ji, Mingcheng Zhuge, Dehong Gao, Deng-Ping Fan, Christos Sakaridis, Luc van Gool
We present a masked vision-language transformer (MVLT) for fashion-specific multi-modal representation.
Image Reconstruction Retrieval