> ML_LITERATURE // ALAYRAC-2022-FLAMINGO-VISUAL-LANGUAGE-MODEL_v1.0
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan · Advances in Neural Information Processing Systems (NeurIPS) (2022)
Principal Contribution
Bridged frozen pretrained vision backbones and frozen autoregressive language models using gated cross-attention layers and a Perceiver Resampler.
Operational Relevance
Serves as qualified reference for implementing task-multimodal, task-question-answering in production systems.
Assumptions
- Underlying spatio-temporal continuity and domain distributional stability hold
Limitations
- Performance scaling and computational footprint depend on receptive field depth, sequence length, and resolution
