Multimodal & Vision-Language Model Architecture

Parent: LLM Models and APIs · researched 2026-05-31T23:24:00.467Z· 19 sources · 10 concepts · skill multimodal-llm-architecture

> Provenance: reference under the ai-agent-engineering hub. Built via /dr deep-research, 2026-05-31. Owns the "how a text-only transformer becomes multimodal" layer — the vision/audio/video front-end

Multimodal & Vision-Language Model (VLM) Architecture

1. Vision encoders (ViT / CLIP / SigLIP / DINOv2 / EVA)

2. The connector / projector (MLP vs Q-Former vs cross-attention)

3. Fusion strategy (unified/early-fusion vs cross-attention vs late fusion)

4. Native any-to-any & image tokenization (Chameleon, VQ-VAE/VQGAN, Fuyu)

5. High-resolution & dynamic tiling (AnyRes, InternVL tiles, NaViT)

6. Audio, video, speech (Whisper encoders, audio tokens, frame sampling, omni)

7. Multimodal position encoding (2-D RoPE, M-RoPE)

8. VLM training stages (projector-align → visual instruction tuning → multimodal preference/DPO)

9. The VLM landscape (architectural map, mid-2026)

10. Multimodal evaluation & hallucination (MMMU, MMBench, DocVQA, MathVista, POPE/MMHal)

Practical patterns

Anti-patterns

Troubleshooting

Cross-references (ai-agent-engineering hub)

References (2024-2026 primary papers + model cards)

Children

Frontier under this node: Audio/video/speech modalities (Whisper encoder, Thinker-Talker, frame sampling), Fusion strategies (unified-concat vs cross-attention vs late), High-resolution & dynamic tiling (AnyRes/NaViT/native-resolution ViT), Modality connector/projector (MLP/Q-Former/gated cross-attention), Multimodal benchmarks & hallucination (MMMU/MMBench/DocVQA/MathVista/POPE/MMHal), Multimodal position encoding (2-D RoPE, M-RoPE), Native any-to-any & image tokenization (Chameleon VQ-VAE, Fuyu), The VLM model landscape (LLaVA/Qwen-VL/InternVL/Pixtral/Llama-3.2-Vision/Chameleon/GPT-4o), VLM training stages (projector-align, visual instruction tuning, multimodal DPO), Vision encoders (ViT/CLIP/SigLIP/DINOv2/EVA)

← the whole tree · 3D view· how to read this page