Vision-Model / Multimodal-LLM Design Critique Technique

Parent: Multimodal & Vision-Language Model Architecture · Topic entry · 10 branches · skill ai-mcp-sdk-prompting/references/vision-model-design-critique.md

A published reference is not available for this topic yet.

Children

Frontier under this node: Judge biases: position, verbosity, self-preference, sycophancy (sycophantic modality gap), Multi-image & before/after comparison (image labeling, image-order/position bias, order-swap), Object/element hallucination in critique (POPE/AMBER/MMHal, co-occurrence & affirmative bias), Region grounding of findings (Set-of-Mark, OmniParser OCR+icon pre-pass, points>boxes), Reliability practices (few-shot anchors, self-consistency, LLM-as-jury panels, abstention, temp 0), Rubric-based visual critique prompting (describe-then-judge, rubric-as-prompt, supplied definitions), Structured/JSON findings output (json_schema/responseSchema/tool-use; the format tax — reason-first), Text-in-image / OCR limits in UI critique (OCRBench, resolution/detail param, external OCR pre-pass), VLM-as-judge / MLLM-as-judge calibration vs human designers (rank-not-score, pairwise>pointwise), Weak fine-grained spatial reasoning & counting (VSR/BLINK/SpatialEval/CountBench)

← the whole tree · 3D view· how to read this page