LLM Compression (Quantization, Distillation, Pruning, Merging)

Parent: LLM Models and APIs · researched 2026-05-31T20:39:48.899Z· 30 sources · 13 concepts · skill ai-agent-engineering

AI & agent-engineering family ROUTER. Split into: ai-agents-orchestration (agent frameworks, multi-agent, memory, planning, guardrails, coding/GUI agents, autonomous loops, eval); ai-rag-retrieval (RA

ai-agent-engineering

Children

Frontier under this node: AWQ, BitNet & EfficientQAT, Compressed-model evaluation (perplexity vs KL vs flips), FP8/INT4 & MX/MXFP4 microscaling, GGUF llama.cpp k-quants & imatrix, GPTQ, KV-cache quantization (KIVI/KVQuant), Knowledge distillation (logit/feature/sequence/on-policy GKD/MiniLLM/DistiLLM), Model merging (TIES/DARE/SLERP/task-arithmetic/soups/MergeKit), PTQ vs QAT, Pruning & sparsity (SparseGPT, Wanda, 2:4 N:M), SmoothQuant (W8A8 activation quant), bitsandbytes NF4 & LLM.int8

← the whole tree · 3D view· how to read this page