Model picks for the 16GB VRAM + 64GB RAM tier (2026)
Parent: Local LLM Inference on Consumer Hardware (RTX 5080 eGPU, Linux, Remote Access) · Topic entry · 10 branches · skill ai-llm-model-layer
A published reference is not available for this topic yet.
Children
- Qwen3.8-27B dense VL (frontier)
- Qwen3.6-35B-A3B MoE (frontier)
- Qwen3-Coder-Next 80B-A3B (frontier)
- gpt-oss-120b and gpt-oss-20b MXFP4 (frontier)
- Gemma 4 26B-A4B and 12B (frontier)
- GLM-4.7-Flash (frontier)
- Mistral Small 4 119B-A6B (frontier)
- Hybrid-attention KV budgets (frontier)
- Unsloth Dynamic GGUF quant choice (frontier)
- KV-cache quantization at 16GB (frontier)
Frontier under this node: GLM-4.7-Flash, Gemma 4 26B-A4B and 12B, Hybrid-attention KV budgets, KV-cache quantization at 16GB, Mistral Small 4 119B-A6B, Qwen3-Coder-Next 80B-A3B, Qwen3.6-35B-A3B MoE, Qwen3.8-27B dense VL, Unsloth Dynamic GGUF quant choice, gpt-oss-120b and gpt-oss-20b MXFP4