DDR5 and Host Tuning for CPU-Side LLM Inference and MoE Offload
Parent: Local LLM Inference on Consumer Hardware (RTX 5080 eGPU, Linux, Remote Access) · Topic entry · 10 branches · skill ai-llm-model-layer
A published reference is not available for this topic yet.
Children
- DDR5 theoretical vs measured bandwidth (channels x MT/s x 8, IFOP/FCLK/UCLK limits) (frontier)
- Measuring memory bandwidth on Linux (mbw, STREAM, likwid-bench, Intel MLC, dmidecode configured speed) (frontier)
- XMP/EXPO profiles and JEDEC vs platform-supported speeds (Intel and AMD 1DPC/2DPC matrices) (frontier)
- 4-DIMM downclocking and 2x32 vs 2x48 vs 4x32 capacity trade-offs (frontier)
- Decode tokens/s estimate: effective bandwidth over active bytes per token (frontier)
- llama.cpp thread count, SMT, Intel E-cores, CCD and taskset/cpu-mask affinity (frontier)
- Hugepages, THP, mlock/--load-mode and NUMA settings for CPU inference (frontier)
- CPU governor and amd-pstate EPP for memory-bound LLM decode (frontier)
- Prefill compute-bound vs decode memory-bound and expert placement (frontier)
- Upgrade paths: faster kits, 96/128GB, Strix Halo, quad/8-channel HEDT and Xeon 600 (frontier)
Frontier under this node: 4-DIMM downclocking and 2x32 vs 2x48 vs 4x32 capacity trade-offs, CPU governor and amd-pstate EPP for memory-bound LLM decode, DDR5 theoretical vs measured bandwidth (channels x MT/s x 8, IFOP/FCLK/UCLK limits), Decode tokens/s estimate: effective bandwidth over active bytes per token, Hugepages, THP, mlock/--load-mode and NUMA settings for CPU inference, Measuring memory bandwidth on Linux (mbw, STREAM, likwid-bench, Intel MLC, dmidecode configured speed), Prefill compute-bound vs decode memory-bound and expert placement, Upgrade paths: faster kits, 96/128GB, Strix Halo, quad/8-channel HEDT and Xeon 600, XMP/EXPO profiles and JEDEC vs platform-supported speeds (Intel and AMD 1DPC/2DPC matrices), llama.cpp thread count, SMT, Intel E-cores, CCD and taskset/cpu-mask affinity