Local LLM Troubleshooting Playbook and Recommended Settings (Ollama, LM Studio, llama-server)
Parent: Local LLM Inference on Consumer Hardware (RTX 5080 eGPU, Linux, Remote Access) · Topic entry · 13 branches · skill ai-llm-model-layer
A published reference is not available for this topic yet.
Children
- Ollama VRAM-tier default context and silent chat truncation (frontier)
- llama-server --fit auto context shrink and ctx-per-slot (frontier)
- LM Studio context overflow policy (frontier)
- Ollama CPU/GPU split diagnosis and env-var tuning
- GPU-not-detected CPU fallback (nvidia_uvm after suspend, container toolkit) (frontier)
- Cold-load latency and keep_alive / TTL residency (frontier)
- Chat template and EOS failures (--jinja, Go TEMPLATE) (frontier)
- gpt-oss Harmony format serving (frontier)
- Reasoning-format parsing and think-tag leakage (frontier)
- Local tool calling and grammar-constrained JSON (frontier)
- Per-family vendor sampling defaults and GGUF sampling metadata (frontier)
- Repetition loops vs repeat_penalty vs presence_penalty (frontier)
- 16GB-card baseline config for Ollama, LM Studio, llama-server (frontier)
Frontier under this node: 16GB-card baseline config for Ollama, LM Studio, llama-server, Chat template and EOS failures (--jinja, Go TEMPLATE), Cold-load latency and keep_alive / TTL residency, GPU-not-detected CPU fallback (nvidia_uvm after suspend, container toolkit), LM Studio context overflow policy, Local tool calling and grammar-constrained JSON, Ollama VRAM-tier default context and silent chat truncation, Per-family vendor sampling defaults and GGUF sampling metadata, Reasoning-format parsing and think-tag leakage, Repetition loops vs repeat_penalty vs presence_penalty, gpt-oss Harmony format serving, llama-server --fit auto context shrink and ctx-per-slot