<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llm-benchpacks-workload-level-agent-benchmarks/ · pack 2026-10-05 · ~2360 tokens -->

# llm-benchpacks workload-level agent benchmarks

> A pack manifest owns request defaults, stream flag, warmup count and repetition count. Warmups use the same adapter, endpoint, model and prompt as measured runs, write raw files, and are excluded from `run.jsonl`.

Parent: [Mac local LLMs: Benchmarking and comparisons](https://llms-explorer.com/tree/mac-local-llms-benchmarking-and-comparisons/) · 1 facets · 37 facts · page: https://llms-explorer.com/tree/llm-benchpacks-workload-level-agent-benchmarks/

## Facts

- A pack manifest owns request defaults, stream flag, warmup count and repetition count. Warmups use the same adapter, endpoint, model and prompt as measured runs, write raw files, and are excluded from `run.jsonl`. — source: `asserted`
- Adapters: `ollama-generate` (native /api/generate) and `openai-chat` (streaming optional). In streaming mode TTFT is the time to the first non-empty content delta. — source: `asserted`
- `benchpack compare` reports median prompt tokens beside median cached prompt tokens and derives a case-level "prefill parity" status. `prefill_tps` is shown only for cases whose status is comparable. — source: `asserted`
- Repo-task packs run an adapter call, then optionally an external agent harness against a prepared workspace, capture the workspace patch, and run deterministic verifiers. — source: `asserted`
- April 2026: runtime comparison packs for mlx-lm and llama-server (result sets 2026-04-28 and 2026-04-29). — source: `asserted`
- 2026-05-05: Qwen3.6 M4/M5 sweep across llama.cpp, Ollama and MLX. — source: `asserted`
- By 2026-09-22 (latest commit read): a SQLite registry with provenance-labelled bundles, a static results site, an external-agent harness with optional model-call telemetry, and a hard one-shot "wrap django-resume in Electron" benchmark for hosted coding agents. — source: `asserted`
- Servers that reject `stream_options.include_usage` need `--openai-stream-usage omit`; streamed output and TTFT are preserved, but usage-derived token counts and rates may stay null. — source: `asserted`
- Ollama native rows report decode timing but not cached-prompt fields, so a cached prefill cannot be excluded and the report warns of incomplete cache metadata. — source: `asserted`
- The prompt-only `desktop-django-wrap` pack scores by regex and `patch-from-failure` is a tiny smoke test, so neither is a broad coding-agent quality claim. — source: `asserted`
- Power, thermal state and background load were not captured or controlled in the 2026-05-05 sweep. — source: `asserted`
- Tmux matrix runs execute packs sequentially so they do not contend for the same local runtime; if one pack fails, later windows report they were skipped. — source: `asserted`
- Whether raw tok/s or workload pass rate should headline: the project says both are needed, since coding agents also depend on TTFT, prefill speed, cache reuse, long-context stability and tool-call formatting. The 2026-05-05 sweep shows why one number misleads: MLX led throughput on both models and hosts, yet every runtime failed the `patch-from-failure` verifier. — source: `asserted`
- Whether streaming TTFT from the first non-empty content delta excludes reasoning-field deltas for thinking models (the sources say "content" only). — source: `asserted`
- Whether the registry's self-reported bundles can be cross-checked: provenance is a label, not a verification. — source: `asserted`
- llm-benchpacks describes itself as portable benchmark packs for local LLM runtimes and coding-agent workloads, and asks whether direct mlx-lm beats Ollama's MLX path and how llama-server compares with mlx_lm.server on the same Mac. — [source](https://github.com/ephes/llm-benchpacks)
- The pack list includes smoke-chat, runtime-sweep, desktop-django-wrap, patch-from-failure, endpoint-python-correctness, python-regression-fix, django-dashboard-regression-fix, mini-project-completion and tool-json. — [source](https://github.com/ephes/llm-benchpacks)
- The default tmux matrix runs smoke-chat, runtime-sweep, desktop-django-wrap and patch-from-failure, and a `coding-tasks` pack set adds patch-from-failure, python-regression-fix and django-dashboard-regression-fix as optional exploratory signal. — [source](https://github.com/ephes/llm-benchpacks)
- A separate `coding-tasks-external-agent` pack set runs the same fixtures and verifiers but tells an external agent to edit the prepared workspace directly through BENCHPACK_EXTERNAL_AGENT_ARGV. — [source](https://github.com/ephes/llm-benchpacks)
- The external-agent harness runs without a shell, passes a runner-owned JSON context file, and on timeout stops the subprocess process group with a bounded terminate-then-kill policy. — [source](https://github.com/ephes/llm-benchpacks)
- Optional external-agent model-call telemetry is a JSONL file of allowlisted safe fields and is kept out of run.jsonl, with prompts, responses, headers and credentials excluded by convention. — [source](https://github.com/ephes/llm-benchpacks)
- The hard one-shot benchmark clones a source repo, runs one unattended agent session, captures the model-authored diff before verification, then runs `npm --prefix electron install`, Electron Node tests and a packaged smoke test. — [source](https://github.com/ephes/llm-benchpacks)
- The one-shot helper supports runners codex-yolo, claude-yolo and pi, with `--reasoning-effort none` as the no-reasoning lane for Codex and `--claude-effort` for Claude Code because that CLI exposes effort levels and no literal thinking-off switch. — [source](https://github.com/ephes/llm-benchpacks)
- Curated one-shot rows live in data/agent-wrap-oneshot-results.json and import into a SQLite registry whose static site filters by result, harness, provider, model and thinking mode and encodes every filter in the query string. — [source](https://github.com/ephes/llm-benchpacks)
- Measured repetition count and warmup count live in the pack manifest defaults, not in CLI flags, and warmups are excluded from run.jsonl. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/decisions.md)
- Backend-reported cached prompt-token counts are normalized as `tokens.cached_prompt` in run.jsonl. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/decisions.md)
- `benchpack compare` gives each case a prefill-parity status with priority missing-case, prompt-missing, prompt-diff, cache-missing, cache-diff, then comparable, and prints prefill_tps only for comparable cases. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/decisions.md)
- openai-chat streaming measures TTFT from the first non-empty content delta, and a pack may leave streaming off to keep non-streaming smoke coverage. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/benchpack-format.md)
- `--openai-stream-usage include` is the default and sends stream_options.include_usage, while `omit` preserves streamed output and TTFT but may leave usage-derived token counts and rates null. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/apple-silicon-m4-m5-runbook.md)
- The 2026-05-05 sweep used an M5 Max MacBook Pro (Mac17,7, 64 GB) and an M4 Max Mac Studio (Mac16,9, 128 GB) with Qwen3.6-35B-A3B (GGUF UD-Q4_K_M, MLX MXFP4) and Qwen3.6-27B (GGUF Q4_K_M, MLX 4-bit). — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- The sweep ran llama.cpp build 9020 with reasoning off, Ollama 0.20.5 native /api/generate with num_ctx 4096, and mlx_lm.server 0.31.3 with enable_thinking=false. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- Median total tok/s on the MoE for short, medium and long cases on the M5 Max were llama.cpp 91.87, 91.41, 91.60; Ollama 49.98, 46.75, 47.22; MLX 102.54, 106.84, 103.65. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- Median total tok/s on the dense 27B for the three cases on the M5 Max were llama.cpp 24.96, 24.25, 23.69; Ollama 16.75, 13.86, 13.88; MLX 29.46, 29.92, 29.58. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- On the M4 Max the MoE medians were llama.cpp 65.98, 72.54, 69.35; Ollama 40.56, 38.05, 38.95; MLX 89.81, 88.66, 90.44, and the dense medians were llama.cpp 21.66, 21.78, 21.25; Ollama 15.02, 13.08, 13.63; MLX 25.46, 25.95, 26.02. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- In the sweep, desktop-django-wrap regex scoring passed for MLX and llama.cpp and failed for Ollama on both models and hosts, and patch-from-failure verifiers failed for every runtime, model and host combination. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- The sweep did not capture power or thermal state and did not control background load. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
- For remote M4 runs only run.jsonl, summary.md, hardware.json and run-metadata.json were pulled back, and raw responses, workspaces and verifier artifacts stayed on the host. — [source](https://raw.githubusercontent.com/ephes/llm-benchpacks/main/docs/qwen36-m4-m5-benchmark-summary.md)
