llama.cpp RPC -sm tensor PR 26610
Parent: Mac local LLMs: Clusters, RDMA, exo and ds4 · Published reference · snapshot 2026-10-05
↓ Facts as markdownall context files
Existing coverage: exo-cluster-software.md records the PR's design bullets and the Spark pp2048 619.36 / tg128 19.75; its statement 'Not Mac-measured' is contradicted below
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Facts
- Existing coverage: exo-cluster-software.md records the PR's design bullets and the Spark pp2048 619.36 / tg128 19.75; its statement 'Not Mac-measured' is contradicted below [source]
- On the same two Macs -sm layer gave tg2048 23.05 and pp2048 270.3, so RPC tensor split was 2.4x slower in decode and 1.6x slower in prefill than layer split on Mac at the time of that test [source]
- The author reported -sm layer on master crashing on the Spark setup with 'Remote RPC server crashed or returned malformed response' (ggml-rpc.cpp:519), and recalled about 400 pp and 15 tg for layer split, so no same-box layer baseline is in the PR table [source]
- The tensor path hangs after tensors are sharded when more than one RPC backend is used and the workers cannot reach each other at the addresses the client passed (tested on Qwen3 0.6B, Qwen3 27B and DeepSeek V4) [source]
- Workaround: GGML_RPC_NO_COMM=1 disables the worker-to-worker comm channel and relays all-reduce through the client, which works over RDMA hardware but loses the direct path; using LAN IPs works but takes over 10 s per token [source]
- On Thunderbolt point-to-point links a worker cannot derive a peer's TB address from the address the client gave it, so worker-to-worker RDMA on Mac needs a mesh and a way to pass peer TB IPs to each rpc server; the Spark setup works because all nodes share one routable RDMA fabric [source]
- The comm channel listens on rpc port plus 1000 and a reviewer (rgerganov) asked for an explicit command-line flag instead [source]
- The custom all-reduce only engages for exactly 2 backends; with more it falls back to the Meta backend all-reduce, which the author calls slow but working, and one rank per endpoint is enforced because a server processes its socket sequentially [source]
- MTP speculative decoding (dspark) does not work with the RPC tensor path because an add across the split is unsupported; the author's MTP benchmark on 2 Sparks after PRs 28387 and 28390 showed 41.5 tok/s on python code at 0.82 accept rate and 54.1 percent aggregate acceptance [source]
- Reviewer rgerganov asked to bump the RPC version to 6.0.0 and drop backward-compatibility gating; a downstream cherry-pick of the graph-cache and 2D tensor parts bumped the RPC protocol to 7, and a tester asked for RPC_PROTO_MAJOR_VERSION to be bumped because mixed server and client versions misbehave [source]
- Over gigabit Ethernet, tensor mode on Qwen3.8 27B IQ3_S with a 4060 Ti plus two BC250 Vulkan nodes gave 8 tok/s pp and 2.5 tok/s tg against 100 and 5 for layer mode [source]
- Status on Oct 4 2026: still open, ggerganov rebased the branch (force-push Oct 3) onto stack PR 29414, new commits 'rpc: allow -sm tensor' and 'fix flush for apple rdma' landed, two approving reviews are required and none is shown; the author volunteered to maintain it on Sep 25 [source]
- Related merges: PR 28387 (ggml: allow backend inputs to not create another split) and PR 28569 (re-enable -sm tensor for qwen4exp); NCCL multi-node tensor parallel PR 28967 is a draft [source]
- Reviewer JohannesGaessler rejected assigning stable uids to aux graphs because the API lets users pass arbitrary graphs; the author said those changes were not required [source]
Corrections and disagreements
- CONTRADICTS exo-cluster-software.md: a Mac measurement exists. ryan5rdx ran -sm tensor on 2 M3 Ultra over RDMA (with PR 26421) on DeepSeek V4 Flash MXFP4 and got tg2048 9.61 tok/s and pp2048 166.05 tok/s [source]
Children
- No children recorded.