the short version
- ds4 targets the 284-billion-parameter DeepSeek V4 Flash on high-memory machines using asymmetric 2-bit quantization. The routed experts are compressed and the critical shared paths are kept precise.
- On the published q2 table, an M5 Max generates about twice as fast as a DGX Spark, and the Spark prefills about twice as fast at 65,536 tokens of context.
- The ds4 sources contain no head-to-head benchmark against llama.cpp or Ollama, so any speed comparison with those runners has to be measured locally.
- ds4 supports a short list of models through one engine with a CLI, an HTTP server exposing /v1/chat/completions, /v1/messages and /v1/responses, and an agent. It does not try to be a general model runner.
The ds4 LLM engine, DwarfStar 4, is an MIT-licensed C inference engine for high-memory Apple Silicon, CUDA and ROCm machines. It runs DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next locally. Its benchmark table reports 39.4 tokens per second of generation and 790.2 tokens per second of prefill at q2 on a 128 GB M5 Max with a 2,048-token context. The table does not name the model, but the site presents V4 Flash Q2 as the 128 GB baseline. The sources contain no benchmark against llama.cpp or Ollama, so that comparison has to be run locally.
The site says the usual path for a 284-billion-parameter mixture-of-experts model like DeepSeek V4 Flash is remote serving, and that ds4 starts from the opposite constraint. It supports a short list of large open-weight models on one machine. One set of binaries covers chat, a local API server and a coding agent. For an engineer choosing a local serving stack, the first questions are whether the model you need is on that list and whether your hardware matches one of the target classes.
What the ds4 LLM stack actually runs
The supported families are DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next. The site lists both text and vision models. The engine is written in C with Metal, CUDA and ROCm backends. The site names quantization levels (Q2, Q4) and download keys such as ds4f-q2, but it does not name an on-disk weight format. It is therefore not established whether ds4 reads files from other runners.
- Apple Silicon Mac, 64 GB and up depending on model, via Metal. The site advertises Qwen on 64 GB.
- NVIDIA DGX Spark or a generic CUDA Linux box.
- AMD Strix Halo machines such as the Framework Desktop and similar, via ROCm.
- Mac Studio 512 GB for V4.1 Q4 or PRO-class headroom. The memory matrix goes up to two 512 GB machines.
At 128 GB, the site's hardware matrix lists V4 Flash Q2 as the baseline. GLM 5.3 Q2 and Qwen Q4 also fit at that size, and V4.1 Q2 streams from SSD. Other listed features are SSD streaming, tensor parallelism, session batching, DSPARK + MTP and vision input.
git clone https://github.com/antirez/ds4
cd ds4 && ./download_model.sh ds4f-q2
make # macOS, Metal
make cuda-spark # Linux, DGX Spark
./ds4
./ds4-server --ctx 100000 # or serve an APIHow 2-bit expert quantization fits 284B parameters
The method is asymmetric 2-bit quantization. It targets the routed experts and preserves the critical shared paths. The site says this is how the supported routed-MoE builds fit their target machines, and it describes the result as compressed rather than lobotomized. The site does not publish an accuracy benchmark for these quants.
Compress the routed experts, keep critical shared paths precise.
ds4 also treats the KV cache as something that can live on disk. Long prefixes are saved to SSD and resumed by prompt hash, so a restart does not have to mean a full re-prefill. The three interfaces (./ds4, ./ds4-server and ./ds4-agent) share the same model state and cache. One commenter on Hacker News noted that ds4-agent is append-only and never rewrites message history, which keeps the KV cache prefix reusable.
M5 Max generates faster, DGX Spark prefills faster
The benchmark table compares two 128 GB machines running q2. On the M5 Max, generation falls from 39.4 to 27.6 tokens per second as context grows from 2,048 to 65,536 tokens, and prefill drops from 790.2 to 398.5. On the DGX Spark, prefill barely changes (825.8 to 823.0), but generation is lower and falls from 18.1 to 13.8.
The Mac generates roughly twice as fast at both context lengths, and the Spark prefills about twice as fast at 65,536 tokens. Dividing by those rates, ingesting a 65,536-token prompt takes about 164 seconds on the Mac and about 80 seconds on the Spark. That cost is what the SSD-resident KV cache is designed to avoid repeating. The site's hardware-matrix figure for the M5 Max at 32K context is 34.4 tokens per second of generation and 557 of prefill, which it describes as estimates from its benchmark table.
How ds4 compares with llama.cpp and Ollama
The source material has no benchmark of ds4 against llama.cpp or Ollama, and Ollama is not mentioned at all. Any throughput comparison has to be run on your own hardware with matched quantization and context. The difference the sources do support is scope: ds4 supports a few model families and a few hardware classes, while general runners aim to cover many models. Its server exposes /v1/chat/completions, /v1/messages and /v1/responses, and the site lists OpenCode, Claude Code, Codex CLI and Pi alongside those endpoints.
Commenters on Hacker News argued that this narrowness is the point. They said a general runner can be misconfigured in ways that break tool calling or reduce performance, and that ds4 avoids this by optimizing a small set of models for one category of hardware. The tradeoff is that ds4 cannot run your target model if it is not DeepSeek V4/V4.1 Flash, GLM 5.x or Qwen3.8 Flash Next.
it only supports a small set of carefully chosen models, but it supports them really well
What practitioners are pushing back on
- Quant quality. One commenter linked a benchmark in which ds4 quants beat Unsloth quants. Another asked which ds4 checkpoint and quantization level were tested, and the excerpt shows no answer.
- Hardware coverage. Commenters asked for Intel support and better AMD support. The site lists ROCm only for Strix Halo-class machines, and Intel is absent.
- Long context. One commenter described a fused TQ change (PR 1115) for 1M-token context with Qwen 3.8 Flash Next on a 128 GB M5 Max. Another said 1M context already worked on the same setup. The sources do not settle which is right.
- Provenance. One commenter called it a vibe-coded knock-off of llama.cpp. Others replied that the emphasis is performance and usable coding and agentic ability on consumer AI hardware.
- Language choice. Commenters reported that antirez finds Rust less ergonomic and sees security-critical code as the reason to choose it. That bears on who contributes, not on what the engine runs.
Downstream work has also appeared. One commenter maintains a fork that packages ds4 as shared libraries for FFI, with Go bindings (ds4go) that added Vision and Qwen support as ds4 added them. Another commenter wrote a separate engine, inspired by DwarfStar, for 32 GB Intel Xe-LP laptops; it currently supports only a quantized Gemma-4. A third reported about 22 tokens per second for that engine on an Intel Ultra 7 255H iGPU after a small patch.
What is still unknown about ds4
Four questions remain open. First, what the weight format is and whether ds4 can load files used by other runners. Second, how accurate the 2-bit asymmetric quants are on a named benchmark compared with Q4 builds of the same model. Third, what the numbers are for GLM 5.x, Qwen3.8 Flash Next and the ROCm path: the table shown covers only q2 rows on M5 Max and DGX Spark and does not name the model. Fourth, how ds4 performs in a matched comparison with llama.cpp. Until those exist, treat 39.4 tokens per second as the best case in the published table and run your own long-context workloads before relying on ds4.
Questions this raises
What models does the ds4 LLM engine support?
ds4 supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, including text and vision models. It is narrow by design, so it cannot run models outside that list.
Is ds4 faster than llama.cpp or Ollama?
The sources contain no benchmark comparing ds4 with llama.cpp or Ollama. Any comparison has to be run on your own hardware with matched quantization and context length.
How fast is ds4 on M5 Max vs DGX Spark?
At q2 on 128 GB machines, the M5 Max generates roughly twice as fast (39.4 vs 18.1 t/s at 2,048 tokens). The DGX Spark prefills about twice as fast at 65,536 tokens (823.0 vs 398.5 t/s).
These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.
