Lucebox Mix
Qwen for speed and DeepSeek V4 for deeper thinking, at once.
Speed Quality 116.7 GBQwen + DeepSeek V4 togetherEngine / open source
The open source inference engine on every Lucebox, tuned for its R9700 and Ryzen AI MAX+. Apache 2.0, on GitHub.
Models
Qwen for speed and DeepSeek V4 for deeper thinking, at once.
Speed Quality 116.7 GBQwen + DeepSeek V4 togetherCareful answers for agents, research and long documents.
Speed Quality 101.5 GBR9700 + Strix Halo128K contextDeepSeek that also reads images, for up to 4 people.
Speed Quality 93.7 GBStrix Halo + R9700, 4 at once16K contextThe fastest: code, quick answers, agents that call it often.
Speed Quality 15.2 GBR9700, 8 at once128K contextAlso reads images (charts, screenshots, photos), for up to 4 people.
Speed Quality 16 GBR9700, 4 at once16K contextThe newest and deepest DeepSeek, with a long 128K context.
Speed Quality 368.4 GBR9700 + Strix Halo, SSD with 128 GB128K contextInstall any of them from Manage, or load any GGUF you bring. The box runs Ubuntu Server with ROCm and root access, so other open source engines install as on any Linux machine.
Speed
The best run in each report: decode while writing code, prefill on prompts of 1.4K to 32K tokens. Qwen runs on the R9700, DeepSeek across the R9700 and the Ryzen AI MAX+.
Qwen 3.8 27BDeepSeek V4 FlashDeepSeek V4.1 FlashBatching
| Series | 1 users | 2 users | 3 users | 4 users | 5 users | 6 users | 7 users | 8 users |
|---|---|---|---|---|---|---|---|---|
| Qwen 3.8 27B, code | 106.6 | 188.8 | 209.4 | 262.1 | 300.9 | Not measured | Not measured | Not measured |
| Qwen 3.8 27B, images | 84 | Not measured | Not measured | 170 | Not measured | Not measured | Not measured | 195 |
| DeepSeek V4 Flash, text | 23.8 | 35.4 | 42.9 | 48.4 | Not measured | Not measured | Not measured | Not measured |
Tokens per second across the box, prompt reading included, so one user sits below the decode peaks above. DeepSeek runs on the Ryzen AI MAX+ alone, without its drafter.
Continuous batchingImage questionsEstimate, Qwen 3.8 27B coding agents
Agents pause between turns to run tools and wait for input. If each writes 30% of the day, 10 agents need more than 5 requests at once under 5% of the time, and at 5 each still gets about 60 tok/s. At 2 to 3 agents per developer, that is 3 to 5 developers.
Manage serves up to 8 at once, so a busy moment slows agents instead of queueing them.
Speculative
Guesses a block of up to 16 tokens; the model checks them in one pass.
ReportDeepSeek drafter; its confidence head sets how many tokens each step verifies.
ReportA small model picks the parts of a long prompt worth reading. Lossy.
The decode drafters are lossless: the model keeps only the tokens it would have written itself.
Optimizations
Reuses the saved tool definitions on every agent turn.
ReportPages cold KV to host RAM, bit‑exact, for long context on one card.
ReportPlaces experts in VRAM, unified memory or on the SSD by use.
ReportRuns one model on the R9700 and the Ryzen AI MAX+ at once.
ReportPaged attention and one scheduler for many requests at once.
ReportReads an image prompt a few layers at a time, so other users keep writing.
ReportTwo processors
3× faster prompt reading on DeepSeek V4.1 Flash.
2× faster writing on the same model.
Each keeps its speed while the other serves: own chip, own memory.
Chip alone is Framework's published run on the same 192 GB Ryzen AI MAX+ PRO 495, with its own engine and quantization.
DeepSeek V4.1 FlashTwo models at onceCompared
Qwen3.8-27B on the R9700
Decode speed on HumanEval coding prompts, averaged. The same drafter in llama.cpp reaches 54.6 tok/s.
DeepSeek V4 Flash, 284B
One model split across the R9700 and the Ryzen AI MAX+ unified memory. Measured on the Ryzen AI MAX+ 395 build.
Qwen3.8-27B, image questions
Eight users asking about charts at once, with the same model file, projector and drafter on both machines.