Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy: up to 3x faster with the R9700
DeepSeek V4.1 Flash is a 383 GB model. The Lucebox Zero 495 runs it across three kinds of memory: the 32 GB on the Radeon AI PRO R9700, the unified memory of the Ryzen AI MAX+ PRO 495 and, on the 128 GB model, the SSD. Each expert sits where its use pays for it. Compared with the same 192 GB chip without the R9700, a long prompt gets its answer about 3x sooner, and it writes code at up to 49 tok/s.
TL;DR
- A 383 GB model on one desktop. The Lucebox Zero 495 runs DeepSeek V4.1 Flash with a 128K context: up to 49 tok/s writing code with 192 GB of memory, and up to 26 tok/s with 128 GB.
- The R9700 makes it up to 3x faster. On the same 192 GB Ryzen AI MAX+ PRO 495, adding the R9700 reads long prompts three times faster and writes twice as fast.
- Every expert in the right place. The GPU holds the experts the model uses most and unified memory holds the rest. On the 128 GB model the SSD keeps the rarely used ones, and a cache that looks ahead hides most of the reads.
- A better quantized model. Geometric's version answers 164 of 177 hard reasoning questions, against 141 for antirez's Q2 of the same size.
- One click to install. Pick DeepSeek V4.1 Flash on the Engine page of your dashboard and press Install. The server is ready in under a minute.
With and without the R9700
Framework tested DeepSeek V4.1 Flash on the 192 GB Ryzen AI MAX+ PRO 495 on its own, with no discrete GPU. It is the same chip the Lucebox Zero 495 uses, so the comparison shows what the R9700 adds.
Long prompts gain the most. At 32K tokens, the Ryzen AI MAX+ alone reads about 340 tokens a second; with the R9700 it reads about 1,000. Writing goes from about 16 tokens a second to 33. Put together, a 30K-token prompt gets its answer in about 40 seconds instead of almost two minutes.
| Without the R9700 | With the R9700 | Faster | |
|---|---|---|---|
| Long prompt and answer (30K tokens in, 256 out) | almost 2 min | about 40 s | about 3x |
| Prompt reading at 32K tokens | about 340 tok/s | about 1,000 tok/s | 3x |
| Prompt reading at 2K tokens | about 350 tok/s | about 800 tok/s | 2.3x |
| Decode | about 16 tok/s | about 33 tok/s | 2x |
Without the R9700: Framework's published figures, with the DwarfStar engine, antirez's Q2 and the DSpark drafter. With the R9700: our engine and our own quantized model, which is about as close to the full model. Framework reads 2K new tokens on top of existing context; we read the whole prompt.
With 192 GB, every expert stays in memory
The 192 GB model has room for all of the experts. The R9700 keeps the layers every token goes through, the drafter and the experts used most. The integrated GPU of the Ryzen AI MAX+ keeps all the others, around 150 GB, so nothing comes from the SSD while the model runs.
That lets the engine do things the SSD tier rules out:
- Both chips work on the prompt at once. The prompt goes through the model in bands. While the Ryzen AI MAX+ works on the experts of one band, the R9700 already runs attention for the next.
- The running state stays on the R9700. It no longer crosses the PCIe link between layers.
- Matrix cores read the prompt. Attention and the 2-bit expert math run on the GPUs' matrix units.
- Decode keeps its pace. Each step reads the cached history in 16-bit and splits the work eight ways, so decode stays above 30 tok/s even after a 30K-token prompt. On the 128 GB model it drops to about 14 at 20K.
Coding agents benefit too: with the prefix cache, each new turn of a session is read at about 750 tokens a second. These runs use a smaller build of the drafter, about 5 GB instead of 8, which leaves the R9700 room to read the prompt in bigger chunks, and an expert placement fitted to this configuration.
With 128 GB: three tiers
On the 128 GB model the experts no longer all fit in memory, so the SSD becomes a third tier. The three are far apart in speed. The R9700 reads its 32 GB at 640 GB/s, the integrated GPU reads unified memory at around 240 GB/s, and the SSD holds everything but reads only a few GB a second. DeepSeek V4.1 Flash picks 6 of its 384 experts for every token in each of its 40 layers, so where each expert lives decides how fast the model writes.
The 128 GB figures here come from the Ryzen AI MAX+ 395 build with the same R9700. The Zero 495 keeps that GPU and adds faster memory.
Experts are placed by how often they are used. Most tokens go to a small group of experts, so those sit in the fast tiers and the long tail goes to the SSD. With a 4K-token prompt, the R9700 holds 6% of the experts and serves almost a quarter of the calls, while the SSD holds a third of them and serves less than a tenth.
From under 1 to 26 tok/s
Our first version kept the hottest experts on the R9700 and read the rest from the SSD. It gave the right answers, at under one token a second. Each step after that kept the answers the same and removed one bottleneck.
- Unified memory as a second tier. The integrated GPU keeps its own set of experts and computes them while the R9700 does the same.
- Ranked placement. The most used experts go on the fastest tier.
- Expert cache. Experts read from the SSD go through a cache in unified memory, and the next layer's experts are guessed and fetched ahead of time.
- One graph and DSpark. Each decode step runs as one graph across the three tiers, and the DSpark drafter proposes tokens that the model checks in one pass. The checked tokens are identical to plain decoding.
- Pinned memory. The unified-memory tier grows past the part reserved for the GPU into pinned system memory, which moves thousands of experts off the SSD.
- Placement and routing for this box. The placement is fitted to this machine, and the router picks the expert in fast memory when two choices are close.
Long prompts
The 128 GB profile reads prompts in 4K-token chunks and runs consecutive chunks layer by layer, so each expert from the SSD is read once per pass instead of once per chunk. It reads about 120 tokens a second from 4K to 20K tokens. Decode slows as the context grows, because every new token looks back over more history.
In agent sessions, the prefix cache keeps what was already read, so later turns start answering in seconds instead of close to a minute.
Loading 383 GB
On the 128 GB model, about 120 GB of the model is read before the first token: the layers every token uses and the experts kept in memory. The engine reads them on eight threads straight from the SSD, past the operating system's file cache, so loading doesn't push everything else out of memory. The server is ready in under a minute, down from five.
A better quantized model, from Geometric
Geometric built a second quantized model for our engine, tuned for the quality of the answers. Two choices set it apart:
- The dense weights stay in FP8, the format DeepSeek ships them in. A new MXFP8 type stores them exactly, at a little over 8 bits per weight.
- The experts get more bits where they matter. Each layer mixes the standard IQ2 and IQ3 formats, within the size of antirez's Q2.
On a set of 177 hard reasoning questions it answers 164 correctly. antirez's Q2, the same size, gets 141.
With the drafter it scores a little lower, mostly because a few long answers run into the 16K-token limit. Geometric also contributed the engine work the model needs: the MXFP8 type, faster IQ2 and IQ3 expert kernels for both GPUs, and a verify step whose drafted tokens match plain decoding exactly.
It runs on both memory sizes. It writes about as fast as ours but reads prompts more slowly, because our fast matrix-core path doesn't cover its formats yet. Pick our quantized model for speed and Geometric's for the best answers.
Run it
On a Lucebox, open the dashboard, pick DeepSeek V4.1 Flash on the Engine page and press Install. It downloads the model and the drafter, about 390 GB, so check that there is room, then builds the engine and checks every file. Start server has it answering in about a minute.
Elsewhere, the engine runs it with one profile for each memory size. With 128 GB:
luce_server DeepSeek-V4.1-Flash-ROCMFP2S.gguf \
--profile ds41-lucebox \
--draft DeepSeek-V4.1-Flash-DSpark-draft-MXFP4-Q8.gguf With 192 GB:
luce_server DeepSeek-V4.1-Flash-ROCMFP2S.gguf \
--profile ds41-gorgon \
--draft DeepSeek-V4.1-Flash-DSpark-draft-MXFP4-Q8.gguf \
--ds4-expert-placement placement.json \
--ds4-router-bias router_bias.bin Our quantized models and the DSpark drafter are on Hugging Face. For the best answers, use Geometric's ds41-lucebox-final.gguf from the same repository in place of ROCMFP2S; ours reads prompts faster, and it is the one measured in this post. Run it as the Lucebox model service so the unified-memory tier can pin its memory. On the 192 GB model the integrated GPU maps around 150 GB of experts, so raise its GTT limit; the engine docs give the values we use and every setting the profiles turn on.
Which model to pick
A Lucebox runs all three, one at a time, or Qwen next to DeepSeek V4 in the Lucebox Mix. DeepSeek V4.1 Flash is the largest and newest, at up to 49 tok/s on code with 192 GB. DeepSeek V4 Flash writes code at over 50 tok/s and takes a third of the disk space. Qwen 3.8 27B is the fastest, around 200 tok/s on code, for quick answers and coding agents that call the model often. Pick V4.1 for the largest model, and V4 or Qwen when speed matters more.
What is next
On the 128 GB model, decode still waits on the SSD for experts nobody predicted, and prompt reading uses only a small part of what the R9700 can compute. On the 192 GB model, the next step is the matrix-core prompt path for Geometric's formats, so the better model reads prompts as fast as ours.
Hardware and setup
| 192 GB | AMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ PRO 495 with 192 GB of unified memory, Ubuntu, ROCm 7.2, --profile ds41-gorgon |
|---|---|
| 128 GB | AMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ 395 with 128 GB of unified memory (the build before the Zero 495), Crucial P310 2 TB SSD, Ubuntu, ROCm 7.2, --profile ds41-lucebox |
| Model | DeepSeek V4.1 Flash: our quantized model ROCMFP2S and the DSpark drafter, and Geometric's ds41-lucebox-final.gguf, all on Hugging Face |
| Workload | Code writing and prompts up to about 30K tokens, 256-token answers, greedy decoding |
Bottom line
DeepSeek V4.1 Flash is too big for any one chip in the box, so the engine uses all of them. The R9700 does the work every token needs and keeps the experts used most, unified memory holds the rest, and on the 128 GB model the SSD keeps the long tail and is read ahead of time. With 192 GB, that makes the model up to three times faster than the same chip without the R9700, at up to 49 tok/s on one desktop.
Measured by us and rounded. The 192 GB figures come from a Ryzen AI MAX+ PRO 495 with 192 GB in October 2026, the 128 GB figures from the Ryzen AI MAX+ 395 build with the same R9700 in September 2026, with the file cache emptied before each start. Decode speeds are for the first request after start. The figures without the R9700 are Framework's, from their post of September 30, 2026. The quality results for Geometric's model are Geometric's.