By Davide Ciffa

Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy: up to 3x faster with the R9700

DeepSeek V4.1 Flash is a 383 GB model. The Lucebox Zero 495 runs it across three kinds of memory: the 32 GB on the Radeon AI PRO R9700, the unified memory of the Ryzen AI MAX+ PRO 495 and, on the 128 GB model, the SSD. Each expert sits where its use pays for it. Compared with the same 192 GB chip without the R9700, a long prompt gets its answer about 3x sooner, and it writes code at up to 49 tok/s.

An open Lucebox with its Radeon AI PRO R9700 under the CPU cooler, on a wooden deck under a starry sky between a silver pyramid and a plush DeepSeek whale, with the Lucebox and Geometric logos above
The R9700 keeps the experts used most, unified memory the rest.
3xfaster on a long prompt than the same 192 GB Ryzen AI MAX+ without the R9700
2xfaster decode, about 33 tok/s instead of 16
up to 49 tok/swriting code, with every expert in memory

TL;DR

With and without the R9700

Framework tested DeepSeek V4.1 Flash on the 192 GB Ryzen AI MAX+ PRO 495 on its own, with no discrete GPU. It is the same chip the Lucebox Zero 495 uses, so the comparison shows what the R9700 adds.

DeepSeek V4.1 Flash on a 192 GB Ryzen AI MAX+ PRO 495, with and without the R9700 Prompt reading at 32K tokens: about 340 tok/s without the R9700, about 1,000 with it. Decode: about 16 tok/s without the R9700, about 33 with it. prompt reading, 32K tokens without R9700 ~340 tok/s with R9700 ~1,000 tok/s 3x decode without R9700 ~16 tok/s with R9700 ~33 tok/s 2x
The same 192 GB Ryzen AI MAX+ PRO 495. Without the R9700: Framework's published run. With the R9700: the Lucebox Zero 495.

Long prompts gain the most. At 32K tokens, the Ryzen AI MAX+ alone reads about 340 tokens a second; with the R9700 it reads about 1,000. Writing goes from about 16 tokens a second to 33. Put together, a 30K-token prompt gets its answer in about 40 seconds instead of almost two minutes.

Without the R9700With the R9700Faster
Long prompt and answer (30K tokens in, 256 out)almost 2 minabout 40 sabout 3x
Prompt reading at 32K tokensabout 340 tok/sabout 1,000 tok/s3x
Prompt reading at 2K tokensabout 350 tok/sabout 800 tok/s2.3x
Decodeabout 16 tok/sabout 33 tok/s2x

Without the R9700: Framework's published figures, with the DwarfStar engine, antirez's Q2 and the DSpark drafter. With the R9700: our engine and our own quantized model, which is about as close to the full model. Framework reads 2K new tokens on top of existing context; we read the whole prompt.

With 192 GB, every expert stays in memory

The 192 GB model has room for all of the experts. The R9700 keeps the layers every token goes through, the drafter and the experts used most. The integrated GPU of the Ryzen AI MAX+ keeps all the others, around 150 GB, so nothing comes from the SSD while the model runs.

That lets the engine do things the SSD tier rules out:

DeepSeek V4.1 Flash with 128 GB and with 192 GB of unified memory Decode writing code: about 26 tok/s with 128 GB, about 49 tok/s with 192 GB. Prompt reading at 14K to 25K tokens: about 120 tok/s with 128 GB, about 1,000 tok/s with 192 GB. decode, writing code 128 GB ~26 tok/s 192 GB ~49 tok/s prompt reading, 14K to 25K tokens 128 GB ~120 tok/s 192 GB ~1,000 tok/s
The 128 GB and 192 GB models, same quantized model and drafter. Prompt reading at 20K tokens on 128 GB and 14K to 25K on 192 GB.

Coding agents benefit too: with the prefix cache, each new turn of a session is read at about 750 tokens a second. These runs use a smaller build of the drafter, about 5 GB instead of 8, which leaves the R9700 room to read the prompt in bigger chunks, and an expert placement fitted to this configuration.

With 128 GB: three tiers

On the 128 GB model the experts no longer all fit in memory, so the SSD becomes a third tier. The three are far apart in speed. The R9700 reads its 32 GB at 640 GB/s, the integrated GPU reads unified memory at around 240 GB/s, and the SSD holds everything but reads only a few GB a second. DeepSeek V4.1 Flash picks 6 of its 384 experts for every token in each of its 40 layers, so where each expert lives decides how fast the model writes.

The 128 GB figures here come from the Ryzen AI MAX+ 395 build with the same R9700. The Zero 495 keeps that GPU and adds faster memory.

Where DeepSeek V4.1 Flash's routed experts live on one Lucebox The R9700 holds about 1,000 experts (10 GB) and reads at 640 GB/s; unified memory holds about 9,000 experts (100 GB) and reads at about 240 GB/s; the SSD holds about 5,000 experts (50 GB) and reads at a few GB/s. reads at R9700 ~10 GB, ~1,000 experts 640 GB/s Unified memory ~100 GB, ~9,000 experts ~240 GB/s SSD ~50 GB, ~5,000 experts ~4.5 GB/s
Where the experts live on the 128 GB model at a 128K context, and how fast each tier reads. Unified memory includes system memory pinned so it is never swapped out.

Experts are placed by how often they are used. Most tokens go to a small group of experts, so those sit in the fast tiers and the long tail goes to the SSD. With a 4K-token prompt, the R9700 holds 6% of the experts and serves almost a quarter of the calls, while the SSD holds a third of them and serves less than a tenth.

Share of experts each tier holds against share of routed calls it serves With 4K tokens of prompt: the R9700 holds 6% of the experts and serves 24% of the calls, unified memory holds 61% and serves 68%, the SSD holds 32% and serves 9%. Routed calls served Experts held R9700 24% of calls 6% of experts Unified memory 68% of calls 61% of experts SSD 9% of calls 32% of experts
With a 4K-token prompt: each tier's share of the experts against its share of the calls.

From under 1 to 26 tok/s

Our first version kept the hottest experts on the R9700 and read the rest from the SSD. It gave the right answers, at under one token a second. Each step after that kept the answers the same and removed one bottleneck.

Decode speed of DeepSeek V4.1 Flash on one Lucebox, step by step Decode speed writing code, first request after start: R9700 + SSD under 1 tok/s, unified-memory tier about 3, ranked placement about 4, expert cache about 9, fused + DSpark about 11, pinned host memory about 23, Lucebox placement about 26 tok/s. R9700 + SSD <1 tok/s + unified memory ~3 tok/s + ranked placement ~4 tok/s + expert cache ~9 tok/s + fused + DSpark ~11 tok/s + pinned memory ~23 tok/s + Lucebox placement ~26 tok/s
Decode speed writing code on the 128 GB model, measured as each step landed. The last bar is today's engine at a 128K context.

Long prompts

The 128 GB profile reads prompts in 4K-token chunks and runs consecutive chunks layer by layer, so each expert from the SSD is read once per pass instead of once per chunk. It reads about 120 tokens a second from 4K to 20K tokens. Decode slows as the context grows, because every new token looks back over more history.

Decode speed by prompt length at the 128K default context Short prompts about 26 tok/s, 4K tokens of prompt about 17 tok/s, 20K tokens about 14 tok/s; prompt reading about 120 tok/s at 4K and 20K. prefill short ~26 tok/s - 4K tokens ~17 tok/s ~120 tok/s ~20K tokens ~14 tok/s ~120 tok/s
Decode speed by prompt length on the 128 GB model, with prompt reading on the right.

In agent sessions, the prefix cache keeps what was already read, so later turns start answering in seconds instead of close to a minute.

Loading 383 GB

On the 128 GB model, about 120 GB of the model is read before the first token: the layers every token uses and the experts kept in memory. The engine reads them on eight threads straight from the SSD, past the operating system's file cache, so loading doesn't push everything else out of memory. The server is ready in under a minute, down from five.

Time from start to ready, DeepSeek V4.1 Flash on one Lucebox Page cache dropped before each start. Before: about 5 minutes. With threaded direct reads: under a minute. faster before ~5 min now <1 min ~6x
Time from start to ready on the 128 GB model, with the file cache emptied before each start.

A better quantized model, from Geometric

Geometric built a second quantized model for our engine, tuned for the quality of the answers. Two choices set it apart:

On a set of 177 hard reasoning questions it answers 164 correctly. antirez's Q2, the same size, gets 141.

Correct answers out of 177 hard reasoning questions Geometric's model 164, with the Lucebox profile and DSpark 158, 10 GB smaller 155, antirez Q2 141. Geometric 164 Geometric, with DSpark 158 Geometric, 10 GB smaller 155 antirez Q2 141
Correct answers out of 177 on Geometric's hard reasoning set, greedy decoding, high reasoning effort. "With DSpark" adds the Lucebox profile and the drafter.

With the drafter it scores a little lower, mostly because a few long answers run into the 16K-token limit. Geometric also contributed the engine work the model needs: the MXFP8 type, faster IQ2 and IQ3 expert kernels for both GPUs, and a verify step whose drafted tokens match plain decoding exactly.

It runs on both memory sizes. It writes about as fast as ours but reads prompts more slowly, because our fast matrix-core path doesn't cover its formats yet. Pick our quantized model for speed and Geometric's for the best answers.

Run it

On a Lucebox, open the dashboard, pick DeepSeek V4.1 Flash on the Engine page and press Install. It downloads the model and the drafter, about 390 GB, so check that there is room, then builds the engine and checks every file. Start server has it answering in about a minute.

Elsewhere, the engine runs it with one profile for each memory size. With 128 GB:

luce_server DeepSeek-V4.1-Flash-ROCMFP2S.gguf \
  --profile ds41-lucebox \
  --draft DeepSeek-V4.1-Flash-DSpark-draft-MXFP4-Q8.gguf

With 192 GB:

luce_server DeepSeek-V4.1-Flash-ROCMFP2S.gguf \
  --profile ds41-gorgon \
  --draft DeepSeek-V4.1-Flash-DSpark-draft-MXFP4-Q8.gguf \
  --ds4-expert-placement placement.json \
  --ds4-router-bias router_bias.bin

Our quantized models and the DSpark drafter are on Hugging Face. For the best answers, use Geometric's ds41-lucebox-final.gguf from the same repository in place of ROCMFP2S; ours reads prompts faster, and it is the one measured in this post. Run it as the Lucebox model service so the unified-memory tier can pin its memory. On the 192 GB model the integrated GPU maps around 150 GB of experts, so raise its GTT limit; the engine docs give the values we use and every setting the profiles turn on.

Which model to pick

A Lucebox runs all three, one at a time, or Qwen next to DeepSeek V4 in the Lucebox Mix. DeepSeek V4.1 Flash is the largest and newest, at up to 49 tok/s on code with 192 GB. DeepSeek V4 Flash writes code at over 50 tok/s and takes a third of the disk space. Qwen 3.8 27B is the fastest, around 200 tok/s on code, for quick answers and coding agents that call the model often. Pick V4.1 for the largest model, and V4 or Qwen when speed matters more.

What is next

On the 128 GB model, decode still waits on the SSD for experts nobody predicted, and prompt reading uses only a small part of what the R9700 can compute. On the 192 GB model, the next step is the matrix-core prompt path for Geometric's formats, so the better model reads prompts as fast as ours.

Hardware and setup

192 GBAMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ PRO 495 with 192 GB of unified memory, Ubuntu, ROCm 7.2, --profile ds41-gorgon
128 GBAMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ 395 with 128 GB of unified memory (the build before the Zero 495), Crucial P310 2 TB SSD, Ubuntu, ROCm 7.2, --profile ds41-lucebox
ModelDeepSeek V4.1 Flash: our quantized model ROCMFP2S and the DSpark drafter, and Geometric's ds41-lucebox-final.gguf, all on Hugging Face
WorkloadCode writing and prompts up to about 30K tokens, 256-token answers, greedy decoding

Bottom line

DeepSeek V4.1 Flash is too big for any one chip in the box, so the engine uses all of them. The R9700 does the work every token needs and keeps the experts used most, unified memory holds the rest, and on the 128 GB model the SSD keeps the long tail and is read ahead of time. With 192 GB, that makes the model up to three times faster than the same chip without the R9700, at up to 49 tok/s on one desktop.


Measured by us and rounded. The 192 GB figures come from a Ryzen AI MAX+ PRO 495 with 192 GB in October 2026, the 128 GB figures from the Ryzen AI MAX+ 395 build with the same R9700 in September 2026, with the file cache emptied before each start. Decode speeds are for the first request after start. The figures without the R9700 are Framework's, from their post of September 30, 2026. The quality results for Geometric's model are Geometric's.

Run DeepSeek V4.1 on your Lucebox

Install it from the Engine page of your dashboard. The engine is open source.

GitHub DeepSeek V4 post Discord