Engine / open source

Lucebox Engine

The open source inference engine on every Lucebox, tuned for its R9700 and Ryzen AI MAX+. Apache 2.0, on GitHub.

Models

What it runs today.

Lucebox Mix

Default

Qwen for speed and DeepSeek V4 for deeper thinking, at once.

Speed Quality 116.7 GBQwen + DeepSeek V4 together

DeepSeek V4 Flash

Careful answers for agents, research and long documents.

Speed Quality 101.5 GBR9700 + Strix Halo128K context

Qwen 3.8 27B

The fastest: code, quick answers, agents that call it often.

Speed Quality 15.2 GBR9700, 8 at once128K context

DeepSeek V4.1 Flash

The newest and deepest DeepSeek, with a long 128K context.

Speed Quality 368.4 GBR9700 + Strix Halo, SSD with 128 GB128K context

Install any of them from Manage, or load any GGUF you bring. The box runs Ubuntu Server with ROCm and root access, so other open source engines install as on any Linux machine.

Speed

Decode and prefill per model.

Decode tok/s, one user
  1. Qwen 3.8 27B 227.8
  2. DeepSeek V4 Flash 86
  3. DeepSeek V4.1, 192 GB 49
  4. DeepSeek V4.1, 128 GB 26
Prefill tok/s, one user
  1. Qwen 3.8 27B 945
  2. DeepSeek V4 Flash 788
  3. DeepSeek V4.1, 192 GB ~1,000
  4. DeepSeek V4.1, 128 GB ~120

The best run in each report: decode while writing code, prefill on prompts of 1.4K to 32K tokens. Qwen runs on the R9700, DeepSeek across the R9700 and the Ryzen AI MAX+.

Qwen 3.8 27BDeepSeek V4 FlashDeepSeek V4.1 Flash

Batching

Throughput as users are added.

Total throughput tok/s, higher is better
  • Qwen 3.8 27B, code2.8× at 5 users
  • Qwen 3.8 27B, images2.3× at 8 users
  • DeepSeek V4 Flash, text2.0× at 4 users
Total throughput, tok/s, higher is better
Series1 users2 users3 users4 users5 users6 users7 users8 users
Qwen 3.8 27B, code106.6188.8209.4262.1300.9Not measuredNot measuredNot measured
Qwen 3.8 27B, images84Not measuredNot measured170Not measuredNot measuredNot measured195
DeepSeek V4 Flash, text23.835.442.948.4Not measuredNot measuredNot measuredNot measured

Tokens per second across the box, prompt reading included, so one user sits below the decode peaks above. DeepSeek runs on the Ryzen AI MAX+ alone, without its drafter.

Continuous batchingImage questions

Estimate, Qwen 3.8 27B coding agents

10 agents, 3 to 5 developers

Agents pause between turns to run tools and wait for input. If each writes 30% of the day, 10 agents need more than 5 requests at once under 5% of the time, and at 5 each still gets about 60 tok/s. At 2 to 3 agents per developer, that is 3 to 5 developers.

Manage serves up to 8 at once, so a busy moment slows agents instead of queueing them.

Developers per box by agents each runs at once
  1. 1 agent each 10
  2. 2 agents each 5
  3. 3 agents each 3
  4. 4 agents each 2

Speculative

Speculative decoding and prefill.

PFlash

Prefill

A small model picks the parts of a long prompt worth reading. Lossy.

  • DeepSeek V4 Flash
Report

The decode drafters are lossless: the model keeps only the tokens it would have written itself.

Optimizations

Engine optimizations.

Tool prefix cache

Prefill

Reuses the saved tool definitions on every agent turn.

Report

KVFlash

Memory

Pages cold KV to host RAM, bit‑exact, for long context on one card.

Report

Three-tier experts

Memory

Places experts in VRAM, unified memory or on the SSD by use.

Report

Two-processor split

Decode

Runs one model on the R9700 and the Ryzen AI MAX+ at once.

Report

Continuous batching

Serving

Paged attention and one scheduler for many requests at once.

Report

Layered image prefill

Serving

Reads an image prompt a few layers at a time, so other users keep writing.

Report

Two processors

One big model, or two at once.

Prefill at 32K tok/s
  1. Chip alone ~340
  2. With R9700 ~1,000

3× faster prompt reading on DeepSeek V4.1 Flash.

Decode tok/s
  1. Chip alone ~16
  2. With R9700 ~33

2× faster writing on the same model.

Two models at once % of solo speed
  1. Qwen, 8 users 88%
  2. DeepSeek, 4 users 99%

Each keeps its speed while the other serves: own chip, own memory.

Chip alone is Framework's published run on the same 192 GB Ryzen AI MAX+ PRO 495, with its own engine and quantization.

DeepSeek V4.1 FlashTwo models at once

Compared

Against other engines and machines.

Qwen3.8-27B on the R9700

6.4× faster answers

Decode speed on HumanEval coding prompts, averaged. The same drafter in llama.cpp reaches 54.6 tok/s.

Token generation tok/s, higher is better
  1. Plain decoding 32.3
  2. llama.cpp 54.6
  3. Lucebox engine 208.1
Details

DeepSeek V4 Flash, 284B

Over 2× a DGX Spark

One model split across the R9700 and the Ryzen AI MAX+ unified memory. Measured on the Ryzen AI MAX+ 395 build.

Token generation tok/s, higher is better
  1. DGX Spark 35.3
  2. Mac M5 Max 39.35
  3. Lucebox 86
Details

Qwen3.8-27B, image questions

2.4× a DGX Spark

Eight users asking about charts at once, with the same model file, projector and drafter on both machines.

Image answers tok/s, higher is better
  1. Spark, plain 65
  2. Spark, DFlash2 81
  3. Lucebox 195
Details