Your data stays private
Client files, code and unreleased ideas never reach a cloud model. Runs fully offline, stores no prompts, needs no account with us.Runs fully offline. No prompts stored, no account with us.
Hardware
Client files, code and unreleased ideas never reach a cloud model. Runs fully offline, stores no prompts, needs no account with us.Runs fully offline. No prompts stored, no account with us.
Point Claude Code, Codex, OpenCode or Pi at its Anthropic and OpenAI compatible API. Repos, tickets and docs stay on your network.Claude Code, Codex or OpenCode on its OpenAI and Anthropic compatible API.
Research bots, automations and agent teams, running 24/7. Pay once, share it with your team, and forget per-token bills and rate limits.Agents running 24/7. Pay once, no per-token bills or rate limits.
The engine, custom kernels, speculative decoding and tuned models come installed for this exact hardware. Plug it in, pair it, and it runs at full speed.Engine, kernels and tuned models come installed for this hardware.
Performance
Unified memory fits big models and a discrete GPU runs them fast, so the best setup pairs the two. DeepSeek V4 Flash has 284B parameters but uses only a few experts per token. The GPU holds what every token needs, unified memory holds the rest, and both work at once.
1 / 2One big model across both
1Each token uses a few experts
4 of 256 experts per layer, most on the GPU
2Both processors work on every token
Both memories read in parallel
3Only small results move
The cache never leaves the GPU
Dense
From one user's DFlash2 run on a long prompt. DGX Spark and Mac M5 Max are public records with their own setups.
DetailsMoE
Measured on Lucebox at 2K context. DGX Spark and Mac M5 Max are public records with their own setups.
DetailsMeasured on the Ryzen AI MAX+ 395 build with the same Radeon AI PRO R9700.
Built in partnership with
Setup
Visit lucebox.com/setup from a nearby laptop or Android device. There is no app or CLI to install.
The onboarding page sends Wi-Fi, account, and optional Tailscale settings to the box over encrypted Bluetooth.
The local dashboard checks the machine, installs the qualified model profile, and starts the private API.
Enable Lucebox Connect once, then open any supported app from Manage with a single click.
Tokens per app and model, recent requests, and live machine meters. Prompts and replies are never stored.
Order
Open source, free forever
Starting at $5,999€6,499£5,999 List price $6,999€7,499£6,999 First 200 units, $1,000 offFirst 200 units, €1,000 offFirst 200 units, £1,000 off USD per machine / taxes, duties and shipping includedUSD per machine / shipping and duties includedEUR per machine / taxes, duties and shipping includedGBP per machine / taxes, duties and shipping included Taxes and shipping includedShipping and duties includedTaxes and shipping includedTaxes and shipping included
Unified memoryLPDDR5X-8533, 273 GB/s
StorageNVMe SSD
ClusteringOptional
One kit per machine. To cluster two Luceboxes, add the kit to both. One kit also links a Lucebox to another 100 GbE machine.
Business For teams and companies: 12 months of Lucebox Engine Pro, white-glove install and onboarding, and direct chat support with the engineering team.
Business details (opens in a new tab)Community
I like your stuff so far, keep going
this guy just cracked 134 tok/s on qwen 3.5-27b dense and 73 on new qwen 3.6-27b on a single 3090. open source moves at godspeed in 2026.
Interesting run w/ Dflash from the lucebox-hub guys
speculative PREFILL?????
I have tested some LLM server software for home PCs for Linux and Windows. Fastest and best for running home is Linux running 145 t/s, Lucebox. @pupposandro @luceboxai
PFlash just killed the 4-minute blank screen problem. 128K token prefill in 25 seconds, same GPU, same model, no compromises
Consumer-grade GPUs actually have sufficient hardware potential, general-purpose frameworks just waste most of it on overhead. Lucebox releases that potential through hand-written kernels, letting even a 2020 RTX 3090 rival Apple's latest chips on efficiency.
Crazy I was litteraly wondering how can I increase my token speed 10 min ago
Crazy what @pupposandro just dropped on Qwen3.5-27B. 207 tok/s on a single 3090 with Q4_K_M and full DFlash speculative? Chinese labs + ggml hacks are just cooking on consumer hardware right now. This is the kind of local win I like to see.
RTX 3090 ready for a new life! Bringing it to @luceboxai team to make some experiments together
github.com/Luce-Org/lucebox-hub looks promising as a way to run "dense" models (eg. Qwen 27B) more efficiently. It's janky, but on my 5090 laptop it seems to be ~2x more tok/s than llama.cpp
This is very very good work BRAVO. I love it
Nearly 10x faster! After finishing Decoding, it starts cranking through Prefill. The previous DFlash was already stunning enough, and now they've added PFlash. Speculative prefill, up to 10x speedup. Go try it right now.
First time reading about speculative prefill, and it's crazy. 257s down to 24s for a 128K prompt on a single RTX 3090. Great article, definitely go ahead and give this a read.
impressive... so this is what it looks like when you focus on a set ram limit and optimizing for a single model above everything
Our open source inference engine is used by engineers at
Blog
DeepSeek V4.1 Flash is a 383 GB model. The Lucebox Zero 495 runs it across the Radeon AI PRO R9700, unified memory and the SSD. With the R9700 it is up to 3x faster than the same Ryzen AI MAX+ PRO 495 alone, at up to 49 tok/s writing code.
At the highest load each machine served, one Lucebox answered 58 image questions a minute (Qwen3.8-27B on its AMD Radeon AI PRO R9700 and the 284B DeepSeek V4 Flash Vision on its Strix Halo, at the same time) and an NVIDIA DGX Spark running llama.cpp 18. With the same Qwen model file and eight questions at once, the R9700 alone finishes 2.4x sooner.
AI inference and cooling designed together on the R9700: the model forecasts the work ahead, a learned thermal model picks the slowest safe fan speed. Chat silent, agent loops 29% and batch work 25% quieter, same peak temperatures and throughput.
Asymmetric parallelism splits the full 284B model across the R9700 and the Strix Halo: 3.63× the decode speed of one DGX Spark, for 36% less than two.
Cold 64-token KV chunks page to host RAM bit-exact, so decode holds at 38.6 tok/s from 64K to 256K with unchanged accuracy.
The full 284B model runs from 128 GB of unified memory: up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.
FAQ