Backed by Y Combinator

Your Inference Workstation.

Order now →

Partners

AMD
Google Gemma
Alibaba Qwen
poolside

Hardware

Private, always on, ready out of the box.

Unified memory
192 GB
128 or 192 GB,
Ryzen AI MAX+ PRO 495 at 273 GB/s
Dedicated GPU
32 GB
AMD Radeon AI PRO R9700,
GDDR6 at 640 GB/s
Storage
4 TB
2 or 4 TB NVMe, up to 7,400 MB/s read
Connectivity
10 GbE
Wi-Fi 7, 2× USB4, optional 100 GbE cluster kit
Power supply
1,000 W
80 PLUS Platinum, SFX
Warranty
1 year
Parts and labor,
after 72 hours of burn-in
01

Your data stays private

Client files, code and unreleased ideas never reach a cloud model. Runs fully offline, stores no prompts, needs no account with us.Runs fully offline. No prompts stored, no account with us.

02

Built for coding agents

Point Claude Code, Codex, OpenCode or Pi at its Anthropic and OpenAI compatible API. Repos, tickets and docs stay on your network.Claude Code, Codex or OpenCode on its OpenAI and Anthropic compatible API.

03

Always on, no meter

Research bots, automations and agent teams, running 24/7. Pay once, share it with your team, and forget per-token bills and rate limits.Agents running 24/7. Pay once, no per-token bills or rate limits.

04

Optimized for speed from the startOptimized for speed

The engine, custom kernels, speculative decoding and tuned models come installed for this exact hardware. Plug it in, pair it, and it runs at full speed.Engine, kernels and tuned models come installed for this hardware.

Check full specs

Performance

Two processors, one engine, built for speed.

Only a heterogeneous setup gives you big models and speed.

Unified memory fits big models and a discrete GPU runs them fast, so the best setup pairs the two. DeepSeek V4 Flash has 284B parameters but uses only a few experts per token. The GPU holds what every token needs, unified memory holds the rest, and both work at once.

1 / 2One big model across both

  1. 1Each token uses a few experts

    4 of 256 experts per layer, most on the GPU

  2. 2Both processors work on every token

    Both memories read in parallel

  3. 3Only small results move

    The cache never leaves the GPU

Over 2× the reported DGX Spark decode speed, at the same price.

Dense

Qwen3.8-27B

Prefill
903 tok/s
Decode
61.3 tok/s

From one user's DFlash2 run on a long prompt. DGX Spark and Mac M5 Max are public records with their own setups.

Details
Prompt processing tok/s, higher is better
  1. Lucebox 903
  2. DGX Spark Not reported
  3. Mac M5 Max 682.3
Token generation tok/s, higher is better
  1. Lucebox 61.3
  2. DGX Spark 36.59
  3. Mac M5 Max 56.4

MoE

DeepSeek V4 Flash

Prefill
788 tok/s
Decode
86 tok/s

Measured on Lucebox at 2K context. DGX Spark and Mac M5 Max are public records with their own setups.

Details
Prompt processing tok/s, higher is better
  1. Lucebox 788
  2. DGX Spark 1,076.8
  3. Mac M5 Max 790.18
Token generation tok/s, higher is better
  1. Lucebox 86
  2. DGX Spark 35.3
  3. Mac M5 Max 39.35

Measured on the Ryzen AI MAX+ 395 build with the same Radeon AI PRO R9700.

Built in partnership with AMD
See all results

Setup

Easy setup. Use it in the apps on your laptop.

Open setup
02 / 11
Setup
Manage
UI, tabs and steps are clickable
  1. 01

    Open the setup page

    Visit lucebox.com/setup from a nearby laptop or Android device. There is no app or CLI to install.

  2. 02

    Pair and connect

    The onboarding page sends Wi-Fi, account, and optional Tailscale settings to the box over encrypted Bluetooth.

  3. 03

    Continue in Manage

    The local dashboard checks the machine, installs the qualified model profile, and starts the private API.

  4. 04

    Connect your apps

    Enable Lucebox Connect once, then open any supported app from Manage with a single click.

  5. 05

    Monitor your Lucebox

    Tokens per app and model, recent requests, and live machine meters. Prompts and replies are never stored.

Works with the tools you already use
Claude Code
Codex
OpenCode
Hermes
OpenClaw
Pi
Oh My Pi
Open WebUI
See the platform

Order

Order your Lucebox Zero 495.

Developer

For individuals

Open source, free forever

Starting at $5,999 List price $6,999 First 200 units, $1,000 off USD per machine / taxes, duties and shipping included Taxes and shipping included

  • Lucebox Engine, open source, with tuned models preinstalled
  • 1-year warranty on parts and labor
  • Worldwide shipping included
  • All taxes and duties included
Order now

Business For teams and companies: 12 months of Lucebox Engine Pro, white-glove install and onboarding, and direct chat support with the engineering team.

Business details (opens in a new tab)
Expected deliveryJanuary 2027
WorldwideTaxes, duties and shipping included
1 yearWarranty covering parts and labor

Community

Proudly open source, and built in public.

Ahmad@TheAhmadOsman

I like your stuff so far, keep going

Sudo su@sudoingX

this guy just cracked 134 tok/s on qwen 3.5-27b dense and 73 on new qwen 3.6-27b on a single 3090. open source moves at godspeed in 2026.

Lotto@LottoLabs

Interesting run w/ Dflash from the lucebox-hub guys

Rijndael@rot13maxi

speculative PREFILL?????

Riku Pasonen@Raitziger

I have tested some LLM server software for home PCs for Linux and Windows. Fastest and best for running home is Linux running 145 t/s, Lucebox. @pupposandro @luceboxai

fahd Mirza@fahdmirza

PFlash just killed the 4-minute blank screen problem. 128K token prefill in 25 seconds, same GPU, same model, no compromises

Geek Lite@QingQ77

Consumer-grade GPUs actually have sufficient hardware potential, general-purpose frameworks just waste most of it on overhead. Lucebox releases that potential through hand-written kernels, letting even a 2020 RTX 3090 rival Apple's latest chips on efficiency.

Takyon∞@Takyon

Crazy I was litteraly wondering how can I increase my token speed 10 min ago

Kyz2ren@ky2renzz

Crazy what @pupposandro just dropped on Qwen3.5-27B. 207 tok/s on a single 3090 with Q4_K_M and full DFlash speculative? Chinese labs + ggml hacks are just cooking on consumer hardware right now. This is the kind of local win I like to see.

Ivan Fioravanti@ivanfioravanti

RTX 3090 ready for a new life! Bringing it to @luceboxai team to make some experiments together

vitalik.eth@VitalikButerin

github.com/Luce-Org/lucebox-hub looks promising as a way to run "dense" models (eg. Qwen 27B) more efficiently. It's janky, but on my 5090 laptop it seems to be ~2x more tok/s than llama.cpp

CAPET@Capetlevrai

This is very very good work BRAVO. I love it

nash_su@nash_su

Nearly 10x faster! After finishing Decoding, it starts cranking through Prefill. The previous DFlash was already stunning enough, and now they've added PFlash. Speculative prefill, up to 10x speedup. Go try it right now.

AJ@ItsmeAjayKV

First time reading about speculative prefill, and it's crazy. 257s down to 24s for a 128K prompt on a single RTX 3090. Great article, definitely go ahead and give this a read.

Twon.@Web3Twon

impressive... so this is what it looks like when you focus on a set ram limit and optimizing for a single model above everything

Our open source inference engine is used by engineers at

Netflix
Intel
Google
Microsoft
60+Contributors
735+Pull requests
90K+Model downloads
View on GitHub

Blog

We share what we learn.

An open Lucebox with its Radeon AI PRO R9700 on a wooden deck under a starry sky between a silver pyramid and the DeepSeek whale, with the Lucebox and Geometric logos above

Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy: up to 3x faster with the R9700

DeepSeek V4.1 Flash is a 383 GB model. The Lucebox Zero 495 runs it across the Radeon AI PRO R9700, unified memory and the SSD. With the R9700 it is up to 3x faster than the same Ryzen AI MAX+ PRO 495 alone, at up to 49 tok/s writing code.

The Lucebox tower on a wooden deck under a starry sky, the Qwen bear holding a magnifying glass over a printed bar chart and the DeepSeek whale with a printed diagram

Vision LLM inference: Lucebox has 3.2x the throughput of NVIDIA DGX Spark

At the highest load each machine served, one Lucebox answered 58 image questions a minute (Qwen3.8-27B on its AMD Radeon AI PRO R9700 and the 284B DeepSeek V4 Flash Vision on its Strix Halo, at the same time) and an NVIDIA DGX Spark running llama.cpp 18. With the same Qwen model file and eight questions at once, the R9700 alone finishes 2.4x sooner.

An AMD Radeon AI PRO R9700 and a 120 mm case fan on a wooden deck under a starry sky

Inference-aware cooling: 29% less fan speed on the AMD R9700, same temperatures

AI inference and cooling designed together on the R9700: the model forecasts the work ahead, a learned thermal model picks the slowest safe fan speed. Chat silent, agent loops 29% and batch work 25% quieter, same peak temperatures and throughput.

Cover of the DeepSeek V4 Flash report: Lucebox against the NVIDIA DGX Spark

3.63× the DGX Spark on DeepSeek V4 Flash

Asymmetric parallelism splits the full 284B model across the R9700 and the Strix Halo: 3.63× the decode speed of one DGX Spark, for 36% less than two.

Cover of the KVFlash report: paged KV cache on a 24 GB GPU

KVFlash: 256K context with 72 MiB of KV

Cold 64-token KV chunks page to host RAM bit-exact, so decode holds at 38.6 tok/s from 64K to 256K with unchanged accuracy.

Cover of the DeepSeek V4 Flash on Strix Halo report

DeepSeek V4 Flash on the Ryzen AI MAX+ 395

The full 284B model runs from 128 GB of unified memory: up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.

Read all articles

FAQ

Frequently Asked Questions

Mission

We're here to make powerful AI something you own, not something you rent. Our mission is to commoditize access to fast, powerful local models, through hardware and software co-design.