<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Lucebox engineering blog</title><description>Engineering notes on local inference, heterogeneous computing, GPU kernels, speculative decoding, and hands-on benchmarks.</description><link>https://www.lucebox.com/</link><language>en-us</language><item><title>Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy: up to 3x faster with the R9700</title><link>https://www.lucebox.com/blog/deepseek-v41-lucebox-memory-hierarchy/</link><guid isPermaLink="true">https://www.lucebox.com/blog/deepseek-v41-lucebox-memory-hierarchy/</guid><description>DeepSeek V4.1 Flash is a 383 GB model. The Lucebox Zero 495 runs it across the Radeon AI PRO R9700, unified memory and the SSD. With the R9700 it is up to 3x faster than the same Ryzen AI MAX+ PRO 495 alone, at up to 49 tok/s writing code.</description><pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate></item><item><title>Continuous batching: Qwen3.8-27B at 301 tok/s across five clients, up to 1.97× llama.cpp</title><link>https://www.lucebox.com/blog/continuous-batching/</link><guid isPermaLink="true">https://www.lucebox.com/blog/continuous-batching/</guid><description>One scheduler contract serves Qwen3.8-27B with DFlash2 and DeepSeek V4 Flash: 106.6 tok/s at one client grows to 300.9 at five on an AMD R9700, up to 1.97× llama.cpp at equal concurrency, on stable slots and paged KV.</description><pubDate>Mon, 28 Sep 2026 12:00:00 GMT</pubDate></item><item><title>Lucebox × Geometric: DeepSeek V4 Flash 0731 reaches 32.7 tok/s on AMD Strix Halo</title><link>https://www.lucebox.com/blog/deepseek-v4-flash-0731/</link><guid isPermaLink="true">https://www.lucebox.com/blog/deepseek-v4-flash-0731/</guid><description>DeepSeek V4 Flash 0731 on 128 GB AMD Strix Halo: 82/92 quality, up to 32.7 tok/s decode, and 173 tok/s sparse prefill at 60K context.</description><pubDate>Mon, 28 Sep 2026 12:00:00 GMT</pubDate></item><item><title>Vision LLM inference: Lucebox has 3.2x the throughput of NVIDIA DGX Spark</title><link>https://www.lucebox.com/blog/vision-llm-inference/</link><guid isPermaLink="true">https://www.lucebox.com/blog/vision-llm-inference/</guid><description>At the highest load each machine served, one Lucebox answered 58 image questions a minute (Qwen3.8-27B on its AMD Radeon AI PRO R9700 and the 284B DeepSeek V4 Flash Vision on its Strix Halo, at the same time) and an NVIDIA DGX Spark running llama.cpp 18. With the same Qwen model file and eight questions at once, the R9700 alone finishes 2.4x sooner.</description><pubDate>Fri, 25 Sep 2026 12:00:00 GMT</pubDate></item><item><title>Inference-aware cooling: 29% less fan speed on the AMD R9700, same temperatures</title><link>https://www.lucebox.com/blog/fan-forecast/</link><guid isPermaLink="true">https://www.lucebox.com/blog/fan-forecast/</guid><description>AI inference and cooling designed together on the AMD Radeon AI PRO R9700: the model forecasts the work ahead and a learned thermal model turns it into the slowest safe fan speed. Chat silent, agent loops 29% and batch work 25% quieter, at the same peak temperatures and identical throughput.</description><pubDate>Thu, 17 Sep 2026 12:00:00 GMT</pubDate></item><item><title>ROCm beats Vulkan on Strix Halo</title><link>https://www.lucebox.com/blog/rocm-beats-vulkan-strix-halo/</link><guid isPermaLink="true">https://www.lucebox.com/blog/rocm-beats-vulkan-strix-halo/</guid><description>DeepSeek V4 Flash on one Strix Halo, both engines measured on the same box on the same day: Lucebox ROCm is 17 to 47% faster on prefill and 28 to 47% faster on speculative decode than llama.cpp Vulkan v0.7.5 at 8K, 32K and 123K prompt tokens, with plain decode and the adaptive verify width measured on the same box.</description><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate></item><item><title>Ling 3.0 Flash on DGX Spark: Up to 141.9 tok/s with Adaptive DSpark and FlashKDA</title><link>https://www.lucebox.com/blog/ling3-flash-dgx-spark/</link><guid isPermaLink="true">https://www.lucebox.com/blog/ling3-flash-dgx-spark/</guid><description>Our Ling 3.0 Flash result on one DGX Spark: up to 36.4% faster prompt reading, near-tied ordinary generation, and 141.9 tok/s in the best matched DSpark case.</description><pubDate>Fri, 28 Aug 2026 12:00:00 GMT</pubDate></item><item><title>Qwen3.8-27B on the AMD R9700: up to 227 tok/s</title><link>https://www.lucebox.com/blog/qwen38-r9700/</link><guid isPermaLink="true">https://www.lucebox.com/blog/qwen38-r9700/</guid><description>The lucebox engine serves Qwen3.8-27B on a single AMD Radeon AI PRO R9700 with the DFlash2 block-diffusion drafter: up to 227 tok/s on code, 208 tok/s HumanEval average, and 3.8x llama.cpp decode running the same drafter, at 8-bit-class quality.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate></item><item><title>Tool prefix caching: 48× faster warm prefill for agent loops</title><link>https://www.lucebox.com/blog/tool-prefix-cache/</link><guid isPermaLink="true">https://www.lucebox.com/blog/tool-prefix-cache/</guid><description>Agent turns often resend thousands of tokens of tool definitions. Lucebox can now restore that stable prefix and process only the new conversation. On Qwen3.6-27B, median warm prefill was 1.04 seconds after a 50.35 second cold turn.</description><pubDate>Mon, 27 Jul 2026 12:00:00 GMT</pubDate></item><item><title>Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) Beats NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed</title><link>https://www.lucebox.com/blog/deepseek-v4-asymmetric-parallelism/</link><guid isPermaLink="true">https://www.lucebox.com/blog/deepseek-v4-asymmetric-parallelism/</guid><description>51.1 tok/s on Lucebox versus our 14.09 tok/s average on one NVIDIA DGX Spark for the full 284B model. The complete $5,999 Lucebox costs 36% less than two DGX Sparks.</description><pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate></item><item><title>DeepSeek V4 Flash: 284B model, up to 32 tok/s on AMD Ryzen AI MAX+ 395</title><link>https://www.lucebox.com/blog/deepseek-v4-strix-halo/</link><guid isPermaLink="true">https://www.lucebox.com/blog/deepseek-v4-strix-halo/</guid><description>The full DeepSeek V4 Flash target runs locally on AMD Ryzen AI MAX+ 395 with 128 GB unified memory, reaching up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.</description><pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate></item><item><title>Laguna XS 2.1 on a RTX 3090: 296 tok/s peak, 152 tok/s at 256K context</title><link>https://www.lucebox.com/blog/laguna-xs21/</link><guid isPermaLink="true">https://www.lucebox.com/blog/laguna-xs21/</guid><description>poolside&apos;s coding MoE with its official DFlash drafter: 296 tok/s peak at short context, a flat 152 tok/s at 256K tokens, and prefill at 3,500 tok/s. Lossless speculative decoding, KVFlash paging, and two model-agnostic engine optimizations run on one 24 GB card.</description><pubDate>Sat, 11 Jul 2026 12:00:00 GMT</pubDate></item><item><title>Luce KVFlash: 256K context with 72 MiB of KV on the GPU</title><link>https://www.lucebox.com/blog/kvflash/</link><guid isPermaLink="true">https://www.lucebox.com/blog/kvflash/</guid><description>On Qwen3.6-27B the KV cache costs 4.6 GiB at 256K and drags decode to 13 tok/s. KVFlash pages cold 64-token chunks to host RAM bit-exact, holding decode at 38.6 tok/s from 64K to 256K with unchanged accuracy.</description><pubDate>Fri, 12 Jun 2026 12:00:00 GMT</pubDate></item><item><title>Lucebox in a container: one image for every supported GPU</title><link>https://www.lucebox.com/blog/docker/</link><guid isPermaLink="true">https://www.lucebox.com/blog/docker/</guid><description>A prebuilt image spans the RTX 2080 Ti through RTX 5090. The fat-binary compile happens once in CI instead of on your box, with two host dependencies, self-tuning, and build provenance included.</description><pubDate>Mon, 08 Jun 2026 12:00:00 GMT</pubDate></item><item><title>Luce Spark: fit Qwen3.6 35B and Laguna XS.2 on a 16 GB GPU</title><link>https://www.lucebox.com/blog/spark/</link><guid isPermaLink="true">https://www.lucebox.com/blog/spark/</guid><description>A 33–35B MoE fires only a fraction of its experts per token but normally pays for all of them in VRAM. Spark keeps the active experts resident and swaps the rest, fitting Qwen3.6 35B-A3B and Laguna XS.2 on 16 GB GPUs.</description><pubDate>Fri, 05 Jun 2026 12:00:00 GMT</pubDate></item><item><title>Gemma 4 26B edges out DeepSeek V4 Flash (284B) on ds4-eval-92, at 5x the speed</title><link>https://www.lucebox.com/blog/gemma-vs-deepseek/</link><guid isPermaLink="true">https://www.lucebox.com/blog/gemma-vs-deepseek/</guid><description>On ds4-eval-92, Gemma 4 26B on a 24 GB RTX 5090 Laptop ties DeepSeek V4 Flash on a 192 GB Mac at 78.3% and decodes about five times faster.</description><pubDate>Wed, 20 May 2026 12:00:00 GMT</pubDate></item><item><title>Launch and tune Lucebox with real agent harnesses</title><link>https://www.lucebox.com/blog/client-harnesses/</link><guid isPermaLink="true">https://www.lucebox.com/blog/client-harnesses/</guid><description>Real-client profiles, launch scripts, and TQ3/DDTree results for OpenCode, Hermes, OpenClaw, Open WebUI, Codex, Claude Code, and Pi.</description><pubDate>Fri, 15 May 2026 12:00:00 GMT</pubDate></item><item><title>DFlash + PFlash on AMD Strix Halo: 2.5× end-to-end vs llama.cpp HIP</title><link>https://www.lucebox.com/blog/amd/</link><guid isPermaLink="true">https://www.lucebox.com/blog/amd/</guid><description>DFlash and PFlash on the Ryzen AI MAX+ 395 iGPU reach 26.85 tok/s speculative decode and 20.2 seconds of prefill at 16K, a 2.51× end-to-end speedup over vanilla llama.cpp HIP on the same silicon.</description><pubDate>Tue, 12 May 2026 12:00:00 GMT</pubDate></item><item><title>Laguna XS.2 on a 3090: 111 tok/s, 5.4x prefill, first MoE target for PFlash</title><link>https://www.lucebox.com/blog/laguna/</link><guid isPermaLink="true">https://www.lucebox.com/blog/laguna/</guid><description>Poolside Laguna XS.2 was ported into DFlash and PFlash as the first MoE target supported by PFlash, reaching about 107 tok/s decode and 5.4× faster 128K prefill than llama.cpp on one RTX 3090.</description><pubDate>Fri, 08 May 2026 12:00:00 GMT</pubDate></item><item><title>PFlash: 10× prefill speedup over llama.cpp at 128K on a RTX 3090</title><link>https://www.lucebox.com/blog/pflash/</link><guid isPermaLink="true">https://www.lucebox.com/blog/pflash/</guid><description>PFlash compresses a 128K prompt to 2.6K tokens with a small drafter before DFlash sees it, reducing cold time to first token from about 257 seconds to 24.8 seconds while preserving measured retrieval accuracy.</description><pubDate>Tue, 28 Apr 2026 12:00:00 GMT</pubDate></item><item><title>DFlash on ggml: up to 207 tok/s Qwen3.5-27B on a RTX 3090</title><link>https://www.lucebox.com/blog/dflash27b/</link><guid isPermaLink="true">https://www.lucebox.com/blog/dflash27b/</guid><description>A standalone C++ and ggml speculative decoder for Qwen3.5-27B Q4_K_M with a DFlash block-diffusion draft model and DDtree verifier, reaching up to 207 tok/s and supporting 128K context on 24 GB.</description><pubDate>Thu, 16 Apr 2026 12:00:00 GMT</pubDate></item><item><title>NVIDIA eGPU on macOS: RTX 3090 and 5090 benchmarks</title><link>https://www.lucebox.com/blog/egpu-myth/</link><guid isPermaLink="true">https://www.lucebox.com/blog/egpu-myth/</guid><description>tinygrad wrote an NVIDIA driver from scratch. We tested real models on an RTX 3090 over USB4 to measure whether an inexpensive eGPU dock can turn a Mac into an AI workstation.</description><pubDate>Tue, 14 Apr 2026 12:00:00 GMT</pubDate></item><item><title>Megakernel: Matching Apple Silicon Efficiency at 2x the Throughput on a RTX 3090</title><link>https://www.lucebox.com/blog/megakernel/</link><guid isPermaLink="true">https://www.lucebox.com/blog/megakernel/</guid><description>The first megakernel for hybrid DeltaNet and Attention LLMs fuses all 24 layers into one CUDA dispatch, reaching 1.87 tok/J and matching M5 Max efficiency at about twice the throughput on an RTX 3090.</description><pubDate>Mon, 13 Apr 2026 12:00:00 GMT</pubDate></item></channel></rss>