September 2026

By Davide Ciffa

Inference-aware cooling: 29% less fan speed on the AMD R9700, same temperatures

A fan controller only knows the past: the heat that has already happened. The AI model running on the card knows the future: how many tokens are queued, how long the prompt is, whether an agent is about to call again. On the AMD Radeon AI PRO R9700 we let the model tell the fan what is coming, so the fan spins for the work ahead instead of the heat behind. Chat runs with the blower at its quietest setting, agent loops at 29% less fan speed, batch work at 25% less, at the same peak temperatures. No power cap, no undervolt, no slower clocks.

An AMD Radeon AI PRO R9700 and a 120 mm case fan on a wooden deck under a starry sky
The R9700 blower is the loudest part of a Lucebox. The fix is not a bigger fan; it is telling the fan what the engine is about to do.
The same agent workload, twice, on one box: the stock firmware curve in the middle, the engine's forecast driving the fan on the right, and on the left what the model is writing while both fans decide what to do. The two fan panels are the recorded runs from the table below, replayed at 6× speed and lined up on the burst pattern; the text is the same workload's own output, streamed at the speed it arrived.
1,734 RPMchat: the blower never leaves its floor (stock firmware: 1,906, pulsing)
−29%agent loops: 2,343 vs 3,297 RPM at the same 88 C peak
−25%sustained batch: 2,686 vs 3,590 RPM, 90 vs 91 C peak

TL;DR

Why a fan controller is always late

The R9700 is a 300 W card with a 75 mm radial blower. Its firmware targets 80 C on the hotspot and pushes the fan to about 3,700 RPM whenever the die gets there. That policy is tuned for a server rack, and on a desk it produces the two sounds everyone knows: the pulse after every short request, when the fan ramps for heat that has already left, and the steady roar through an agent session, because the firmware cannot know that the pause between tool calls lasts eight seconds.

We recorded the stock behaviour on three synthetic traces that stand in for real use. Chat: two requests of 1,600 prompt tokens and 200 output tokens, then a 20 s pause. Agent: 30 s bursts of back-to-back requests with 8 s pauses. Sustained: back-to-back forever. Every run is six minutes, sampled at 1 Hz from the card's own sensors.

Blower RPM over six minutes of chat traffic Stock firmware pulses to about 2,245 RPM after each burst and settles near 1,900. The co-design stays flat at 1,734 RPM and never leaves its floor. 1,600 1,800 2,000 2,200 2,400 0 60 120 180 240 300 360 Seconds Blower RPM Stock firmware Co-design
Chat trace, blower RPM over six minutes. Stock firmware (grey) ramps 20 s after each burst, when the die is already back at 45 C. The co-design (gold) never moves.

The physics you cannot argue with

Before building anything, we measured what the fan actually buys. With the card pinned at 300 W we capped the blower at fixed speeds and watched the die settle.

Steady-state die temperature versus blower speed at 300 W At 300 W, 3,664 RPM holds 82 C, 2,821 RPM holds 88 C, and 2,376 RPM holds 93 C. Below about 2,400 RPM the die never settles: at 1,990 RPM it passed 100 C and the run was aborted after 70 seconds. 80 86 92 98 104 1,800 2,200 2,600 3,000 3,400 3,800 Blower RPM at 300 W Die temperature (C) firmware throttle our target never settles: 101 C in 70 s Steady state reached
Steady-state die temperature versus blower speed at 300 W, chassis fan on. Below about 2,400 RPM the die does not settle at all.

Two rules fall out of that curve. The first is a floor: under a sustained 300 W load, holding the die around 90 C needs 2,700 to 3,100 RPM, and nothing a controller does changes it. The second is why the floor is worth fighting for anyway: fan noise scales with roughly the fifth power of speed, so a 29% slower fan drops about 7 dB of sound power, most of the way to half as loud.

ΔL  =  50⋅log⁡10 ⁣(n1n0)\Delta L \;=\; 50 \cdot \log_{10}\!\left(\frac{n_1}{n_0}\right)

The fan law: the sound a fan makes grows with about the fifth power of its speed, so a small cut in RPM is a real cut in noise. We did not put a microphone on the box; the column below is what the law says about sound power, not a measured dBA.

What the measured speeds mean by the fan law. A drop of 10 dB is roughly half as loud to the ear.
BlowerChangeSound power (fan law)Where it comes from
1,906 → 1,734 RPM−9%−2.1 dBchat, stock against the forecast
3,031 → 2,751 RPM−9%−2.1 dBagent loop on a hot box, measured today
3,297 → 2,343 RPM−29%−7.4 dBagent loop, the widest pair we recorded
3,300 → 1,750 RPM−47%−13.8 dBfull blow against the blower's floor

The consequence for control is simple. Any second the fan spins faster than the die needs is wasted noise, and any second it spins slower than needed is a temperature overshoot the firmware will correct loudly. The only way to be right on both sides is to know what is coming.

Co-design: the engine forecasts its own heat

The inference engine has everything a forecast needs and a thermometer has none of it. So the server now publishes, once per second, a small JSON object: how many seconds of full-power work lie ahead.

busy ahead  =  tokens to readread rate+tokens left to writewrite rate⏟work already accepted  +  duty60s⋅(60s−known)⏟work likely to follow \text{busy ahead} \;=\; \underbrace{\frac{\text{tokens to read}}{\text{read rate}} + \frac{\text{tokens left to write}}{\text{write rate}}}_{\text{work already accepted}} \;+\; \underbrace{\text{duty}_{60s} \cdot (60\text{s} - \text{known})}_{\text{work likely to follow}}

The first half is arithmetic on what the server is already holding: the prompt tokens still to read, net of the prefix cache, and the tokens still to write, at the rates the engine measures for itself. The second half is the guess: a session that kept the card busy most of the last minute will probably keep doing it, and a request that arrives carrying tool definitions is an agent loop from its first token.

A final answer without a tool call, or a client that disconnects, tells the fan service the session is over, and the forecast drops to zero at once. Under concurrent serving the scheduler publishes the exact remaining work of every live slot. None of this reads a temperature: the engine stays portable, and the fan service owns the thermal side.

What the engine forecasts at the start of a request, against what the request really took Across four request shapes the forecast lands within a second of the real remaining time, except on a 24K-token prompt where it says 27 seconds against a real 30.4. Forecast at request start What it really took 0 7 14 21 28 35 s Long answer, first request 5.5 s 26.7 s Long answer, later ones 27.0 s 27.0 s RAG, 6K prompt 7.3 s 7.3 s Long context, 24K prompt 27.0 s 30.4 s
What the engine says at the moment a request arrives, against what the request really took. The forecast has to hold across shapes that behave nothing alike: a short prompt with a long answer, a 6K-token retrieval prompt, a 24K-token document. The one case it under-calls is the very first request of a shape it has never served, where it has no history to draw on; from the second request on it knows.

A thermal model at the sensor's noise floor

The forecast says how long the card will burn 300 W; the fan service has to say how hot the die gets, for a candidate fan speed, at the end of that horizon. We started with a hand-fitted two-node model, which is the physics: a fast hotspot term that follows power within seconds, and a slow heatsink node with a 45 s time constant.

Tdie=Theatsink+30∘ ⁣C⋅P300 WdTheatsinkdt=T∞−Theatsinkτ,τ≈45 s T_\text{die} = T_\text{heatsink} + 30^\circ\!\text{C} \cdot \frac{P}{300\,\text{W}} \qquad \frac{dT_\text{heatsink}}{dt} = \frac{T_\infty - T_\text{heatsink}}{\tau}, \quad \tau \approx 45\,\text{s}

The die is the heatsink plus a hotspot that follows power within seconds. The heatsink itself is slow: it drifts towards wherever the current power and fan speed would leave it, with a time constant of about 45 seconds. That lag is the whole opportunity.

T∞  ∝  PnT_\infty \;\propto\; \frac{P}{\sqrt{n}}

And where it ends up: the temperature rise scales with the power you pour in, divided by the square root of the fan speed. Doubling the fan does not halve the temperature; it buys about 30%. This is why the floor exists.

That model got the shape right and the number wrong: on held-out runs it under-predicted the hot samples by 4 to 6 C at the horizons the controller plans over, which is exactly how much the die overshot in the afternoon runs. So we replaced it with a learned one. It is a ridge regression, small enough to evaluate hundreds of times a second, and its inputs are what you would write down by hand: the die, memory and board temperatures now and 5 and 10 seconds ago, the power and fan speed the plan implies over the horizon, that power divided by the square root of that fan speed, and the share of the horizon spent reading a prompt rather than writing an answer, because a long prompt runs hotter than decoding at the same watts. The first fit used a narrower set of those inputs and stopped about a degree short; adding the rest took it to the sensor floor.

Mean absolute error of the die prediction on held-out hot samples The hand-fitted physics errs 3.2 to 6.5 C depending on how far ahead it looks. The model we ship errs 1.6 to 2.0 C, at the floor set by the sensor's own jitter of 2.4 C. 10 s ahead 30 s ahead 60 s ahead 0 1.4 2.8 4.2 5.6 7 C Hand-fitted physics 3.2 5.6 6.5 First learned model 2.1 3.2 3.4 The model we ship 1.6 1.8 2.0 Sensor jitter floor 2.4 2.4 2.4
Mean absolute error of the die prediction on held-out hot samples. Our first fit already beat the physics; the version we ship sits at the floor set by the sensor itself, whose junction reading moves 2.4 C on average from one second to the next under a perfectly steady load.

The controller is then one line. Every second it searches the fan duty from 50% upward and keeps the first one whose predicted peak stays under the target at every horizon the forecast reaches:

fan floor  =  min⁡{ f  :  max⁡H∈{10,30,60}sT^H(f)  ≤  Tlimit } \text{fan floor} \;=\; \min \left\{\, f \;:\; \max_{H \in \{10, 30, 60\}\text{s}} \hat{T}_H(f) \;\le\; T_\text{limit} \,\right\}

In words: the slowest fan speed whose predicted peak stays under the target, at every horizon the forecast reaches. The floor is written into the first two points of the card's own fan curve, so the curve's upper points and the firmware's 100 C throttle stay underneath as the safety net. It rises at once and fades gently, 1% every 10 seconds while the forecast says idle.

This is where the thermal mass gets used on purpose. A 30 s job from a cool die never reaches 90 C, so the fan stays at its floor and the job ends before the heat matters. A 60 s job does, so the floor goes straight to the level the end of the job needs, in one move, with no overshoot and no hunting. Long work settles at the physics floor.

Results

Everything below is the production service, the engine with the forecast, and a 92 C target, on one Lucebox with the R9700 and a Strix Halo, serving Qwen3.8-27B with its DFlash2 drafter. The target was chosen because it reproduces the stock firmware's peak temperatures.

Blower RPM over six minutes of agent traffic Stock firmware idles near 3,500 RPM through every pause. A temperature-only custom curve hunts between about 1,730 and 3,000 RPM every cycle. The co-design holds a steady band between about 2,200 and 2,600 RPM after the first minute. 1,600 2,100 2,600 3,100 3,600 0 60 120 180 240 300 360 Seconds Blower RPM Stock firmware Temperature-only curve Co-design
Agent trace. Stock (grey) idles at 3,500 RPM through every pause. A temperature-only custom curve (dim gold) is quieter on average but hunts 1,730 to 3,000 every cycle, which the ear hates most. The co-design (gold) holds a steady band and fades out at the end.
Mean blower RPM per workload, six-minute runs Stock firmware runs 1,906 to 3,590 RPM depending on the workload. The co-design runs 1,734 to 2,686 RPM at the same or a lower peak die temperature in every case with a stock comparison. Stock firmware Co-design Die peak 0 800 1600 2400 3200 4000 rpm Chat 1,906 76 C 1,734 73 C Agent loop 3,297 88 C 2,343 88 C Sustained 3,590 91 C 2,686 90 C Long outputs 2,595 90 C RAG 1,852 86 C Long context 2,584 93 C
Mean blower speed per workload, six-minute runs of the same request trace. Labels give the peak die temperature. The stock runs were recorded on an earlier, slower engine build; see the table for what each run actually served.
Six-minute runs on the same box, same model, same requests; stock firmware curve versus the co-design at a 92 C target
WorkloadStock blowerCo-design blower, mean / peakDie peakServed (co-design / stock)
Chat, 5 s bursts / 20 s pauses1,906 RPM, pulsing1,734 / 1,741, never leaves the floor73 C (stock 76)28 req at 65.6 tok/s / 26 at 47.1
Agent loop, 30 s / 8 s3,297 RPM2,343 / 2,64788 C (stock 88)94 req at 65.1 tok/s / 68 at 46.6
Long outputs, 1,500-token answers—2,595 / 3,04490 C11 req at 50.4 tok/s
RAG, 6K-token prompts, unique—1,852 / 2,07886 C20 req at 18.1 tok/s
Long context, 24K-token prompts, unique—2,584 / 3,03893 C8 req at 3.1 tok/s
Sustained batch3,590 RPM2,686 / 3,03190 C (stock 91)117 req at 64.7 tok/s / 82 at 45.2
Sustained, 4 concurrent streams—2,697 / 2,93592 C320 req, 177 tok/s aggregate

The die runs a couple of degrees warmer on average than under stock, 75 C against 74 on the agent trace, 84 against 81 sustained, because the controller lets the card use its thermal mass instead of blowing it down early. The peaks, which are what protect the silicon, are the same. Memory stayed at 76 to 80 C throughout.

The same box, twice, minutes apart

The numbers above come from runs recorded on separate days. To see the two policies under conditions that cannot drift, we ran one continuous session with the load never stopping and switched the fan policy every 400 seconds: warm up, firmware, forecast, firmware, forecast, firmware. The box was hot throughout, which is the state a box doing agent work all day is actually in.

The same hot box, minutes apart: the firmware curve against the forecast In one continuous session with the load never stopping, the stock firmware curve swings the blower between about 2,200 and 3,800 RPM on every cycle of the workload, averaging 3,031. The engine forecast holds it between about 2,400 and 3,100, averaging 2,771, with the die peaking 2 C higher. 2,000 2,500 3,000 3,500 4,000 0 60 120 180 240 300 Seconds inside the segment Blower RPM Stock firmware curve, mean 3,031 Engine forecast, mean 2,771
One session, five segments, the load never stopping. The firmware curve surges on every cycle of the workload; the forecast holds a band a third as wide. Repetitions agreed within 1%: 3,031, 3,004 and 3,045 RPM for the firmware, 2,771 and 2,779 for the forecast.

On this box, on this day, that is 8% less fan speed and a swing cut from 1,621 RPM to 696, at a die peak of 85 C against 83. The steadiness is the part that reproduces every time, and it is the part you hear: a fan that surges twice a minute draws attention in a way a steady one never does. The 29% at the top of this post came from the widest pair we recorded, on a day when the same workload pushed the die to 88 C instead of 83 and the firmware curve had to work much harder. How much you save depends on how hot your box already is; how much steadier it gets does not.

The target is the knob, and it is honest now. With the learned model the controller lands within 2 C of what it promises, so the target becomes a real trade: about 100 RPM per degree on agent loads and 113 on sustained work. 89 C runs the die 3 C cooler than stock's peaks for 300 RPM more; 92 C reproduces stock's peaks. We ship 92.

What batching does to heat

Concurrency changes the efficiency side without touching the noise side. With four streams the card still sits at 300 W and the blower at the same speed, but it delivers 2.7 times the tokens.

Sustained load under the co-design, one versus four concurrent streams, six minutes each
StreamsRequestsAggregate tok/sEnergy per tokenBlower, mean / peakDie peak
1117654.5 J2,686 / 3,031 RPM90 C
43201771.6 J2,697 / 2,935 RPM92 C

For a given amount of work the box therefore spends 63% less time at full blow, and each token costs a third of the energy. Under a continuous multi-user load the steady noise is unchanged, and the forecast handles the concurrent case with the exact per-slot progress the scheduler publishes.

Long context, the honest trade

A 24K-token prompt is 30 s of prefill at 300 W from a cool die, about one thermal time constant. That is the one workload where the forecast and the physics collide: the engine now predicts the 30 s correctly, and the model knows prefill runs hotter than decode, so holding a hard 92 C through it costs 3,300 to 3,600 RPM during the prefill. The previous model ran 400 RPM quieter by silently letting the peak reach 93 or 94 C. Which one a box should do is a config line, and for most clients the second is the right answer.

It has to fail safe

A fan controller that can be wrong about the future needs a floor it cannot fall through. The curve's upper points and the firmware's 100 C throttle are always there; on top, the service forces 100% fan if the die sits at 97 C, resets the curve on every exit, and falls back to the hand-fitted physics model with an extra 5 C of margin if its learned model is missing. We injected six faults under a sustained 300 W load:

Fault injection, production service, sustained load; the die never exceeded 90 C
FaultBehaviour
Hint file corrupted for 20 sread as no hint, floor faded one step, die 88 C
Hint stale for 40 sreactive fallback, floor re-planned the moment the hint returned
Engine killed, restarted after 25 sservice kept running, faded the floor, re-planned seconds after the load resumed
Service killed under loadcurve reset instantly, firmware took over at 3,549 RPM, restart re-applied a 61% floor within 20 s
Started with no model filebuilt-in physics model with margin
Reset command afterwardsidempotent, curve at firmware default

The same service was then run on the production STHT1 board with the DeepSeek V4 Flash hybrid stack on the R9700 and the Strix Halo: fan interface present, zero GPU errors, the curve reset on exit.

Hardware and setup

BoxLucebox: AMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ 395 (Strix Halo), Ubuntu, kernel 7.0 amdgpu, ROCm 7.2
Fan interfaceamdgpu.ppfeaturemask with the overdrive bit set exposes gpu_od/fan_ctrl: a 5-point hotspot curve (50 to 100%), acoustic targets, minimum duty
Model under loadQwen3.8-27B UD-IQ4_XS with the DFlash2 drafter, 300 W pinned in both prefill and decode
Engine hinta small JSON object the server publishes once per second, read over the private API or from a file on the box
Thermal modelridge regression over three horizons, fitted on 64 recorded runs (22,738 one-second rows); refit nightly on the box from its own logs
Chassisa 120 mm fan on the card's shroud at 100%: worth 4 C on memory and nothing on the die; the blower remains the only thing that moves air through the fin stack

Bottom line

A fan does not need to be smarter; it needs to be told. The inference engine already knows the next minute of its own heat, and a small learned model turns that into the slowest fan speed the physics allows. On the R9700 that is silence for chat, a third less fan for agent work and a quarter less for batch work, at the same peak temperatures, while the quiet runs served more requests than the stock baselines did. Below that, only a bigger fan or a passive card goes.


All numbers measured by us on one Lucebox (R9700 + Strix Halo) on September 15 to 17, 2026, at 1 Hz from the card's own sensors, six minutes per run, the same request trace on both sides of every pair. The stock baselines were recorded on an earlier engine build than the co-design runs, so the two sides of a pair served different amounts of work; the table says how much. Fan noise is quoted in RPM; the fan-law conversion to dB is the standard n⁵ rule, not a microphone measurement.

A quiet box that runs Qwen3.8-27B at 200+ tok/s

Open-source engine. The fan follows the engine, not the thermometer.

GitHub R9700 post Discord