Qwen3.8 27B
Q4_K_M · 16.8 GB · GGUF
Tested on llama.cpp, with a small LM Studio check.
Model card & downloadsA one-row weight edit to help reasoning models deliver an answer before their token budget runs out.
200 MATH-500 problems · 4,096 output tokens
One run per configuration
The --reasoning-budget flag also eliminated empty answers and scored about the same at this budget. These are selected, reused test sets. Read the limitations.
September 29 update: earlier medium-effort results added; untested use cases clarified.
A reasoning model writes its thinking first, then its answer. Both share one output budget. If thinking uses it all, the reply ends before the answer begins.
Illustration, not measured token lengths. Delivering an answer does not guarantee it is correct.
In my 200-problem MATH test at 4,096 tokens, every empty answer from stock (31 on 27B, 29 on Flash-Next) was the model still thinking when the budget ran out. Measured
I trained a change to just one output-layer row: the one that scores </think>.
The aim is to help the model end its thinking and deliver an answer before the budget runs out.
The edit is stored in the GGUF file, with no other bytes changed and no extra runtime control required.
Fitting details are below.
Of 200 problems at a 4,096-token output budget. The reasoning-budget flag also removed all empty answers.
Paired over the same problems, the edit solved 21 that stock missed and missed 2 that stock solved (exact McNemar p = 0.0001). Against the 2,048 flag it was 9 vs 7 (p = 0.80); against 3,072, 9 vs 6 (p = 0.61).
Of 200 problems at a 4,096-token output budget. The reasoning-budget flag also removed all empty answers.
Against stock, the edit solved 25 that stock missed and missed 7 that stock solved (p = 0.002). The 2,048 flag solved 8 problems the edit missed and the edit solved 4 the flag missed (p = 0.39).
On HumanEval (164 problems, 4,096 tokens) the edit passed 157 vs 151 for stock on 27B (7 vs 1, p = 0.07) and 155 vs 151 on Flash-Next (6 vs 2, p = 0.29). Neither difference is statistically significant. Every p-value on this page is exploratory; see the limits below.
llama.cpp's --reasoning-budget flag ends the thinking after a set number of tokens. In my tests it did about as well as the
edit at 4,096 tokens, and better at a larger budget.
At 16,384 output tokens, both returned answers on all 100 problems. The accuracy difference was not statistically significant.
The flag solved 4 problems the edit missed, and the edit solved none that the flag missed (p = 0.13: not significant, but every difference went the same way). Using both together did not help: the edit plus the 2,048 flag scored 163 of 200 at 4,096 tokens, against 169 for the edit alone (5 vs 11, p = 0.21).
That the benefit carries over to such apps is an expectation. I checked it only in LM Studio, on 27B, on 12 problems. Expectation
Added September 29, 2026. I tested medium effort during development but omitted this comparison from the initial write-up. The released weights and final-release results are unchanged.
Medium effort changes the chat-template instructions. It is different from the --reasoning-budget token cutoff compared above.
These are earlier Qwen3.8 27B development runs: stock Q4_K_M and an earlier edit (cb2), one H100, 16 parallel requests, and a 4,096-token output limit. The 200 math prompts match the release tests, but the main release math runs used an RTX 3090 with 8 parallel requests and the released edit is round 6. This is a separate comparison.
| Configuration | Correct | Empty answers | Stray tags |
|---|---|---|---|
| Stock, default effort | 147 | 32 | 0 |
| Stock, medium effort | 156 | 14 | 0 |
| Earlier edit, cb2 | 167 | 1 | 4 |
| Configuration | Passed | Empty answers |
|---|---|---|
| Stock, default effort | 147 | 11 |
| Stock, medium effort | 156 | 0 |
| Earlier edit, cb2 | 154 | 0 |
Medium helped, but did not eliminate math blanks in this run. All 14 medium-effort math blanks hit the output limit. Medium eliminated coding blanks and slightly outscored the earlier edit on HumanEval. This does not establish that the edit is universally better than medium effort.
I recovered the launcher and raw records, checked shared prompt IDs and text, and recomputed the original math scoring. Coding counts are from saved execution outcomes, not a fresh execution. Settings, provenance limits and per-item evidence. No medium-effort comparison is reported for Flash-Next.
</think> into its answer:
27B in 2 of 200 MATH answers at 4,096 tokens and
4 of 100 at 16,384; Flash-Next in 1
of 200 MATH and 1 of 164 HumanEval answers. None were
seen in stock runs without the flag. The 2,048 flag also produced them at 16,384 tokens
(5 of 100).| App | 27B (Q4_K_M) | Flash-Next (UD-IQ4_XS) |
|---|---|---|
| llama.cpp | Main release evals (PrismML fork, commit 5d80cff). Mainline load check (commit 5262471): 30 of 30 MATH answered, 24 correct. | All evals (a commit on the branch of llama.cpp PR #27742). Mainline load check (commit 5262471): 30 of 30 answered, 26 correct. |
| LM Studio | Small check on 12 problems (details below). | Untested. |
| Ollama | Untested. | Untested. |
| Other apps | Untested. | Untested. |
Q4_K_M · 16.8 GB · GGUF
Tested on llama.cpp, with a small LM Studio check.
Model card & downloadsUD-IQ4_XS · About 94 GB · Three GGUF shards
Tested on llama.cpp. Other apps remain untested.
Model card & downloadsRead each model card for settings, compatibility, licensing and evidence before downloading.
At each step, the output layer gives every possible next token a score. The score for </think>
comes from one row of weights: 5,120 numbers in 27B and 2,560 in Flash-Next.
I trained a change to that row, and only that row. It was fitted on the model's own final hidden states from
469 prompts that are not in any test, with the aim of raising the </think> score when the model is ready to
answer or is going in circles, while keeping it low mid-thought and inside the answer. Fitting ran in rounds; each round
added the places where the previous version stopped too late, wrote a stray </think> into its
answer, or ended its turn with no reply.
The fitted row was stored back in the file's own format. A full byte comparison with the stock file shows nothing else changed: 4,161 of the row's 4,200 bytes differ in 27B, and 2,084 of 2,100 in Flash-Next (shard 2 only). The result loads like any GGUF and needs no flag, plugin or extra code at run time.
LM Studio 0.4.16 on Windows, engine llama.cpp-win-x86_64-nvidia-cuda12-avx2 2.47.0, one RTX 5080
(16 GB) with 60% of the model on the GPU (--gpu 0.6), context 20,480, 4 parallel requests, through its
OpenAI-compatible server. Each request sent the MATH prompt with max_tokens 4,096 and seed: 0,
and no sampler settings, so LM Studio used its own defaults; I did not record which values those were.
The problems were the first 12 of the 31 where stock 27B gave an empty answer on llama.cpp. Because they were chosen where stock failed, this checks that the fix carries over to LM Studio. This selected subset does not estimate general accuracy.
| 27B Q4_K_M, 12 problems, LM Studio | answered | correct | ran out while thinking |
|---|---|---|---|
| Row edit | 12 | 4 | 0 |
| Stock (same file, original row) | 1 | 1 | 11 |
2 of the edit's answers reached the 4,096-token limit after the answer had started. No stray tags in either run. On llama.cpp the edit also answered all 12 of these problems and got 4 right. Per-problem results.
-ngl 99 --flash-attn on --jinja, several requests in parallel. Each request sent
the messages, max_tokens and seed: 0, nothing else.xhigh). Sampler: the file's defaults
(temperature 1.0, top-k 20, top-p 0.95) plus llama.cpp's min-p 0.05.\boxed{} answer. The scorer is strict, so compare rows with each
other, not with published MATH-500 numbers.| Build | passed | empty |
|---|---|---|
| 27B stock (Q4_K_M) | 151 | 8 |
| 27B stock + flag 3072 | 157 | 0 |
| 27B row edit | 157 | 0 |
| Flash-Next stock | 151 | 11 |
| Flash-Next row edit | 155 | 0 |
| stock | edit | |
|---|---|---|
| StrongREJECT (150), judged a refusal | 145 | 147 |
| SimpleSafetyTests (100), judged a refusal | 82 | 83 |
| XSTest safe prompts (100), full compliance | 95 | 98 |
| XSTest safe prompts, partial refusal | 1 | 2 |
| XSTest safe prompts, full refusal | 1 | 0 |
| XSTest safe prompts, empty answer | 3 | 0 |
| 250 harmful prompts: empty answer, model ended its turn | 1 | 2 |
| 250 harmful prompts: empty answer, ran out of tokens | 1 | 0 |
| stock | edit | |
|---|---|---|
| StrongREJECT (150), judged a refusal | 143 | 143 |
| SimpleSafetyTests (100), judged a refusal | 85 | 84 |
| XSTest safe prompts (100), full compliance | 88 | 92 |
| XSTest safe prompts, partial refusal | 7 | 7 |
| XSTest safe prompts, full refusal | 1 | 1 |
| XSTest safe prompts, empty answer | 4 | 0 |
| 250 harmful prompts: empty answer, model ended its turn | 1 | 2 |
| 250 harmful prompts: empty answer, ran out of tokens | 1 | 0 |
Aggregate rates. Per-prompt labels are in each evidence bundle. The 250 harmful prompts are the StrongREJECT and SimpleSafetyTests rows together.
On 120 hard practice problems (MATH Level 5 training problems, not in any test), stock 27B left 47 empty at 4,096
tokens. In 4 of those 47, the thinking already contained the correct answer in a \boxed{}; in 27, the
correct answer string appeared somewhere in the thinking. String presence alone does not show that the model had
completed a correct solution. Measured
My reading is that some runs were ready to answer but did not stop, and others were genuinely unfinished; the edit can only help with the first kind. Interpretation
469 fitting prompts: 150 MATH training problems, 80 MBPP training problems, 40 Alpaca instructions and 199 harmful requests (AdvBench, JailbreakBench). Any prompt that matched or closely resembled (word overlap of 50% or more) a prompt from any test here was removed before fitting.
For 27B, the backslash in \boxed{} was lost from the 150 MATH fitting prompts because of a string
escaping bug, so the model saw a control character there. Test prompts were not affected. Flash-Next used the same
prompts with the defect repaired.
The prompt texts are not published because many are harmful requests. The evidence bundles list each prompt's source and SHA-256.
Before fitting, I check that a standalone computation of the </think> score from saved hidden
states reproduces the server. On Flash-Next the median difference was 0.0001 nats, but 8 of 36 check points were off by
0.1 to 1.1 nats, apparently because this architecture's hidden states depend on how requests are batched. The fit may
be a little noisier than on 27B. Test results are measured directly on the server and are not affected.
Qwen3.8-27B-Reasoning-Termination-Fix-Q4_K_M.gguf, SHA-256
2da7bb4517e5bd19e61ff2d8b894832d148b720afb1f06b98e5ab6f19f6c844b. Stock file: my own llama-quantize
Q4_K_M (no imatrix) from Unsloth's BF16 GGUF.
Evidence · summary counts ·
hashes.6d17a23f9334848f10829fc36cc5e5a169e0af7d34a62c0a2d8577e526bd0439.
Evidence · summary counts ·
hashes.Most fixes for runaway thinking act at run time. Others edit weights inside the network.
--reasoning-budget.</think> score and what follows it:
ThinkBrake,
Entropy After </Think>.</think> score, and report that simply scaling that score at run time made decoding unstable.I have not found published work that fits the stop decision into the </think> output row
itself and ships it as a plain GGUF, but my search was not systematic. An earlier version of this row edit is part
of my Bonsai 2 V2 release.