Research noteJosh Bolding (@boldingbuilds) · September 2026

Helping Qwen
Finish Thinking

A one-row weight edit to help reasoning models deliver an answer before their token budget runs out.

Empty answers · stock → edited

200 MATH-500 problems · 4,096 output tokens
One run per configuration

Qwen3.8 27B31→0
Qwen3.8 Flash-Next29→0

The --reasoning-budget flag also eliminated empty answers and scored about the same at this budget. These are selected, reused test sets. Read the limitations.

September 29 update: earlier medium-effort results added; untested use cases clarified.

Thinking needs room for an answer.

A reasoning model writes its thinking first, then its answer. Both share one output budget. If thinking uses it all, the reply ends before the answer begins.

Budget exhausted4,096 tokens
Thinking continues to the limit
No answer returned
Answer delivered4,096 tokens
Thinking
</think>
Answer
Thinking ends with room to respond

Illustration, not measured token lengths. Delivering an answer does not guarantee it is correct.

Empty answer
Nothing after the thinking. The model did not decline and did not get it wrong; it never answered. Graders score it as a miss.
Refusal
The model answers, but declines the request.
Wrong answer
The model answers, and the answer is incorrect.

In my 200-problem MATH test at 4,096 tokens, every empty answer from stock (31 on 27B, 29 on Flash-Next) was the model still thinking when the budget ran out. Measured

01 / METHODOne row. An ordinary GGUF.

I trained a change to just one output-layer row: the one that scores </think>. The aim is to help the model end its thinking and deliver an answer before the budget runs out. The edit is stored in the GGUF file, with no other bytes changed and no extra runtime control required. Fitting details are below.

02 / RESULTSWhat changed at 4,096 tokens

Qwen3.8 27BQ4_K_M GGUF. 200 MATH-500 problems, 4,096 output tokens, one run per row.
Both bar scales: 0–200 problemsOutlined dot = zero empty answers
Empty answers · stock to edited31→0
Correct answers · stock to edited150→169

Of 200 problems at a 4,096-token output budget. The reasoning-budget flag also removed all empty answers.

Paired results and statistical detail

Paired over the same problems, the edit solved 21 that stock missed and missed 2 that stock solved (exact McNemar p = 0.0001). Against the 2,048 flag it was 9 vs 7 (p = 0.80); against 3,072, 9 vs 6 (p = 0.61).

Qwen3.8 Flash-NextUnsloth UD-IQ4_XS GGUF. Same 200 problems, 4,096 output tokens, one run per row.
Both bar scales: 0–200 problemsOutlined dot = zero empty answers
Empty answers · stock to edited29→0
Correct answers · stock to edited150→168

Of 200 problems at a 4,096-token output budget. The reasoning-budget flag also removed all empty answers.

Paired results and statistical detail

Against stock, the edit solved 25 that stock missed and missed 7 that stock solved (p = 0.002). The 2,048 flag solved 8 problems the edit missed and the edit solved 4 the flag missed (p = 0.39).

On HumanEval (164 problems, 4,096 tokens) the edit passed 157 vs 151 for stock on 27B (7 vs 1, p = 0.07) and 155 vs 151 on Flash-Next (6 vs 2, p = 0.29). Neither difference is statistically significant. Every p-value on this page is exploratory; see the limits below.

03 / COMPARISONEdit, flag, or medium effort?

llama.cpp's --reasoning-budget flag ends the thinking after a set number of tokens. In my tests it did about as well as the edit at 4,096 tokens, and better at a larger budget.

Qwen3.8 27B at a larger budgetQ4_K_M GGUF. First 100 of the same problems, 16,384 output tokens. Not tested on Flash-Next.
Both bar scales: 0–100 problemsOutlined dot = zero empty answers
Correct with the 2,048-token flag88 / 100
Correct with the edited model84 / 100

At 16,384 output tokens, both returned answers on all 100 problems. The accuracy difference was not statistically significant.

Paired results and statistical detail

The flag solved 4 problems the edit missed, and the edit solved none that the flag missed (p = 0.13: not significant, but every difference went the same way). Using both together did not help: the edit plus the 2,048 flag scored 163 of 200 at 4,096 tokens, against 169 for the edit alone (5 vs 11, p = 0.21).

The edit may suit you if

  • your app or API has no reasoning-budget control;
  • you want a file that needs no settings.

That the benefit carries over to such apps is an expectation. I checked it only in LM Studio, on 27B, on 12 problems. Expectation

The flag is the better tool if

  • you run llama.cpp yourself;
  • you use larger output budgets;
  • you want to tune the budget per task;
  • you prefer unmodified weights.

Earlier development comparison: medium reasoning effort

Added September 29, 2026. I tested medium effort during development but omitted this comparison from the initial write-up. The released weights and final-release results are unchanged.

Medium effort changes the chat-template instructions. It is different from the --reasoning-budget token cutoff compared above.

These are earlier Qwen3.8 27B development runs: stock Q4_K_M and an earlier edit (cb2), one H100, 16 parallel requests, and a 4,096-token output limit. The 200 math prompts match the release tests, but the main release math runs used an RTX 3090 with 8 parallel requests and the released edit is round 6. This is a separate comparison.

Development MATH-500 comparison · 200 problems
ConfigurationCorrectEmpty answersStray tags
Stock, default effort147320
Stock, medium effort156140
Earlier edit, cb216714
Same development campaign · HumanEval, 164 problems
ConfigurationPassedEmpty answers
Stock, default effort14711
Stock, medium effort1560
Earlier edit, cb21540

Medium helped, but did not eliminate math blanks in this run. All 14 medium-effort math blanks hit the output limit. Medium eliminated coding blanks and slightly outscored the earlier edit on HumanEval. This does not establish that the edit is universally better than medium effort.

I recovered the launcher and raw records, checked shared prompt IDs and text, and recomputed the original math scoring. Coding counts are from saved execution outcomes, not a fresh execution. Settings, provenance limits and per-item evidence. No medium-effort comparison is reported for Flash-Next.

04 / LIMITATIONSWhere the evidence stops

Where each file was tested

App27B (Q4_K_M)Flash-Next (UD-IQ4_XS)
llama.cppMain release evals (PrismML fork, commit 5d80cff). Mainline load check (commit 5262471): 30 of 30 MATH answered, 24 correct. All evals (a commit on the branch of llama.cpp PR #27742). Mainline load check (commit 5262471): 30 of 30 answered, 26 correct.
LM StudioSmall check on 12 problems (details below).Untested.
OllamaUntested.Untested.
Other appsUntested.Untested.

05 / MODELSExplore the files

Qwen3.8 Flash-Next

UD-IQ4_XS · About 94 GB · Three GGUF shards

Tested on llama.cpp. Other apps remain untested.

Model card & downloads

Read each model card for settings, compatibility, licensing and evidence before downloading.

06 / DETAILSMethod & reproduction

How the row was fitted

At each step, the output layer gives every possible next token a score. The score for </think> comes from one row of weights: 5,120 numbers in 27B and 2,560 in Flash-Next.

I trained a change to that row, and only that row. It was fitted on the model's own final hidden states from 469 prompts that are not in any test, with the aim of raising the </think> score when the model is ready to answer or is going in circles, while keeping it low mid-thought and inside the answer. Fitting ran in rounds; each round added the places where the previous version stopped too late, wrote a stray </think> into its answer, or ended its turn with no reply.

The fitted row was stored back in the file's own format. A full byte comparison with the stock file shows nothing else changed: 4,161 of the row's 4,200 bytes differ in 27B, and 2,084 of 2,100 in Flash-Next (shard 2 only). The result loads like any GGUF and needs no flag, plugin or extra code at run time.

LM Studio check (27B only)

LM Studio 0.4.16 on Windows, engine llama.cpp-win-x86_64-nvidia-cuda12-avx2 2.47.0, one RTX 5080 (16 GB) with 60% of the model on the GPU (--gpu 0.6), context 20,480, 4 parallel requests, through its OpenAI-compatible server. Each request sent the MATH prompt with max_tokens 4,096 and seed: 0, and no sampler settings, so LM Studio used its own defaults; I did not record which values those were.

The problems were the first 12 of the 31 where stock 27B gave an empty answer on llama.cpp. Because they were chosen where stock failed, this checks that the fix carries over to LM Studio. This selected subset does not estimate general accuracy.

27B Q4_K_M, 12 problems, LM Studioansweredcorrectran out while thinking
Row edit1240
Stock (same file, original row)1111

2 of the edit's answers reached the 4,096-token limit after the answer had started. No stray tags in either run. On llama.cpp the edit also answered all 12 of these problems and got 4 right. Per-problem results.

How the main llama.cpp tests were run
  • llama-server with -ngl 99 --flash-attn on --jinja, several requests in parallel. Each request sent the messages, max_tokens and seed: 0, nothing else.
  • Thinking on, at the chat template's default effort (xhigh). Sampler: the file's defaults (temperature 1.0, top-k 20, top-p 0.95) plus llama.cpp's min-p 0.05.
  • MATH-500: the first 200 problems after a seeded shuffle (seed 7); the first 100 of those at 16,384 tokens. Scored by exact match of the last \boxed{} answer. The scorer is strict, so compare rows with each other, not with published MATH-500 numbers.
  • HumanEval: all 164 problems, pass@1, one sample each.
  • Refusals: StrongREJECT (first 150), SimpleSafetyTests (100), XSTest safe prompts (first 100), graded by an LLM judge that is not the model under test (a Q8_0 file from OBLITERATUS/Qwen3.8-27B-OBLITERATED, thinking off, temperature 0). An uncensored judge was used because a safety-tuned judge can refuse to read harmful replies.
  • An empty answer has two causes, counted separately in the evidence: the budget ran out during thinking, or the model ended its turn with no answer.
  • 27B tests ran on RTX 3090s. Flash-Next tests ran on rented H200s, except its mainline load check (one RTX 3090).
HumanEval and refusal tables

HumanEval, 164 problems, 4,096 tokens

Buildpassedempty
27B stock (Q4_K_M)1518
27B stock + flag 30721570
27B row edit1570
Flash-Next stock15111
Flash-Next row edit1550

Refusals, Qwen3.8 27B

stockedit
StrongREJECT (150), judged a refusal145147
SimpleSafetyTests (100), judged a refusal8283
XSTest safe prompts (100), full compliance9598
XSTest safe prompts, partial refusal12
XSTest safe prompts, full refusal10
XSTest safe prompts, empty answer30
250 harmful prompts: empty answer, model ended its turn12
250 harmful prompts: empty answer, ran out of tokens10

Refusals, Qwen3.8 Flash-Next

stockedit
StrongREJECT (150), judged a refusal143143
SimpleSafetyTests (100), judged a refusal8584
XSTest safe prompts (100), full compliance8892
XSTest safe prompts, partial refusal77
XSTest safe prompts, full refusal11
XSTest safe prompts, empty answer40
250 harmful prompts: empty answer, model ended its turn12
250 harmful prompts: empty answer, ran out of tokens10

Aggregate rates. Per-prompt labels are in each evidence bundle. The 250 harmful prompts are the StrongREJECT and SimpleSafetyTests rows together.

Why an earlier stop can help

On 120 hard practice problems (MATH Level 5 training problems, not in any test), stock 27B left 47 empty at 4,096 tokens. In 4 of those 47, the thinking already contained the correct answer in a \boxed{}; in 27, the correct answer string appeared somewhere in the thinking. String presence alone does not show that the model had completed a correct solution. Measured

My reading is that some runs were ready to answer but did not stop, and others were genuinely unfinished; the edit can only help with the first kind. Interpretation

Fitting data, and a known defect

469 fitting prompts: 150 MATH training problems, 80 MBPP training problems, 40 Alpaca instructions and 199 harmful requests (AdvBench, JailbreakBench). Any prompt that matched or closely resembled (word overlap of 50% or more) a prompt from any test here was removed before fitting.

For 27B, the backslash in \boxed{} was lost from the 150 MATH fitting prompts because of a string escaping bug, so the model saw a control character there. Test prompts were not affected. Flash-Next used the same prompts with the defect repaired.

The prompt texts are not published because many are harmful requests. The evidence bundles list each prompt's source and SHA-256.

A measurement quirk on Flash-Next

Before fitting, I check that a standalone computation of the </think> score from saved hidden states reproduces the server. On Flash-Next the median difference was 0.0001 nats, but 8 of 36 check points were off by 0.1 to 1.1 nats, apparently because this architecture's hidden states depend on how requests are batched. The fit may be a little noisier than on 27B. Test results are measured directly on the server and are not affected.

Files, hashes and reproduction
  • 27B: Qwen3.8-27B-Reasoning-Termination-Fix-Q4_K_M.gguf, SHA-256 2da7bb4517e5bd19e61ff2d8b894832d148b720afb1f06b98e5ab6f19f6c844b. Stock file: my own llama-quantize Q4_K_M (no imatrix) from Unsloth's BF16 GGUF. Evidence · summary counts · hashes.
  • Flash-Next: three shards; only shard 2 differs from Unsloth's UD-IQ4_XS. Shard 2 SHA-256 6d17a23f9334848f10829fc36cc5e5a169e0af7d34a62c0a2d8577e526bd0439. Evidence · summary counts · hashes.
  • The bundles hold per-problem results (IDs, scores and flags, no model text), scripts, settings, judge labels by prompt index and bake logs. They are enough to check every number. Rerunning the fit needs the unpublished fitting prompts, and sampling is not bit-reproducible, so a refit will not give the exact same row.

Related work

Most fixes for runaway thinking act at run time. Others edit weights inside the network.

I have not found published work that fits the stop decision into the </think> output row itself and ships it as a plain GGUF, but my search was not systematic. An earlier version of this row edit is part of my Bonsai 2 V2 release.

Credits