The finding

We sent DeepSeek V4 Pro the same 97 calculation cases through OpenRouter, pinning each run to a single host: accrued interest, required collateral, recall adjustments, each with the calculation conventions stated.

Novita got 99% right. DigitalOcean got 91%. Azure got 7%. DeepInfra got 5%.

Same model name. Same prompts. Same cases. The response on every host identified itself identically, as deepseek/deepseek-v4-pro.

The cause was a single default. Azure and DeepInfra returned zero reasoning tokens; DigitalOcean and Novita spent between 1,700 and 3,500 per call. Two hosts serve the model with reasoning off unless asked, two serve it on. When we requested reasoning explicitly, Azure and DeepInfra reasoned and got five of five cases right each. Nobody is doing anything wrong. The defaults just differ, and the response does not say which one you got.

The implication

Without reasoning, the model still understood every case. Its explanations correctly identified the recall, worked out the post-recall balance, chose the right settlement date. Then it reported a term or nothing instead of computing the interest.

That is the gap this newsletter has written about for months: models read the dispute and fail to produce the number. It did not disappear at the frontier. It sits behind the reasoning budget, and a routing default can quietly take the budget away.

The fix is two lines: request reasoning explicitly on every call, and check the reasoning token count on every response.

The residual finding

We then checked our own published data. Our AAL-D-007 benchmark recorded reasoning tokens on every call, so it could be split after the fact.

About 16% of DeepSeek’s calls had run without reasoning. The effect on accuracy was small, roughly three points, because that benchmark prints every figure and does not require calculation. The effect on consistency was not. Where a case’s repeated runs mixed reasoning on and off, DeepSeek changed its answer on 37% of cases, against 21% where every run reasoned. Kimi K3 showed the same pattern, 15% against 8%.

Part of what we reported as model instability was routing. We have added a disclosure to the D-007 page with the corrected figures, and qualified our earlier statement that DeepSeek was the least reliable model across two benchmarks.

The control consideration

If your firm runs reasoning models through a gateway, aggregator or third-party host, three questions:

Is reasoning requested explicitly, or left to the host’s default?

Do you log reasoning tokens per call, so you would notice if it stopped?

Do you know which host served each answer?

In our test, the answers to those three questions moved accuracy on identical cases from 5% to 99%.

From AI Alpha Labs this week

  • Correction. Last week’s issue attributed GPT-5.6 Sol’s errors on our repo benchmark to a blind spot on recall disputes. The cause was an ambiguous field in our own test data. With it fixed, the model scored 32 of 32.

  • AAL-D-008 is published. Frontier models saturated it, both with computed figures printed and when made to compute them, and every failure we found came from ambiguity in the test, not arithmetic: aialphalabs.ai/research/AAL-D-008

  • AAL-D-007 now carries a disclosure on reasoning mode and flip rates: aialphalabs.ai/research/AAL-D-007