Repository navigation
[Regression] HTTP 400 "invalid request error" is back on chat/completions (deepseek/deepseek-v4.1-flash) - same payload as #952, not a sizing problem #960
Description
Activity
Update — the failure continued after filing, and turned continuous
Two more failures landed after I opened this issue. Same session, same payload shape, nothing changed on my side:
local time (UTC+8) trace_id HTTP 2026-09-30 22:09:30 42271aac2750c1d402e7ca93b50e7e0e400 2026-09-30 22:10:32 6f2007fe77b69f500ecd0d3d4ffd1144400 Complete trace_id list for this session — 6 failures over 16 minutes (21:54:31 → 22:10:32):
21:54:31 a050673b3298a747a15895803095bcf0 21:58:55 cba4769eacbb4c2b923514cf008c0fe5 21:59:46 d1d04640835c25fcb9a82aad1d4f78c3 22:01:09 60ff6649b72991ba5672087345a7aaaf 22:09:30 42271aac2750c1d402e7ca93b50e7e0e 22:10:32 6f2007fe77b69f500ecd0d3d4ffd1144Why this changes the picture. There was exactly one success in the middle of that window (22:00:45), so it is not a hard content lock. But after 22:00:54 the session failed on every single subsequent request — 5 consecutive turns, each dying on its first step. If the rejection were random background noise (the ~0.3% rate implied by the day's 1,530 successes vs 4 failures), five consecutive hits would be vanishingly unlikely. So something is pinning this session specifically.
Two explanations fit what I can see from the client side, and I cannot distinguish them without your backend logs:
- Session affinity to a bad backend. The session keeps getting routed to one instance that rejects it, while my probe traffic (29 requests, run 22:1x–22:3x, while this session was still failing) all went to healthy instances and returned 200. The single success at 22:00:45 would be the one request that happened to land elsewhere.
- Something in this session's accumulated request body that the validator rejects. The client appends the error text to the session history on each failed attempt, so the payload mutates on every retry — the self-reinforcing loop already described in Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1049647 tokens (985647 in the messages, 64000 in the completion). #892.
Recovery is still impossible from the client side. The turn dies immediately, the message is not recognised as retryable, and compaction does not fire (the session sits around 179k tokens, far below this client's compaction trigger). The user's only escape is to start a new session — which is exactly the "long session unrecoverable" failure mode this project has now seen in #892 and #959.
What I would ask for from these six trace IDs: whether they were served by the same backend instance, and whether that instance also served the 22:00:45 success. If it is affinity, a retry-on-a-different-backend policy would fix this class of failure outright.
Reacted by NeetYukiSame issue when using DeepSeekV4.1Flash for the task. Please fix it.
Reacted by TONYGFXUpdate — the "unsupported parameter" hypothesis is testable, and it fails.
Ran a parameter-acceptance probe against the same endpoint, model and key, with tiny bodies:
max_tokens: 384000→ 200- reasoning effort
max(what our client is configured with) → 200 stream: true→ 200- an unknown parameter (
verbosity) → 200, silently ignored — this API does not reject stray parameters
And the control that matters: when a parameter genuinely is invalid, this API names it and returns a
paramfield:{"error":{"message":"Invalid option: expected one of \"low\"|\"medium\"|\"high\"|\"xhigh\"|\"max\"","type":"invalid_request_error","param":"reasoning_effort"}}The 400s reported above carry no
paramand nocode— onlymessageandtype. Separately, the client records its request header (config + tools) once per session with reasoninitialand never changed it, so all 56 requests in the affected session sent identical parameters — and 47 of them returned 200. A parameter problem would have failed all 56.The failure therefore correlates with neither the parameter set nor the transcript content (see the OK/400 interleaving in "Steps to reproduce"). The remaining variable is which backend served the request, and the six trace_ids are the handle on that.
Taking a look

Summary
The
invalid request error trace_id: …failure reported in #952 (2026-09-29, then HTTP 422) reappeared on 2026-09-30 21:54–22:01 CST, now as HTTP 400. It hit 4 times in 8 minutes and made an in-progress session unusable (every new turn died on its first request).I am filing this as a regression / still-unfixed follow-up to #952, and I have included a set of controlled probes showing what it is not — so the on-call engineer does not have to re-run those.
Expected Behavior
Identical requests (same session, same model, same protocol, near-identical size) should behave consistently. When the provider rejects a request, the error body should name the offending parameter and the limit it violated — not a bare
invalid request errorwith no bound and no actual value.Actual Behavior
Client log, verbatim (extracted from the local session journal):
All 4 failures (deduplicated by
trace_id):a050673b3298a747a15895803095bcf0cba4769eacbb4c2b923514cf008c0fe5d1d04640835c25fcb9a82aad1d4f78c360ff6649b72991ba5672087345a7aaafModel
deepseek/deepseek-v4.1-flash· endpointhttps://api.commandcode.ai/provider/v1· protocolchat/completions(not responses mode) · one main session plus 5 concurrent subagent sessions on the same account.Same payload as #952 — only the status code changed:
The inner payload is byte-identical. Only HTTP status (422 → 400) and outer
type(server_error→invalid_request_error) differ.Steps to reproduce the issue
I could not reproduce it on demand. It appears intermittent / backend-dependent — the same session at the same size alternates between success and failure:
Same day, same account, same model, same protocol, from the client's own usage records:
170k+ tokens both succeeds and fails; 539k succeeds. Five sibling sessions running concurrently (136k–193k tokens, same account/model/protocol) had zero failures in the same window.
Command Code Version
Not the
cmdCLI. Reproduced through the provider API by a third-party client (dsh desktop app) using an API key againstPOST /provider/v1/chat/completions.cmd --versionis therefore not applicable. Happy to attempt a reproduction throughcmdif you want one — just say which version to install.Operating System
Windows
Additional context
1. What I already ruled out with live probes (29 requests, 0 failures)
Run against
POST /provider/v1/chat/completionswith the same model and the same API key, ~1 hour after the incident. Every arm returned HTTP 200:max_tokens: 384000(our client's declared value)max_tokens: 384000prompt_tokens: 187625)stream: true+ 384000So this is not a prompt-size limit, not an output-budget overrun, not an image/vision issue, and not simple concurrency pressure. It looks like an intermittent server-side condition.
2. An observation that may help you locate it
Our probes did manage to trigger a different 400 — the parameter-validation path — and its shape is not the shape we hit in production:
Note: no
trace_id, and a concretereasonplus the limit.This one carries a
trace_idand gives no reason. So the production failures come from a different code path than the request-body validator — most likely something upstream being passed through. That is why I am asking you to look the trace IDs up rather than guessing further from the client side.(Side note while we were in there: the hard ceiling for this model is
max_tokens <= 393216. Our client declares384000, i.e. 97.7% of the ceiling. It is legal today, but close enough to the edge that any future lowering of that ceiling would break it silently. Not the cause of this report — just flagging it.)3. Recovery path / impact
When this fires, the turn dies immediately and the session has no way out from the client side: the message is not recognisable as retryable and does not trigger compaction, so the user has to intervene manually. Same as described in #952 and #892.
4. Related issues
#952 (identical payload, HTTP 422, still open, 0 comments after ~28 h) · #955 (intermittent HTTP 422 above ~256K, closed by its own author, not by a fix) · #892 (client
max_tokensnot clamped → permanent 400 loop; assigned to ahmadbilaldev) · #944 (another dsh user: "I hope the Command Code provider can ignore fields that are currently unsupported. The JSON validation is currently too strict.") · #950 / #949 (400s in responses mode) · #959 (400a single path expansion cannot exceed 512 candidates, long session unrecoverable, opened 2026-09-30 13:48 UTC).5. What would help most
trace_ids above and tell us what the backend saw — was the request rejected upstream, or by a specific gateway instance?invalid request erroris impossible to act on from the client side.6. Session file
I can attach a redacted extract with only the 4 error entries plus the surrounding
usagerecords (no user content). I would rather not upload the full session journal: it contains third-party personal data.