Skip to content

[Regression] HTTP 400 "invalid request error" is back on chat/completions (deepseek/deepseek-v4.1-flash) - same payload as #952, not a sizing problem #960

Description

@UNscientific-9

Summary

The invalid request error trace_id: … failure reported in #952 (2026-09-29, then HTTP 422) reappeared on 2026-09-30 21:54–22:01 CST, now as HTTP 400. It hit 4 times in 8 minutes and made an in-progress session unusable (every new turn died on its first request).

I am filing this as a regression / still-unfixed follow-up to #952, and I have included a set of controlled probes showing what it is not — so the on-call engineer does not have to re-run those.

Expected Behavior

Identical requests (same session, same model, same protocol, near-identical size) should behave consistently. When the provider rejects a request, the error body should name the offending parameter and the limit it violated — not a bare invalid request error with no bound and no actual value.

Actual Behavior

Client log, verbatim (extracted from the local session journal):

400: {"message":"{\"message\":\"invalid request error trace_id: a050673b3298a747a15895803095bcf0\",\"type\":\"invalid_request_error\"}\n","type":"invalid_request_error"}

All 4 failures (deduplicated by trace_id):

local time (UTC+8) trace_id HTTP
2026-09-30 21:54:31 a050673b3298a747a15895803095bcf0 400
2026-09-30 21:58:55 cba4769eacbb4c2b923514cf008c0fe5 400
2026-09-30 21:59:46 d1d04640835c25fcb9a82aad1d4f78c3 400
2026-09-30 22:01:09 60ff6649b72991ba5672087345a7aaaf 400

Model deepseek/deepseek-v4.1-flash · endpoint https://api.commandcode.ai/provider/v1 · protocol chat/completions (not responses mode) · one main session plus 5 concurrent subagent sessions on the same account.

Same payload as #952 — only the status code changed:

422: {"message":"{\"message\":\"invalid request error trace_id: f9965636e3da815215348ef4ddf2ae63\",\"type\":\"invalid_request_error\"}\n","type":"server_error"}

The inner payload is byte-identical. Only HTTP status (422 → 400) and outer type (server_error → invalid_request_error) differ.

Steps to reproduce the issue

I could not reproduce it on demand. It appears intermittent / backend-dependent — the same session at the same size alternates between success and failure:

21:54:06  OK    prompt = 168,891 tokens
21:54:31  FAIL  <- one assistant message appended in between
21:58:55  FAIL
21:59:46  FAIL
22:00:45  OK    prompt = 178,436 tokens   <- larger than the request that failed
22:01:09  FAIL  <- one short notification appended in between

Same day, same account, same model, same protocol, from the client's own usage records:

metric (2026-09-30, chat/completions) value
successful requests 1,530
successes > 150k tokens 623
successes > 170k tokens 488
successes > 200k tokens 271
largest single success 539,688 tokens

170k+ tokens both succeeds and fails; 539k succeeds. Five sibling sessions running concurrently (136k–193k tokens, same account/model/protocol) had zero failures in the same window.

Command Code Version

Not the cmd CLI. Reproduced through the provider API by a third-party client (dsh desktop app) using an API key against POST /provider/v1/chat/completions. cmd --version is therefore not applicable. Happy to attempt a reproduction through cmd if you want one — just say which version to install.

Operating System

Windows

Additional context

1. What I already ruled out with live probes (29 requests, 0 failures)

Run against POST /provider/v1/chat/completions with the same model and the same API key, ~1 hour after the incident. Every arm returned HTTP 200:

variable tested probe result
max_tokens: 384000 (our client's declared value) small input + 384000 200 OK
large input 187,625-token text + max_tokens: 384000 200 OK (usage confirms prompt_tokens: 187625)
images 1 image, then 8 images, + 384000 200 OK (0.12 MB / 0.98 MB bodies)
streaming stream: true + 384000 200 OK
concurrency 10 simultaneous requests + 384000 10/10 200
combination 187k tokens + 384000 200 OK

So this is not a prompt-size limit, not an output-budget overrun, not an image/vision issue, and not simple concurrency pressure. It looks like an intermittent server-side condition.

2. An observation that may help you locate it

Our probes did manage to trigger a different 400 — the parameter-validation path — and its shape is not the shape we hit in production:

probe (max_tokens=500000):
{"error":{"message":"{\"code\":400, \"reason\":\"INVALID_REQUEST_BODY\", \"message\":\"max_tokens (current value: 500000) must be between 0 and 393216 \", \"metadata\":{}}","type":"invalid_request_error"}}

Note: no trace_id, and a concrete reason plus the limit.

production failure:
{"message":"{\"message\":\"invalid request error trace_id: a050673b3298a747a15895803095bcf0\",\"type\":\"invalid_request_error\"}\n","type":"invalid_request_error"}

This one carries a trace_id and gives no reason. So the production failures come from a different code path than the request-body validator — most likely something upstream being passed through. That is why I am asking you to look the trace IDs up rather than guessing further from the client side.

(Side note while we were in there: the hard ceiling for this model is max_tokens <= 393216. Our client declares 384000, i.e. 97.7% of the ceiling. It is legal today, but close enough to the edge that any future lowering of that ceiling would break it silently. Not the cause of this report — just flagging it.)

3. Recovery path / impact

When this fires, the turn dies immediately and the session has no way out from the client side: the message is not recognisable as retryable and does not trigger compaction, so the user has to intervene manually. Same as described in #952 and #892.

4. Related issues

#952 (identical payload, HTTP 422, still open, 0 comments after ~28 h) · #955 (intermittent HTTP 422 above ~256K, closed by its own author, not by a fix) · #892 (client max_tokens not clamped → permanent 400 loop; assigned to ahmadbilaldev) · #944 (another dsh user: "I hope the Command Code provider can ignore fields that are currently unsupported. The JSON validation is currently too strict.") · #950 / #949 (400s in responses mode) · #959 (400 a single path expansion cannot exceed 512 candidates, long session unrecoverable, opened 2026-09-30 13:48 UTC).

5. What would help most

  1. Look up the four trace_ids above and tell us what the backend saw — was the request rejected upstream, or by a specific gateway instance?
  2. If it is a parameter/field check, name the field and the limit in the error body. A bare invalid request error is impossible to act on from the client side.
  3. If it is capacity/routing (a subset of backends rejecting, others accepting), say so — that would explain the alternating success/failure at identical sizes and would let clients implement a sane retry.

6. Session file

I can attach a redacted extract with only the 4 error entries plus the surrounding usage records (no user content). I would rather not upload the full session journal: it contains third-party personal data.

Activity

  1. UNscientific-9 commented on Sep 30, 2026

    @UNscientific-9
    Author

    Update — the failure continued after filing, and turned continuous

    Two more failures landed after I opened this issue. Same session, same payload shape, nothing changed on my side:

    local time (UTC+8) trace_id HTTP
    2026-09-30 22:09:30 42271aac2750c1d402e7ca93b50e7e0e 400
    2026-09-30 22:10:32 6f2007fe77b69f500ecd0d3d4ffd1144 400

    Complete trace_id list for this session — 6 failures over 16 minutes (21:54:31 → 22:10:32):

    21:54:31  a050673b3298a747a15895803095bcf0
    21:58:55  cba4769eacbb4c2b923514cf008c0fe5
    21:59:46  d1d04640835c25fcb9a82aad1d4f78c3
    22:01:09  60ff6649b72991ba5672087345a7aaaf
    22:09:30  42271aac2750c1d402e7ca93b50e7e0e
    22:10:32  6f2007fe77b69f500ecd0d3d4ffd1144
    

    Why this changes the picture. There was exactly one success in the middle of that window (22:00:45), so it is not a hard content lock. But after 22:00:54 the session failed on every single subsequent request — 5 consecutive turns, each dying on its first step. If the rejection were random background noise (the ~0.3% rate implied by the day's 1,530 successes vs 4 failures), five consecutive hits would be vanishingly unlikely. So something is pinning this session specifically.

    Two explanations fit what I can see from the client side, and I cannot distinguish them without your backend logs:

    1. Session affinity to a bad backend. The session keeps getting routed to one instance that rejects it, while my probe traffic (29 requests, run 22:1x–22:3x, while this session was still failing) all went to healthy instances and returned 200. The single success at 22:00:45 would be the one request that happened to land elsewhere.
    2. Something in this session's accumulated request body that the validator rejects. The client appends the error text to the session history on each failed attempt, so the payload mutates on every retry — the self-reinforcing loop already described in Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1049647 tokens (985647 in the messages, 64000 in the completion). #892.

    Recovery is still impossible from the client side. The turn dies immediately, the message is not recognised as retryable, and compaction does not fire (the session sits around 179k tokens, far below this client's compaction trigger). The user's only escape is to start a new session — which is exactly the "long session unrecoverable" failure mode this project has now seen in #892 and #959.

    What I would ask for from these six trace IDs: whether they were served by the same backend instance, and whether that instance also served the 22:00:45 success. If it is affinity, a retry-on-a-different-backend policy would fix this class of failure outright.

  2. NeetYuki commented on Sep 30, 2026

    @NeetYuki

    Same issue when using DeepSeekV4.1Flash for the task. Please fix it.

  3. UNscientific-9 commented on Oct 1, 2026

    @UNscientific-9
    Author

    Update — the "unsupported parameter" hypothesis is testable, and it fails.

    Ran a parameter-acceptance probe against the same endpoint, model and key, with tiny bodies:

    • max_tokens: 384000 → 200
    • reasoning effort max (what our client is configured with) → 200
    • stream: true → 200
    • an unknown parameter (verbosity) → 200, silently ignored — this API does not reject stray parameters

    And the control that matters: when a parameter genuinely is invalid, this API names it and returns a param field:

    {"error":{"message":"Invalid option: expected one of \"low\"|\"medium\"|\"high\"|\"xhigh\"|\"max\"","type":"invalid_request_error","param":"reasoning_effort"}}
    

    The 400s reported above carry no param and no code — only message and type. Separately, the client records its request header (config + tools) once per session with reason initial and never changed it, so all 56 requests in the affected session sent identical parameters — and 47 of them returned 200. A parameter problem would have failed all 56.

    The failure therefore correlates with neither the parameter set nor the transcript content (see the OK/400 interleaving in "Steps to reproduce"). The remaining variable is which backend served the request, and the six trace_ids are the handle on that.

  4. self-assigned this
    on Oct 1, 2026
  5. ahmadbilaldev commented on Oct 1, 2026

    @ahmadbilaldev

    Taking a look

  6. UNscientific-9 commented on Oct 2, 2026

    @UNscientific-9
    Author

    Image Same issue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions