You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
ClaudeCodeLoggedOut fires with correct env, empty native stderr, and sub-1.5s failure — only inside a real Hermes gateway process, never in any standalone reproduction #50
We hit persistent ClaudeCodeLoggedOut failures (the error added in #23) when running this
provider as primary/fallback inside a real, long-running Hermes gateway process (launchd-supervised
on macOS, messaging surface, cron ticking multiple profiles). Every standalone reproduction we
built — bare CLI, exact production CLI flags, this plugin's own Client class directly, concurrent
calls, calls under artificial heavy system load, calls with a full-size system prompt + 34-tool
manifest, and the async _acreate/asyncio.to_thread dispatch path with a saturated shared
executor — succeeded cleanly every time. The failure only ever reproduces inside the actual gateway
process, and we've been unable to find what's different about it after an extensive investigation.
This may be the same underlying issue as #16 or #17 (gateway/bot-surface class of bugs), but neither
report's enumerated causes match what we found, so filing separately with the new evidence in case
it helps narrow things down.
Environment
Hermes Agent: updated c4a5deef → 59004a62 same session (~11.4k commits, well past requires_hermes: ">=0.21.4")
Claude Code CLI: current as of 2026-09-26 (npm install -g @anthropic-ai/claude-code)
OS: macOS (Apple Silicon), gateway run as a launchd LaunchAgent
(RLIMIT_NOFILE soft-capped at 4096 via the plist, LimitLoadToSessionType: Aqua, Background)
Auth: isolated CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR pointed at a separate, dedicated claude auth login session directory (Claude Team plan seat), kept apart from any other login on
the machine. Baked directly into the launchd plist's EnvironmentVariables so it can't be dropped
by any config-reload path.
The org's own ANTHROPIC_API_KEY-presence conflict (this plugin's own guard) was hit and fixed
separately by scrubbing that key from this one provider's env before the check — unrelated to
this report, mentioning only so it's not mistaken for the same issue.
Symptom
error_type=ClaudeCodeLoggedOut ... summary=Claude Code is installed but has no usable login in the
environment Hermes runs it in. Run `claude auth login` as the user Hermes runs as, set
CLAUDE_CODE_OAUTH_TOKEN (from `claude setup-token`) in Hermes' environment, or point
CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR at a logged-in config directory, then try again.
(native: Not logged in · Please run /login)
Fires on the first attempt of essentially every real turn through the live gateway, retries 3x
(all fail identically), then correctly falls back to the configured fallback provider. Service
continuity is preserved by the fallback, but the primary model is effectively unusable.
What we ruled out, with real tests, not just plausible reasoning
ANTHROPIC_API_KEY conflict (this plugin's own conflict-check, directsdk.py ~line 470) —
real, and unrelated to this report; we scrub the 4 conflicting ANTHROPIC_* vars from this one
provider's env dict before the check, in our own patch. Not the cause of the symptom below.
Env-var delivery — added temporary instrumentation directly in Client._run logging HOME/CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR at the exact moment of the call. Confirmed
correct on 3/3 live failures. Later expanded to a full os.environ diff (190 keys in the real
gateway vs. 17–50 in every standalone test, including a maximally clean env -i shell) —
nothing suspicious found; all plugin-relevant vars correct every time.
Concurrent multi-process access to the shared config dir — ran up to 8 concurrent real
calls through the actual Client class (real admission relay, real per-request loopback ports,
same shared config dir) — all succeeded. Also tried 2 concurrent calls sharing oneClient
instance (same cached working directory via _workdir()) — also succeeded.
The admission relay mechanism itself — single and concurrent calls through the real relay
succeed; ruled out as a variable on its own.
System/CPU load — the gateway machine ran loadavg_1m 9–13 during real failures. Ran the
same standalone reproduction under artificial load up to loadavg_1m 10.47 (matching/exceeding
the real conditions) — still succeeded, 8/8.
Request size/complexity — every earlier test used a trivial single short message with no
tools. Rebuilt a full-size request (~15KB synthetic system prompt, 34 tool schemas matching the
real gateway's tool count, 5-turn conversation history) — single and concurrent, under load —
still succeeded.
The async dispatch path — create() branches on asyncio.get_running_loop(); every test
above had used the synchronous branch. Rebuilt the test using the real _acreate/ asyncio.to_thread path, deliberately saturating the default shared thread-pool executor with
40 unrelated blocking tasks first — still succeeded, 4/4.
What we found instead, live, in the real gateway (never reproducible standalone)
Native's stderr is completely empty on every failure (0 lines captured, via temporary stderr=subprocess.PIPE + a drain thread added directly in Client._run). There is no hidden
diagnostic text anywhere in the process we spawn.
Every failure completes in well under 1.5 seconds (0.48–0.70s observed across 3 real
failures), consistent with admission.used == False — i.e. native fails via what looks like a
purely local check before attempting any network or relay contact at all. This is far faster
than any real network round-trip we saw in a successful call (typically 1–7s).
Two calls in the same real burst (main response + an auxiliary call) started only ~10ms
apart — tighter than anything our own thread-pool-based concurrency tests produced, though
a shared-Client-instance variant with similar timing still didn't reproduce it standalone.
Existing issues checked first
Searched all 49 issues/PRs (open and closed) before filing. Closest related, neither an exact
match:
Per-client native cwd is deleted by Hermes' 24h scratch prune; long-lived clients then fail with ENOENT #43 (per-client native cwd deleted by Hermes' 24h scratch prune, long-lived clients fail
with ENOENT) — same class of bug (a long-running gateway misbehaving in a way no fresh
standalone script reproduces, surfacing as a misleading auth-shaped error). Doesn't fit here:
it requires a client older than ~24h with an idle, pruned scratch dir; every one of our
failures happened within minutes of a fresh gateway restart, and our credentials live in a
separate, non-pruned directory (CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR), not the ephemeral
per-client cwd.
Question for maintainers
Given native fails near-instantly with zero stderr output and admission.used == False, is there
a known local-only precondition Claude Code checks before ever reaching for the network (e.g.
credentials file staleness/format, a lock/session marker in the config dir, a timestamp-based
"last verified" cache) that could plausibly behave differently on a first real invocation inside a
freshly-spawned, launchd-supervised process versus a manually-invoked shell — something that isn't
env, isn't concurrency, isn't load, and isn't request shape? We're out of angles we can test without
attaching a debugger/strace to native's own process during a live failure, which is why we're
filing this now rather than continuing to guess.
Happy to run further tests against a specific hypothesis if one exists — we have the isolated repro
environment already set up and can reproduce the live-gateway conditions on request.
Summary
We hit persistent
ClaudeCodeLoggedOutfailures (the error added in #23) when running thisprovider as primary/fallback inside a real, long-running Hermes gateway process (launchd-supervised
on macOS, messaging surface, cron ticking multiple profiles). Every standalone reproduction we
built — bare CLI, exact production CLI flags, this plugin's own
Clientclass directly, concurrentcalls, calls under artificial heavy system load, calls with a full-size system prompt + 34-tool
manifest, and the async
_acreate/asyncio.to_threaddispatch path with a saturated sharedexecutor — succeeded cleanly every time. The failure only ever reproduces inside the actual gateway
process, and we've been unable to find what's different about it after an extensive investigation.
This may be the same underlying issue as #16 or #17 (gateway/bot-surface class of bugs), but neither
report's enumerated causes match what we found, so filing separately with the new evidence in case
it helps narrow things down.
Environment
c4a5deef→59004a62same session (~11.4k commits, well pastrequires_hermes: ">=0.21.4")claude-subscription-directsdk-experimental, installed viahermes plugins install claude-subscription-directsdk(pinned catalog SHA)npm install -g @anthropic-ai/claude-code)launchdLaunchAgent(
RLIMIT_NOFILEsoft-capped at 4096 via the plist,LimitLoadToSessionType: Aqua, Background)CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIRpointed at a separate, dedicatedclaude auth loginsession directory (Claude Team plan seat), kept apart from any other login onthe machine. Baked directly into the launchd plist's
EnvironmentVariablesso it can't be droppedby any config-reload path.
ANTHROPIC_API_KEY-presence conflict (this plugin's own guard) was hit and fixedseparately by scrubbing that key from this one provider's env before the check — unrelated to
this report, mentioning only so it's not mistaken for the same issue.
Symptom
Fires on the first attempt of essentially every real turn through the live gateway, retries 3x
(all fail identically), then correctly falls back to the configured fallback provider. Service
continuity is preserved by the fallback, but the primary model is effectively unusable.
What we ruled out, with real tests, not just plausible reasoning
ANTHROPIC_API_KEYconflict (this plugin's own conflict-check,directsdk.py~line 470) —real, and unrelated to this report; we scrub the 4 conflicting
ANTHROPIC_*vars from this oneprovider's env dict before the check, in our own patch. Not the cause of the symptom below.
Client._runloggingHOME/CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIRat the exact moment of the call. Confirmedcorrect on 3/3 live failures. Later expanded to a full
os.environdiff (190 keys in the realgateway vs. 17–50 in every standalone test, including a maximally clean
env -ishell) —nothing suspicious found; all plugin-relevant vars correct every time.
calls through the actual
Clientclass (real admission relay, real per-request loopback ports,same shared config dir) — all succeeded. Also tried 2 concurrent calls sharing one
Clientinstance (same cached working directory via
_workdir()) — also succeeded.succeed; ruled out as a variable on its own.
loadavg_1m9–13 during real failures. Ran thesame standalone reproduction under artificial load up to
loadavg_1m 10.47(matching/exceedingthe real conditions) — still succeeded, 8/8.
tools. Rebuilt a full-size request (~15KB synthetic system prompt, 34 tool schemas matching the
real gateway's tool count, 5-turn conversation history) — single and concurrent, under load —
still succeeded.
create()branches onasyncio.get_running_loop(); every testabove had used the synchronous branch. Rebuilt the test using the real
_acreate/asyncio.to_threadpath, deliberately saturating the default shared thread-pool executor with40 unrelated blocking tasks first — still succeeded, 4/4.
What we found instead, live, in the real gateway (never reproducible standalone)
stderr=subprocess.PIPE+ a drain thread added directly inClient._run). There is no hiddendiagnostic text anywhere in the process we spawn.
failures), consistent with
admission.used == False— i.e. native fails via what looks like apurely local check before attempting any network or relay contact at all. This is far faster
than any real network round-trip we saw in a successful call (typically 1–7s).
apart — tighter than anything our own thread-pool-based concurrency tests produced, though
a shared-Client-instance variant with similar timing still didn't reproduce it standalone.
Existing issues checked first
Searched all 49 issues/PRs (open and closed) before filing. Closest related, neither an exact
match:
/api/hellopre-flight #17/fix: explain a logged-out CLI instead of a bare Native API error #23 (theClaudeCodeLoggedOuterror class itself) — same general "works on desktop/CLI, fails ongateway/bot surface" shape, but Provider fails with “Not logged in” because the admission relay does not handle the CLI’s
/api/hellopre-flight #17's own root-cause enumeration (different user/HOME/configdir) doesn't hold here — we confirmed all three correct, live, repeatedly.
with ENOENT) — same class of bug (a long-running gateway misbehaving in a way no fresh
standalone script reproduces, surfacing as a misleading auth-shaped error). Doesn't fit here:
it requires a client older than ~24h with an idle, pruned scratch dir; every one of our
failures happened within minutes of a fresh gateway restart, and our credentials live in a
separate, non-pruned directory (
CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR), not the ephemeralper-client cwd.
Question for maintainers
Given native fails near-instantly with zero stderr output and
admission.used == False, is therea known local-only precondition Claude Code checks before ever reaching for the network (e.g.
credentials file staleness/format, a lock/session marker in the config dir, a timestamp-based
"last verified" cache) that could plausibly behave differently on a first real invocation inside a
freshly-spawned, launchd-supervised process versus a manually-invoked shell — something that isn't
env, isn't concurrency, isn't load, and isn't request shape? We're out of angles we can test without
attaching a debugger/strace to native's own process during a live failure, which is why we're
filing this now rather than continuing to guess.
Happy to run further tests against a specific hypothesis if one exists — we have the isolated repro
environment already set up and can reproduce the live-gateway conditions on request.