Skip to content

ClaudeCodeLoggedOut fires with correct env, empty native stderr, and sub-1.5s failure — only inside a real Hermes gateway process, never in any standalone reproduction #50

Description

@seterhenakbar

Summary

We hit persistent ClaudeCodeLoggedOut failures (the error added in #23) when running this
provider as primary/fallback inside a real, long-running Hermes gateway process (launchd-supervised
on macOS, messaging surface, cron ticking multiple profiles). Every standalone reproduction we
built — bare CLI, exact production CLI flags, this plugin's own Client class directly, concurrent
calls, calls under artificial heavy system load, calls with a full-size system prompt + 34-tool
manifest, and the async _acreate/asyncio.to_thread dispatch path with a saturated shared
executor — succeeded cleanly every time. The failure only ever reproduces inside the actual gateway
process, and we've been unable to find what's different about it after an extensive investigation.

This may be the same underlying issue as #16 or #17 (gateway/bot-surface class of bugs), but neither
report's enumerated causes match what we found, so filing separately with the new evidence in case
it helps narrow things down.

Environment

  • Hermes Agent: updated c4a5deef → 59004a62 same session (~11.4k commits, well past
    requires_hermes: ">=0.21.4")
  • Plugin: claude-subscription-directsdk-experimental, installed via hermes plugins install claude-subscription-directsdk (pinned catalog SHA)
  • Claude Code CLI: current as of 2026-09-26 (npm install -g @anthropic-ai/claude-code)
  • OS: macOS (Apple Silicon), gateway run as a launchd LaunchAgent
    (RLIMIT_NOFILE soft-capped at 4096 via the plist, LimitLoadToSessionType: Aqua, Background)
  • Auth: isolated CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR pointed at a separate, dedicated
    claude auth login session directory (Claude Team plan seat), kept apart from any other login on
    the machine. Baked directly into the launchd plist's EnvironmentVariables so it can't be dropped
    by any config-reload path.
  • The org's own ANTHROPIC_API_KEY-presence conflict (this plugin's own guard) was hit and fixed
    separately by scrubbing that key from this one provider's env before the check — unrelated to
    this report, mentioning only so it's not mistaken for the same issue.

Symptom

error_type=ClaudeCodeLoggedOut ... summary=Claude Code is installed but has no usable login in the
environment Hermes runs it in. Run `claude auth login` as the user Hermes runs as, set
CLAUDE_CODE_OAUTH_TOKEN (from `claude setup-token`) in Hermes' environment, or point
CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR at a logged-in config directory, then try again.
(native: Not logged in · Please run /login)

Fires on the first attempt of essentially every real turn through the live gateway, retries 3x
(all fail identically), then correctly falls back to the configured fallback provider. Service
continuity is preserved by the fallback, but the primary model is effectively unusable.

What we ruled out, with real tests, not just plausible reasoning

  1. ANTHROPIC_API_KEY conflict (this plugin's own conflict-check, directsdk.py ~line 470) —
    real, and unrelated to this report; we scrub the 4 conflicting ANTHROPIC_* vars from this one
    provider's env dict before the check, in our own patch. Not the cause of the symptom below.
  2. Env-var delivery — added temporary instrumentation directly in Client._run logging
    HOME/CLAUDE_SUBSCRIPTION_DIRECTSDK_CONFIG_DIR at the exact moment of the call. Confirmed
    correct on 3/3 live failures. Later expanded to a full os.environ diff (190 keys in the real
    gateway vs. 17–50 in every standalone test, including a maximally clean env -i shell) —
    nothing suspicious found; all plugin-relevant vars correct every time.
  3. Concurrent multi-process access to the shared config dir — ran up to 8 concurrent real
    calls through the actual Client class (real admission relay, real per-request loopback ports,
    same shared config dir) — all succeeded. Also tried 2 concurrent calls sharing one Client
    instance (same cached working directory via _workdir()) — also succeeded.
  4. The admission relay mechanism itself — single and concurrent calls through the real relay
    succeed; ruled out as a variable on its own.
  5. System/CPU load — the gateway machine ran loadavg_1m 9–13 during real failures. Ran the
    same standalone reproduction under artificial load up to loadavg_1m 10.47 (matching/exceeding
    the real conditions) — still succeeded, 8/8.
  6. Request size/complexity — every earlier test used a trivial single short message with no
    tools. Rebuilt a full-size request (~15KB synthetic system prompt, 34 tool schemas matching the
    real gateway's tool count, 5-turn conversation history) — single and concurrent, under load —
    still succeeded.
  7. The async dispatch path — create() branches on asyncio.get_running_loop(); every test
    above had used the synchronous branch. Rebuilt the test using the real _acreate/
    asyncio.to_thread path, deliberately saturating the default shared thread-pool executor with
    40 unrelated blocking tasks first — still succeeded, 4/4.

What we found instead, live, in the real gateway (never reproducible standalone)

  • Native's stderr is completely empty on every failure (0 lines captured, via temporary
    stderr=subprocess.PIPE + a drain thread added directly in Client._run). There is no hidden
    diagnostic text anywhere in the process we spawn.
  • Every failure completes in well under 1.5 seconds (0.48–0.70s observed across 3 real
    failures), consistent with admission.used == False — i.e. native fails via what looks like a
    purely local check before attempting any network or relay contact at all. This is far faster
    than any real network round-trip we saw in a successful call (typically 1–7s).
  • Two calls in the same real burst (main response + an auxiliary call) started only ~10ms
    apart
    — tighter than anything our own thread-pool-based concurrency tests produced, though
    a shared-Client-instance variant with similar timing still didn't reproduce it standalone.

Existing issues checked first

Searched all 49 issues/PRs (open and closed) before filing. Closest related, neither an exact
match:

Question for maintainers

Given native fails near-instantly with zero stderr output and admission.used == False, is there
a known local-only precondition Claude Code checks before ever reaching for the network (e.g.
credentials file staleness/format, a lock/session marker in the config dir, a timestamp-based
"last verified" cache) that could plausibly behave differently on a first real invocation inside a
freshly-spawned, launchd-supervised process versus a manually-invoked shell — something that isn't
env, isn't concurrency, isn't load, and isn't request shape? We're out of angles we can test without
attaching a debugger/strace to native's own process during a live failure, which is why we're
filing this now rather than continuing to guess.

Happy to run further tests against a specific hypothesis if one exists — we have the isolated repro
environment already set up and can reproduce the live-gateway conditions on request.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions