Skip to content

bug: sandbox unit test metadata_loopback_connect_is_relayed_to_supervisor times out when run with other broker tests #4149

Description

@drew

User Story

I run the openshell-sandbox unit tests on my workstation to validate sandbox runtime changes before opening PRs. While validating the native local accept change for #4058, network_broker::tests::metadata_loopback_connect_is_relayed_to_supervisor failed on every full run, including on unmodified main, so a red suite no longer tells me whether my change broke something.

Problem Statement

The test waits 30 seconds for the broker to queue the workload's TCP open to the loopback metadata address (127.0.0.1:8174). When it runs alongside the other broker tests, that wait times out:

thread 'network_broker::tests::metadata_loopback_connect_is_relayed_to_supervisor' panicked at crates/openshell-sandbox/src/network_broker.rs:
called `Result::unwrap()` on an `Err` value: Elapsed(())

Observed on an aarch64 Linux workstation (debug build, cargo test -p openshell-sandbox --lib):

Invocation Result
Full lib suite, default parallelism, on main (a48920a) Fails
Full lib suite, default parallelism, on branch fix/4058-native-local-accept/drew Fails 4 of 4 runs
network_broker::tests:: only, default parallelism Fails
Test alone (--exact) Passes, but takes 22.4 s of its 30 s budget
Full lib suite with --test-threads=1 Passes

The same test binary passed the network_broker::tests:: subset inside lightly loaded RHCOS VMs.

Impact / Why This Matters

The openshell-sandbox unit suite fails locally on main, which hides real regressions in the network broker and trains contributors to ignore failures in this area. Rerunning with --test-threads=1 works around it but takes 148 s instead of about 95 s and is not what cargo test or the default tasks run.

Acceptance Criteria

  • cargo test -p openshell-sandbox --lib passes with default parallelism on a development workstation.
  • The test's timing budget is explained by measured costs, not set to an arbitrary larger timeout.

Reproduction Steps

  1. Check out main.
  2. Run cargo test -p openshell-sandbox --lib.
  3. Observe metadata_loopback_connect_is_relayed_to_supervisor fail with Elapsed(()); rerun it alone with --exact and note its duration.

Environment

  • OpenShell: main at a48920a
  • OS: Linux 6.17, aarch64, 20 cores
  • Runtime, deployment, or integration: unit tests, debug profile (Rust 1.95.0)

Agent Investigation

Unconfirmed hypothesis: the 30 s wait includes executable identity resolution for the connecting thread, which SHA-256-hashes the executable. In tests that executable is the debug test binary (133 MiB). With the dev profile, sha2 0.10 hashes that file in 6.4 s, compared with 0.24 s in release on the same host. Each test constructs its own NetworkBroker and ProcfsIdentityResolver, so the digest cache is not shared and concurrently running broker tests repeat the hashing. Only this test waits on the metadata relay path with a 30 s bound. A profiling run should confirm where the 22 s go before choosing a fix; candidates include optimizing sha2 in the dev profile ([profile.dev.package.sha2] opt-level = 3), sharing the identity cache in tests, or deriving the timeout from measured cost.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions