User Story
I run the openshell-sandbox unit tests on my workstation to validate sandbox runtime changes before opening PRs. While validating the native local accept change for #4058, network_broker::tests::metadata_loopback_connect_is_relayed_to_supervisor failed on every full run, including on unmodified main, so a red suite no longer tells me whether my change broke something.
Problem Statement
The test waits 30 seconds for the broker to queue the workload's TCP open to the loopback metadata address (127.0.0.1:8174). When it runs alongside the other broker tests, that wait times out:
thread 'network_broker::tests::metadata_loopback_connect_is_relayed_to_supervisor' panicked at crates/openshell-sandbox/src/network_broker.rs:
called `Result::unwrap()` on an `Err` value: Elapsed(())
Observed on an aarch64 Linux workstation (debug build, cargo test -p openshell-sandbox --lib):
| Invocation |
Result |
Full lib suite, default parallelism, on main (a48920a) |
Fails |
Full lib suite, default parallelism, on branch fix/4058-native-local-accept/drew |
Fails 4 of 4 runs |
network_broker::tests:: only, default parallelism |
Fails |
Test alone (--exact) |
Passes, but takes 22.4 s of its 30 s budget |
Full lib suite with --test-threads=1 |
Passes |
The same test binary passed the network_broker::tests:: subset inside lightly loaded RHCOS VMs.
Impact / Why This Matters
The openshell-sandbox unit suite fails locally on main, which hides real regressions in the network broker and trains contributors to ignore failures in this area. Rerunning with --test-threads=1 works around it but takes 148 s instead of about 95 s and is not what cargo test or the default tasks run.
Acceptance Criteria
Reproduction Steps
- Check out
main.
- Run
cargo test -p openshell-sandbox --lib.
- Observe
metadata_loopback_connect_is_relayed_to_supervisor fail with Elapsed(()); rerun it alone with --exact and note its duration.
Environment
- OpenShell:
main at a48920a
- OS: Linux 6.17, aarch64, 20 cores
- Runtime, deployment, or integration: unit tests, debug profile (Rust 1.95.0)
Agent Investigation
Unconfirmed hypothesis: the 30 s wait includes executable identity resolution for the connecting thread, which SHA-256-hashes the executable. In tests that executable is the debug test binary (133 MiB). With the dev profile, sha2 0.10 hashes that file in 6.4 s, compared with 0.24 s in release on the same host. Each test constructs its own NetworkBroker and ProcfsIdentityResolver, so the digest cache is not shared and concurrently running broker tests repeat the hashing. Only this test waits on the metadata relay path with a 30 s bound. A profiling run should confirm where the 22 s go before choosing a fix; candidates include optimizing sha2 in the dev profile ([profile.dev.package.sha2] opt-level = 3), sharing the identity cache in tests, or deriving the timeout from measured cost.
User Story
I run the
openshell-sandboxunit tests on my workstation to validate sandbox runtime changes before opening PRs. While validating the native local accept change for #4058,network_broker::tests::metadata_loopback_connect_is_relayed_to_supervisorfailed on every full run, including on unmodifiedmain, so a red suite no longer tells me whether my change broke something.Problem Statement
The test waits 30 seconds for the broker to queue the workload's TCP open to the loopback metadata address (
127.0.0.1:8174). When it runs alongside the other broker tests, that wait times out:Observed on an aarch64 Linux workstation (debug build,
cargo test -p openshell-sandbox --lib):main(a48920a)fix/4058-native-local-accept/drewnetwork_broker::tests::only, default parallelism--exact)--test-threads=1The same test binary passed the
network_broker::tests::subset inside lightly loaded RHCOS VMs.Impact / Why This Matters
The
openshell-sandboxunit suite fails locally onmain, which hides real regressions in the network broker and trains contributors to ignore failures in this area. Rerunning with--test-threads=1works around it but takes 148 s instead of about 95 s and is not whatcargo testor the default tasks run.Acceptance Criteria
cargo test -p openshell-sandbox --libpasses with default parallelism on a development workstation.Reproduction Steps
main.cargo test -p openshell-sandbox --lib.metadata_loopback_connect_is_relayed_to_supervisorfail withElapsed(()); rerun it alone with--exactand note its duration.Environment
mainat a48920aAgent Investigation
Unconfirmed hypothesis: the 30 s wait includes executable identity resolution for the connecting thread, which SHA-256-hashes the executable. In tests that executable is the debug test binary (133 MiB). With the dev profile,
sha20.10 hashes that file in 6.4 s, compared with 0.24 s in release on the same host. Each test constructs its ownNetworkBrokerandProcfsIdentityResolver, so the digest cache is not shared and concurrently running broker tests repeat the hashing. Only this test waits on the metadata relay path with a 30 s bound. A profiling run should confirm where the 22 s go before choosing a fix; candidates include optimizingsha2in the dev profile ([profile.dev.package.sha2] opt-level = 3), sharing the identity cache in tests, or deriving the timeout from measured cost.