Skip to content

docs(cargo-codspeed): clarify BENCHNAME and --bench help - #195

Merged
not-matthias merged 2 commits into
mainfrom
cod-3789-wizard-run-only-the-targeted-benchmark-cases-anchored
Oct 9, 2026
Merged

not-matthias merged 2 commits into
mainfrom
cod-3789-wizard-run-only-the-targeted-benchmark-cases-anchored

Conversation

@not-matthias

Copy link
Copy Markdown
Member

TLDR: cargo codspeed run --help said BENCHNAME selects benches "containing this string". It is an unanchored regex, and what it matches differs by measurement mode.

  • BENCHNAME help: regex semantics; matched against the CodSpeed URI in simulation/memory and the framework name in walltime (divan crate::mod::fn::arg, criterion group/fn/arg); how to pass -- --exact (divan/criterion only); empty selections exit 0.
  • verbatim_doc_comment keeps the mode list on separate lines in --help.
  • --bench help: notes it can be repeated. Wording stays mode-neutral since build and run share it.

Review notes

  • Docs only. The mode-dependent matching itself is unchanged; making walltime match on the URI is a separate change.

@not-matthias
not-matthias marked this pull request as ready for review October 8, 2026 13:16
@greptile-apps

greptile-apps Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Clarifies help text and refactors benchmark measurement markers.

No new blocking issue was found in the help changes or measurement guards.

What we checked:

  • Manual markers do not overlap: The manual path returns before sample() creates its outer guard.

Summary

The PR expands BENCHNAME help with mode-specific names, regex matching, --exact, and a listing example. It also documents repeated --bench options.

  • Despite the docs-only description, it adds start/stop guards around Criterion’s synchronous and asynchronous manual measurements.
  • Regular Criterion sampling now uses the same guard to stop measurement when the scope exits.
  • No new actionable issues were found.

Reviews (2) · Last reviewed commit: "fix(criterion): send benchmark markers a..." · Reviewed by Greptile

Comment thread crates/cargo-codspeed/src/app.rs Outdated
@codspeed

codspeed Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 7.69%

⚠️ 59 benchmarks measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.

Preventing compiler optimizations

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 12 improved benchmarks
❌ 25 (👁 25) regressed benchmarks
✅ 557 untouched benchmarks

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ Simulation rem 232.3 ns 178.1 ns +30.42%
⚡ Simulation div 232.3 ns 178.1 ns +30.42%
⚡ Simulation find_highest_set_bit[42] 283.5 ns 229.3 ns +23.62%
⚡ Simulation find_highest_set_bit[1024] 283.5 ns 229.3 ns +23.62%
⚡ Simulation find_highest_set_bit[255] 283.5 ns 229.3 ns +23.62%
⚡ Simulation find_highest_set_bit[65535] 283.5 ns 229.3 ns +23.62%
⚡ WallTime bench_array1[42] 41 ns 35 ns +17.14%
⚡ WallTime permutations[6] 75.1 µs 69.7 µs +7.84%
⚡ WallTime add_two_integers[(255, 255)] 20 ns 19 ns +5.26%
⚡ WallTime add_two_integers[(42, 13)] 20 ns 19 ns +5.26%
⚡ WallTime n_queens_solver[4] 2.4 µs 2.2 µs +4.81%
⚡ WallTime generate_parentheses[5] 23.2 µs 22.3 µs +3.82%
👁 WallTime rem 5 ns 6 ns -16.67%
👁 WallTime mul 2 ns 3 ns -33.33%
👁 WallTime hamiltonian_cycle[5] 890 ns 923 ns -3.58%
👁 WallTime iter_with_setup 53 ns 56 ns -5.36%
👁 WallTime iter_batched_ref_large_input 5 ns 6 ns -16.67%
👁 WallTime iter_batched_large_input 9 ns 10 ns -10%
👁 WallTime from_elem_decimal[1024] 210 ns 221 ns -4.98%
👁 Simulation iterative[30] 290.4 ns 346 ns -16.06%
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing cod-3789-wizard-run-only-the-targeted-benchmark-cases-anchored (8d24c95) with main (832f252)

Open in CodSpeed

@not-matthias
not-matthias force-pushed the cod-3789-wizard-run-only-the-targeted-benchmark-cases-anchored branch 2 times, most recently from 11ea0e5 to d3bca8e Compare October 8, 2026 14:02

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

olgtm

Comment thread crates/cargo-codspeed/src/app.rs Outdated
Comment thread crates/criterion_compat/criterion_fork/src/codspeed_iter_manual.rs
Comment thread crates/criterion_compat/criterion_fork/src/routine.rs Outdated
The BENCHNAME help claimed a substring match on benchmark names. It is an unanchored regex forwarded to each bench binary, and the matched string differs by mode: the CodSpeed URI in simulation/memory, the framework's own name in walltime. Document that, how to use --exact, and that an empty selection still exits successfully.

The --bench help now mentions it can be repeated.
@not-matthias
not-matthias force-pushed the cod-3789-wizard-run-only-the-targeted-benchmark-cases-anchored branch from d3bca8e to 8d24c95 Compare October 9, 2026 09:57
@not-matthias

Copy link
Copy Markdown
Member Author

@codspeedbot why did so many simulation benches regress?

@codspeed

codspeed Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

@codspeedbot why did so many simulation benches regress?

Short answer: I don't think these are real regressions. Every flagged simulation benchmark is very small (roughly 100–800 ns), and all of them moved by the same fixed amount of time. That pattern comes from the base and head runs landing on different runners and from code layout, not from slower code in this PR.

What the report shows

  • Every change is the same size, in both directions. Each flagged simulation benchmark is about +55.5 ns slower: Iterative 121.5 → 177.1, count_set_bits[*] 170 → 226, iterative[*] 234 → 290, bencher a 568 → 624, iter_manual_simple 839 → 895. The "improvements" (rem/div, find_highest_set_bit[*]) are about −54 ns faster. A real slowdown would grow with the amount of work, not add a fixed amount. A fixed step like this looks like one extra or one fewer simulated cache miss in the code around the measured region. On a 120 ns benchmark, that alone is a 30% change.
  • The base and head ran on different CPUs. The base is the main run from 3 weeks ago on an AMD EPYC 7763. The head ran on an EPYC 9V45 / 9V74, and on Intel Xeon for the bencher and native benchmarks. The simulation mode builds its cache model from the physical CPU, so a different runner can shift results even when the source is identical. The CI simulation jobs use ubuntu-latest, so the hardware isn't fixed from one run to the next.
  • Some regressed benchmarks are in code this PR doesn't touch. That includes divan's math/the_algorithms and bencher_example. The PR only changes Criterion's marker handling and adds an unused-by-divan BenchmarkMarkers type to the shared codspeed crate.

Same-machine check

I ran main (832f252) and this PR's head (8d24c95) one after the other on the same machine, in simulation mode, with divan math + the_algorithms and Criterion fibonaccis + main (comparison):

Benchmark CI report Same machine
Criterion Iterative / Iterative[20] −31% +39%
divan iterative[5..30] −16 to −19% +14 to +22%
iter_manual_simple −6% +8.5%
rem / div +30% +31%

On one machine the shift is again a fixed amount (about 29 ns), and most of the benchmarks that "regressed" in CI come out faster. A few divan benchmarks the PR doesn't touch (add_two_integers[*]) go the other way by the same 29 ns. The direction depends on the machine and the code layout, not on this PR. Benchmarks that run longer stay inside the noise threshold.

Other notes

  • iter_manual_simple is flagged as dominated by syscalls, so its simulation value understates the real cost and is noisy anyway.
  • The 59 "measured no execution time" benchmarks are the same in base and head, so they aren't related to this PR.

Suggested action: acknowledge these as expected on CodSpeed. If these jumps on sub-microsecond benchmarks keep coming up, two options help: give the tiny benchmarks more work per iteration so one cache miss matters less, or run the simulation jobs on fixed hardware.

Full comparison for the PR

Summary
  • Warming up the performance wizard...
  • Investigate simulation regressions — Base and head simulation runs landed on different CPUs (ubuntu-latest runners); changes are a constant ~55 ns shift, including in divan/bencher which the PR doesn't touch
  • Same-machine A/B check (main vs PR head) — On one machine the same benchmarks shift by a constant ~29 ns, mostly in the opposite direction (Iterative +39%, iterative[N] +14–22%). This is a cache/layout artifact, not a real regression.
  • Reply with findings — Reply with findings
  • Handling comment — Comment handling finished successfully

Was this response helpful? React with 👍 or 👎 to help us improve CodSpeedBot.

@not-matthias
not-matthias merged commit d07bbd0 into main Oct 9, 2026
47 checks passed
@not-matthias
not-matthias deleted the cod-3789-wizard-run-only-the-targeted-benchmark-cases-anchored branch October 9, 2026 15:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants