Skip to content

ci: add runner benchmark workflow - #3137

Draft
vishr wants to merge 4 commits into
ci/blacksmithfrom
ci/runner-benchmark
Draft

vishr wants to merge 4 commits into
ci/blacksmithfrom
ci/runner-benchmark

Conversation

@vishr

@vishr vishr commented Oct 1, 2026

Copy link
Copy Markdown
Member

Adds a manual workflow that compares CI runners on identical work, so we can report a fair Blacksmith vs GitHub-hosted comparison.

What it measures

Each runner type runs the same commit 5 times with a cold build cache (GOCACHE in a fresh temp dir, setup-go cache off):

  • go build ./... (cold, including the standard library)
  • go vet ./...
  • go test -race -count=1 ./...
  • all benchmarks with a fixed iteration count (-benchtime=2000x), so a faster CPU finishes sooner instead of running more iterations

Runners: ubuntu-latest and blacksmith-4vcpu-ubuntu-2404 (both 4 vCPUs, like for like) and blacksmith-8vcpu-ubuntu-2404 (the tier regular CI uses).

The report job publishes medians, the speedup relative to ubuntu-latest, runner queue times (job created to started), vCPUs, RAM and CPU model in the job summary, plus a benchstat comparison (full output in the report artifact).

Why a separate workflow

The existing benchmark job runs each benchmark for a fixed time (-count=8 at the default benchtime), so it takes about 20 minutes on any runner, and single CI runs mix queue time, cache state and noise. This workflow isolates CPU speed.

It runs on workflow_dispatch or a push to ci/runner-benchmark, so it stays out of regular CI. Stacked on #3128 (adds blacksmith-4vcpu-ubuntu-2404 to .github/actionlint.yaml).

https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq

vishr added 4 commits October 1, 2026 09:52
Compares ubuntu-latest with Blacksmith 4 and 8 vCPU runners on identical
work: a cold go build, go vet, go test -race and fixed-iteration
benchmarks, five attempts each on the same commit with a fresh build cache.
The report job publishes medians, speedups and runner queue times in the
job summary, plus a benchstat comparison.

Runs on workflow_dispatch or a push to ci/runner-benchmark, so it stays
out of regular CI.

Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
Blacksmith enables a transparent Go build cache through GOCACHEPROG,
which bypasses GOCACHE, so the first run's "cold" builds were warm on
Blacksmith. Cold timings now unset GOCACHEPROG and use an empty
GOCACHE; separate "default setup" timings keep each runner as it comes.

The report also normalizes benchmark names and CPU lines so benchstat
compares the three runners in one table, and runs the benchmarks twice
per attempt (10 samples per runner) so benchstat can report confidence
intervals.

Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
The "default setup" timings ran GitHub-hosted runners with an empty build
cache, while GitHub's recommended setup restores the Actions cache through
actions/setup-go. That overstated Blacksmith's advantage.

A warm-up job now fills the Actions cache first, ubuntu-latest jobs restore
it with setup-go (cache: true), and Blacksmith jobs keep setup-go caching
off, as Blacksmith's docs recommend, since its Go build cache replaces it.
The report adds the Set up Go step duration, which includes the cache
restore on GitHub. Cold timings are unchanged.

Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
setup-go keys its cache on go.sum only. The warm-up job restored a stale
master entry from September, hit the primary key and skipped saving, so
the GitHub "default setup" timings ran effectively uncached.

The warm-up job now saves the Go build and module caches with
actions/cache under a key unique to the run, and ubuntu-latest jobs
restore that exact entry (failing on a miss). The restore is timed with
the Set up Go row. The CPU row now lists every CPU model seen, since
both providers mix hardware across jobs.

Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant