Conversation
Compares ubuntu-latest with Blacksmith 4 and 8 vCPU runners on identical work: a cold go build, go vet, go test -race and fixed-iteration benchmarks, five attempts each on the same commit with a fresh build cache. The report job publishes medians, speedups and runner queue times in the job summary, plus a benchstat comparison. Runs on workflow_dispatch or a push to ci/runner-benchmark, so it stays out of regular CI. Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
Blacksmith enables a transparent Go build cache through GOCACHEPROG, which bypasses GOCACHE, so the first run's "cold" builds were warm on Blacksmith. Cold timings now unset GOCACHEPROG and use an empty GOCACHE; separate "default setup" timings keep each runner as it comes. The report also normalizes benchmark names and CPU lines so benchstat compares the three runners in one table, and runs the benchmarks twice per attempt (10 samples per runner) so benchstat can report confidence intervals. Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
The "default setup" timings ran GitHub-hosted runners with an empty build cache, while GitHub's recommended setup restores the Actions cache through actions/setup-go. That overstated Blacksmith's advantage. A warm-up job now fills the Actions cache first, ubuntu-latest jobs restore it with setup-go (cache: true), and Blacksmith jobs keep setup-go caching off, as Blacksmith's docs recommend, since its Go build cache replaces it. The report adds the Set up Go step duration, which includes the cache restore on GitHub. Cold timings are unchanged. Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
setup-go keys its cache on go.sum only. The warm-up job restored a stale master entry from September, hit the primary key and skipped saving, so the GitHub "default setup" timings ran effectively uncached. The warm-up job now saves the Go build and module caches with actions/cache under a key unique to the run, and ubuntu-latest jobs restore that exact entry (failing on a miss). The restore is timed with the Set up Go row. The CPU row now lists every CPU model seen, since both providers mix hardware across jobs. Claude-Session: https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a manual workflow that compares CI runners on identical work, so we can report a fair Blacksmith vs GitHub-hosted comparison.
What it measures
Each runner type runs the same commit 5 times with a cold build cache (
GOCACHEin a fresh temp dir,setup-gocache off):go build ./...(cold, including the standard library)go vet ./...go test -race -count=1 ./...-benchtime=2000x), so a faster CPU finishes sooner instead of running more iterationsRunners:
ubuntu-latestandblacksmith-4vcpu-ubuntu-2404(both 4 vCPUs, like for like) andblacksmith-8vcpu-ubuntu-2404(the tier regular CI uses).The report job publishes medians, the speedup relative to
ubuntu-latest, runner queue times (job created to started), vCPUs, RAM and CPU model in the job summary, plus abenchstatcomparison (full output in the report artifact).Why a separate workflow
The existing benchmark job runs each benchmark for a fixed time (
-count=8at the defaultbenchtime), so it takes about 20 minutes on any runner, and single CI runs mix queue time, cache state and noise. This workflow isolates CPU speed.It runs on
workflow_dispatchor a push toci/runner-benchmark, so it stays out of regular CI. Stacked on #3128 (addsblacksmith-4vcpu-ubuntu-2404to.github/actionlint.yaml).https://claude.ai/code/session_01QKDYQr53zNKkR7nif2CAAq