Skip to content

docs(submitqueue): [DO NOT REVIEW][DO NOT LAND] Tango conflict research - #763

Draft
sbalabanov wants to merge 10 commits into
mainfrom
research/tango-conflict-graphs
Draft

sbalabanov wants to merge 10 commits into
mainfrom
research/tango-conflict-graphs

Conversation

@sbalabanov

@sbalabanov sbalabanov commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

DRAFT RESEARCH ONLY — DO NOT REVIEW; DO NOT LAND. Leave this PR in draft. Do not request reviewers, mark it ready for review, or enable automatic merging. This branch contains a Go evaluator and a documented Tango/SubmitQueue conflict-analysis proposal; it does not deploy an analyzer.

Report: doc/rfc/submitqueue/tango-conflict-analysis.md. The earlier HTML report (https://personal.uberinternal.com/sergeyb/submitqueue-tango-conflict-analysis/index.html) reflects the superseded fixed-N workload and has not been updated.

Why?

A Tango-backed conflict analyzer stores a target signature per batch and compares each new batch against every in-flight batch. This research estimates, for a cold controller that must fetch and decode all B signatures, at what number of in-flight batches each signature design stops working because of size, materialization time or latency. go-code runs fewer than 100 in-flight batches today; the goal is to pick a design that scales.

What?

  • Captured the complete go-code Bazel graph (2.97M targets) and measured its footprint.
  • Pulled the exact per-request affected-target histogram for the go queue from Hive (58,876 requests; median 42, mean 2,527, max 2.88M). One batch is modeled as one request.
  • Benchmarked three designs on closure-shaped affected sets drawn from that histogram, at B = 100–10,000: TangoSnapshot (baseline: the Tango response with dependency metadata, raw or gzip), NameKey (sorted labels) and ID64 (sorted 64-bit IDs plus an ID→name dictionary for conflict relaxation). For each, it serializes, deserializes and checks one candidate. A per-target Monte Carlo model extends the results to B = 100,000 and reports the first B exceeding example latency and heap budgets, under object-store and KV point-read profiles.
  • Result: TangoSnapshot breaks on a single tail request at any B. NameKey holds to low thousands. ID64 is limited by round trips rather than bytes, so the choice of store matters more than the encoding.
  • Simulated signatures computed once against the newest base. Plain affected-set overlap misses conflicts. Adding new-edge destinations (the production SubmitQueue rule) reduces but does not remove misses. Inheriting the signatures of structural dependencies removes them, at a cost in false conflicts. The report documents this as a policy trade-off.
  • Removed superseded artifacts, the posting-index and cold-scan modes, and the root-level summary file.

Test Plan

  • Evaluator unit tests: go test ./tool/tangograph-eval and ./tool/bazel test //tool/tangograph-eval/....
  • BUILD and dependency checks: make check-gazelle && make check-tidy.
  • Re-run the benchmark and the simulation with the commands under "Reproduce" in the report. The benchmark needs the local, uncommitted 2.8 GB Bazel stream; the last run took 13 min 48 s and peaked at 35 GB RSS.
  • No live Tango endpoint, SubmitQueue controller or storage backend was changed or tested.

Revert Plan

Close this draft PR without merging. If it is ever merged contrary to this instruction, revert the research commits; they add only documentation and a standalone evaluator, not service behavior.

Issue Links

None — exploratory research requested directly; no tracking issue was provided.

🤖 Generated with Claude Code

AI Verification

Validated at f223612 on Sep 30 22:58 UTC · 40 files analyzed · 3s

Validator Status Issues
web-coverage not_applicable 0
android-lint not_applicable 0
web-lint not_applicable 0
web-unit not_applicable 0
merge-conflict not_applicable 0
arc-lint not_applicable 0
app-validation not_applicable 0
ios-lint not_applicable 0
web-repocheck not_applicable 0
diff-template not_applicable 0
fix-disclosure not_applicable 0
android-coverage not_applicable 0
java-lint not_applicable 0
ios-test not_applicable 0
web-typecheck not_applicable 0
go-gazelle not_applicable 0
arc-unit not_applicable 0
go-proto-lint not_applicable 0
go-thrift-lint not_applicable 0
java-coverage not_applicable 0
uber-one not_applicable 0
go-coverage failed 0
artifacts completed 0
ureview completed 0
custom not_applicable 0
go-lint failed 0

0 issues detected

Skipped validators: claude · EngWiki

Prior runs

Run at f223612 on Sep 30 22:57 UTC · 40 files · 3s · 0 issues detected

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@sbalabanov sbalabanov changed the title feat(tool): measure serialized conflict payload loading and checks docs(submitqueue): [DO NOT REVIEW][DO NOT LAND] Tango conflict research Sep 30, 2026
… distribution

Redo the Tango conflict-analysis research against the observed per-request
affected-target histogram instead of fixed N/2N/5N workloads, and answer where
each signature design stops working as in-flight batches grow.

- Benchmark: draw each batch's size from the Hive histogram (one batch = one
  request), build closure-shaped affected sets on the full go-code graph, and
  give TangoSnapshot real direct dependencies. Run B = 100..10,000 and extend
  to 100,000 with a per-target Monte Carlo scale model that reports the first
  B exceeding illustrative latency/heap budgets, under object-store and KV
  point-read profiles.
- Base drift: simulate signatures computed once against the newest base. Show
  that affected-set overlap misses conflicts, that new-edge destinations (the
  production SubmitQueue rule) reduce but do not remove misses, and that
  inheriting structural dependencies' signatures removes them at a cost in
  false conflicts.
- Report: rewrite the RFC around the cold-controller worst case with
  TangoSnapshot as the baseline, a design vocabulary, units in every table
  header and an interpretation under each result table; document the ID
  dictionary's role in conflict relaxation and prior art; drop superseded
  sections, artifacts, the posting-index and cold-scan modes, and the
  root-level summary file.
- Fix per-batch decode timing truncation to whole microseconds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sbalabanov
sbalabanov force-pushed the research/tango-conflict-graphs branch from 5519c18 to fa126c3 Compare October 1, 2026 22:16

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants