You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Runs ~20 options strategies against live market data in shadow mode, records every hypothetical fill under worst/base/optimistic assumptions, and grades each with anytime-valid e-processes. Places no orders.
Which model is really behind your API relay or agent IDE? Behavioral fingerprinting + anytime-valid sequential tests (FPR<=1%). LLMs can't be random - measured on 9 frontier models.
Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.
A/B testing and causal inference scored against known answers: simulations where I set the effect, and a randomised benchmark the observational methods have to recover.
Ships ML models in stages - shadow, then 1/5/25/50% of traffic - and rolls a bad one back automatically. The guardrails stay valid under constant checking: 0.6% false rollback with two identical models, where a repeatedly-read A/B test acts wrongly 36.7% of the time.
Statistically rigorous real-world evaluation for robot policies. Plan how many rollouts you need, run them randomized and blinded from a phone at the robot, let fieldtrial switch LeRobot or openpi policies itself, and get a report that only says "better" when a pre-registered test does.
Rank-targeted nested sequential design for LLM evaluation: reach the same ranking conclusion for less, and see which comparisons the data never supported.
Agent Skill for causal measurement in marketing: sample ratio mismatch, an always-valid sequential test that survives daily peeking, CUPED variance reduction, sample sizing, and geo/holdout designs for channels where you cannot randomise users. Zero dependencies.
A/B testing and experiment statistics built from scratch - exact t-distribution p-values, Wilson intervals, power and sample size, CUPED, and O'Brien-Fleming and Pocock sequential monitoring. Pure Python, no scipy, no statsmodels.
Feature flags and A/B testing, built to show when experiment results can be trusted: sequential testing, sample ratio mismatch checks, and Monte Carlo-validated statistics.