A compact, lossless wire encoding for agent-to-agent structured payloads. Line-oriented UTF-8, schema-once record arrays, counts that double as checksums, and a spec'd row-number field that fixes the one thing every text format fails at: a model addressing "the Nth row".
!eson/1
from=reviewer
to=implementer
kind=code_review
findings[2]{n,severity,file,line,message}
1 high src/auth.js 42 token never expires
2 medium src/api.js 18 missing rate limit
meta{complete,retry}
true null
ESON is a payload encoding, not a protocol — it rides inside whatever transport your agents already use (an orchestrator's tool results, A2A/MCP messages, a queue). Its readers are language models and programs, so the repo ships the two things both need: conformance vectors for parsers and a canonical primer for models.
For end-to-end pipes, two thin layers ride on top — think Content-Encoding,
not a new HTTP: the Honey Wire Profile (eight testable rules
for token-efficient, integrity-checked agent messages; eson lint checks the
per-message ones) and wire negotiation (capability
tokens, self-identifying payloads, mandatory compact-JSON fallback — adoption
is never a compatibility bet).
Token cost on realistic handoff documents (o200k tokenizer, npm run bench:formats):
| Format | Valid JSON? | vs compact JSON |
|---|---|---|
| JSON (pretty) | yes | +55% ← what models emit unprompted |
| JSON (compact) | yes | 0% |
| TOON | no | −20% |
| JSON (columnar) | yes | −22% |
| ESON | no | −28% |
Model comprehension (Claude + GPT, small and frontier): ties at 100% across
formats for every realistic access pattern — key lookup, field match, nested
extraction. The only failures are format-independent model limits: positional
access and filtered counting. ESON addresses the first in-spec (the reserved
n field, §5 — 0–17% → 100% for ~+6–9% tokens) and forbids relying on the
second (aggregate in code, never in the model).
| Your pipe | Use |
|---|---|
| Low volume, no prompt caching, or scalar-heavy envelopes | compact JSON — ESON's ~126-token primer never amortizes |
| Must stay valid JSON (stdlib-only readers, strict tooling) | columnar JSON (−22%): keys once — {"#c":[cols],"#r":[rows]} |
| High-volume, cached, record-array-heavy, both ends yours | ESON (−28%, −35% on record arrays) |
| Models may address rows positionally ("the 3rd finding") | add the n field (any format; spec'd + checksum-verified in ESON) |
| Auth, money, migrations, deletes, anything irreversible | schema-validated JSON — never a compact format |
Break-even vs columnar JSON: ~2 record-heavy messages with prompt caching;
never without caching; never for scalar-heavy traffic (npm run bench:primer).
If you take one thing from this table: stop emitting pretty-printed JSON
between agents (+55% for nothing).
npm install # dev deps are only for the benchmarks
npm test # JS + Python suites + shared conformance vectors
echo '{"findings":[{"sev":"high","msg":"no auth"}]}' | node bin/eson.js encode --number
node bin/eson.js decode < doc.eson
node bin/eson.js lint < message.json # Honey Wire Profile: W1/W4/W5/W6, exit 1 on MUST violationconst { encode, decode, tryDecode } = require("eson-format");
const wire = encode({ findings }, { number: true });
const { ok, value } = tryDecode(wire); // non-throwing, for message routersimport eson
wire = eson.encode({"findings": findings}, number=True)
data = eson.decode(wire) # raises eson.ESONError on malformed inputBoth implementations are dependency-free single files (js/index.js, py/eson.py) — vendoring is fine.
ESON pays off wherever one model's structured output becomes another model's input, at volume, behind a prompt cache:
- Orchestrator ↔ subagent handoffs — a reviewer agent returns 30 findings to an orchestrator; a scout agent returns a file/symbol map. Record arrays, many messages per session: ESON's sweet spot (−35% on record arrays).
- Tool results fed to a model — a scanner, linter, or search tool returns hundreds of uniform rows that go straight into an agent's context. Encode once at the tool boundary; every call saves.
- Queue/batch pipelines between LLM workers — classify → extract → verify stages passing record batches through a queue. Both ends are yours, traffic is high, the primer is cached: full −28%.
- Multi-agent fleets with a shared system prompt — one canonical primer in the fleet's base prompt covers every agent; per-message savings are pure.
Where it does not pay (use compact/columnar JSON): one-shot calls, scalar envelopes, uncached pipes, third-party receivers you can't put a primer in front of, and anything irreversible (auth, money, migrations, deletes — see the decision table above).
Four steps. The codec handles the program side; the primer handles the model side.
1. Get the codec into both ends. npm install eson-format, or vendor the
single file (js/index.js or py/eson.py — zero deps).
2. Put the canonical primer in the receiving model's system prompt. Copy it verbatim from PRIMER.md (~126 tokens, cacheable). If you emit numbered rows, append the numbered-rows addendum too. Don't paraphrase it — the canonical text is what the comprehension benchmarks were run with.
3. Sender: encode at the boundary. Use number: true whenever a model
might address rows positionally ("the 3rd finding"):
// orchestrator hands findings to a fixer agent
const { encode } = require("eson-format");
const payload = encode({ from: "reviewer", findings }, { number: true });
await spawnAgent({ system: BASE_PROMPT + "\n" + ESON_PRIMER, message: payload });4. Receiver: decode in the program, verify, fall back. Programs decode before routing; models read the wire text directly (that's the point — no decode step in the model's context):
const { tryDecode } = require("eson-format");
const { ok, value } = tryDecode(msg); // non-throwing
if (!ok) handleAsCompactJSON(msg); // mandatory fallback (NEGOTIATION.md)import eson
try:
data = eson.decode(msg)
except eson.ESONError:
data = json.loads(msg) # fallback pathThe [N] count is a checksum — a primed model must report corruption on
mismatch instead of answering from a truncated payload.
Hardening for production pipes:
- Lint outbound messages in CI or at send time:
node bin/eson.js lint < message.json(Honey Wire Profile W1/W4/W5/W6, exit 1 on MUST violation). - Mixed fleet? Negotiate: advertise
eson/1as a capability token, send compact JSON to receivers that didn't ack (NEGOTIATION.md). - Prove your implementation with the conformance vectors: vectors/vectors.json — pass them all or it isn't ESON.
- Shell pipes work too, no library integration needed:
your-tool --json | node bin/eson.js encode --number.
Comprehension was benchmarked on Claude and GPT families — these all use
the same two moves: primer in the system-side prompt, encode at the boundary.
Claude Code. Paste the primer into CLAUDE.md (or a skill) — every
session and subagent can then read ESON. Two hook points:
- Subagent returns: instruct agents in
.claude/agents/*.mdto return their final result as ESON — the orchestrator pays −28% on every handoff it reads. - Tool output: pipe bulky JSON through the CLI before it enters context:
gh api ... | node bin/eson.js encode --number.
Claude Agent SDK. Append the primer to the system prompt; feed encoded payloads as the message:
import { query } from "@anthropic-ai/claude-agent-sdk";
const ESON_PRIMER = "..."; // the base primer, verbatim from PRIMER.md
for await (const msg of query({
prompt: encode({ findings }, { number: true }),
options: { systemPrompt: { type: "preset", preset: "claude_code", append: ESON_PRIMER } },
})) { /* ... */ }Anthropic API (raw). Put the primer in a cached system block — the cache is what amortizes it:
client.messages.create(
model="claude-sonnet-5",
system=[{"type": "text", "text": ESON_PRIMER,
"cache_control": {"type": "ephemeral"}}],
messages=[{"role": "user", "content": eson.encode({"findings": findings}, number=True)}],
)OpenAI (Codex / GPT). Same pattern: primer in AGENTS.md for Codex, or in
the system/instructions field for the API; encode tool outputs and
inter-agent messages. GPT models scored identically to Claude on the
comprehension suite — including the n-field fix for positional access.
MCP servers. Return ESON as the tool result's text content for
record-heavy tools (search hits, scan findings, query rows); the client agent
carries the primer. Keep the negotiation rule: only emit ESON when the client
declared eson/1, else compact JSON (NEGOTIATION.md).
Custom harnesses (LangGraph, CrewAI, hand-rolled). Encode on every
model-bound edge, decode in program code at router/aggregator nodes
(tryDecode → compact-JSON fallback). Aggregate counts and filters in the
node, never in the model — pass results, not raw rows to count.
- SPEC.md — normative spec, v1.1 (
!eson/1wire format) - PROFILE.md — the Honey Wire Profile: rules W1–W8 for any agent pipe
- NEGOTIATION.md — encoding negotiation and fallback
- PRIMER.md — canonical model primer; part of the wire contract
- vectors/vectors.json — conformance vectors; an
implementation conforms iff it passes them (
vectors/generate.jsregenerates) - js/, py/ — reference implementations + tests
- bench/ — deterministic token benchmarks with committed results:
FORMATS.md (
npm run bench:formats), PRIMER-COST.md (npm run bench:primer), and the live end-to-end WIRE.md (node bench/wire.mjs, needsANTHROPIC_API_KEY: primer + dispatch line → 100% valid ESON, −24% wire tokens at 100% recovery); the model-comprehension methodology lives in the honey benchmark suite
Extracted from Honey, where ESON is the Lever-3 (agent-to-agent) format, validated against JSON/columnar/TOON on token efficiency and model comprehension across Claude and GPT model families.
MIT © Green-PT