Skip to content

Repository files navigation

Honey ESON — Efficient Structured Object Notation

A compact, lossless wire encoding for agent-to-agent structured payloads. Line-oriented UTF-8, schema-once record arrays, counts that double as checksums, and a spec'd row-number field that fixes the one thing every text format fails at: a model addressing "the Nth row".

!eson/1
from=reviewer
to=implementer
kind=code_review
findings[2]{n,severity,file,line,message}
1	high	src/auth.js	42	token never expires
2	medium	src/api.js	18	missing rate limit
meta{complete,retry}
true	null

ESON is a payload encoding, not a protocol — it rides inside whatever transport your agents already use (an orchestrator's tool results, A2A/MCP messages, a queue). Its readers are language models and programs, so the repo ships the two things both need: conformance vectors for parsers and a canonical primer for models.

For end-to-end pipes, two thin layers ride on top — think Content-Encoding, not a new HTTP: the Honey Wire Profile (eight testable rules for token-efficient, integrity-checked agent messages; eson lint checks the per-message ones) and wire negotiation (capability tokens, self-identifying payloads, mandatory compact-JSON fallback — adoption is never a compatibility bet).

Why (measured, not vibes)

Token cost on realistic handoff documents (o200k tokenizer, npm run bench:formats):

Format Valid JSON? vs compact JSON
JSON (pretty) yes +55% ← what models emit unprompted
JSON (compact) yes 0%
TOON no −20%
JSON (columnar) yes −22%
ESON no −28%

Model comprehension (Claude + GPT, small and frontier): ties at 100% across formats for every realistic access pattern — key lookup, field match, nested extraction. The only failures are format-independent model limits: positional access and filtered counting. ESON addresses the first in-spec (the reserved n field, §5 — 0–17% → 100% for ~+6–9% tokens) and forbids relying on the second (aggregate in code, never in the model).

When to use what — honest decision table

Your pipe Use
Low volume, no prompt caching, or scalar-heavy envelopes compact JSON — ESON's ~126-token primer never amortizes
Must stay valid JSON (stdlib-only readers, strict tooling) columnar JSON (−22%): keys once — {"#c":[cols],"#r":[rows]}
High-volume, cached, record-array-heavy, both ends yours ESON (−28%, −35% on record arrays)
Models may address rows positionally ("the 3rd finding") add the n field (any format; spec'd + checksum-verified in ESON)
Auth, money, migrations, deletes, anything irreversible schema-validated JSON — never a compact format

Break-even vs columnar JSON: ~2 record-heavy messages with prompt caching; never without caching; never for scalar-heavy traffic (npm run bench:primer). If you take one thing from this table: stop emitting pretty-printed JSON between agents (+55% for nothing).

Use it

npm install          # dev deps are only for the benchmarks
npm test             # JS + Python suites + shared conformance vectors

echo '{"findings":[{"sev":"high","msg":"no auth"}]}' | node bin/eson.js encode --number
node bin/eson.js decode < doc.eson
node bin/eson.js lint < message.json   # Honey Wire Profile: W1/W4/W5/W6, exit 1 on MUST violation
const { encode, decode, tryDecode } = require("eson-format");
const wire = encode({ findings }, { number: true });
const { ok, value } = tryDecode(wire); // non-throwing, for message routers
import eson
wire = eson.encode({"findings": findings}, number=True)
data = eson.decode(wire)   # raises eson.ESONError on malformed input

Both implementations are dependency-free single files (js/index.js, py/eson.py) — vendoring is fine.

Use cases

ESON pays off wherever one model's structured output becomes another model's input, at volume, behind a prompt cache:

  • Orchestrator ↔ subagent handoffs — a reviewer agent returns 30 findings to an orchestrator; a scout agent returns a file/symbol map. Record arrays, many messages per session: ESON's sweet spot (−35% on record arrays).
  • Tool results fed to a model — a scanner, linter, or search tool returns hundreds of uniform rows that go straight into an agent's context. Encode once at the tool boundary; every call saves.
  • Queue/batch pipelines between LLM workers — classify → extract → verify stages passing record batches through a queue. Both ends are yours, traffic is high, the primer is cached: full −28%.
  • Multi-agent fleets with a shared system prompt — one canonical primer in the fleet's base prompt covers every agent; per-message savings are pure.

Where it does not pay (use compact/columnar JSON): one-shot calls, scalar envelopes, uncached pipes, third-party receivers you can't put a primer in front of, and anything irreversible (auth, money, migrations, deletes — see the decision table above).

How to implement it

Four steps. The codec handles the program side; the primer handles the model side.

1. Get the codec into both ends. npm install eson-format, or vendor the single file (js/index.js or py/eson.py — zero deps).

2. Put the canonical primer in the receiving model's system prompt. Copy it verbatim from PRIMER.md (~126 tokens, cacheable). If you emit numbered rows, append the numbered-rows addendum too. Don't paraphrase it — the canonical text is what the comprehension benchmarks were run with.

3. Sender: encode at the boundary. Use number: true whenever a model might address rows positionally ("the 3rd finding"):

// orchestrator hands findings to a fixer agent
const { encode } = require("eson-format");
const payload = encode({ from: "reviewer", findings }, { number: true });
await spawnAgent({ system: BASE_PROMPT + "\n" + ESON_PRIMER, message: payload });

4. Receiver: decode in the program, verify, fall back. Programs decode before routing; models read the wire text directly (that's the point — no decode step in the model's context):

const { tryDecode } = require("eson-format");
const { ok, value } = tryDecode(msg);        // non-throwing
if (!ok) handleAsCompactJSON(msg);           // mandatory fallback (NEGOTIATION.md)
import eson
try:
    data = eson.decode(msg)
except eson.ESONError:
    data = json.loads(msg)                   # fallback path

The [N] count is a checksum — a primed model must report corruption on mismatch instead of answering from a truncated payload.

Hardening for production pipes:

  • Lint outbound messages in CI or at send time: node bin/eson.js lint < message.json (Honey Wire Profile W1/W4/W5/W6, exit 1 on MUST violation).
  • Mixed fleet? Negotiate: advertise eson/1 as a capability token, send compact JSON to receivers that didn't ack (NEGOTIATION.md).
  • Prove your implementation with the conformance vectors: vectors/vectors.json — pass them all or it isn't ESON.
  • Shell pipes work too, no library integration needed: your-tool --json | node bin/eson.js encode --number.

Platform recipes

Comprehension was benchmarked on Claude and GPT families — these all use the same two moves: primer in the system-side prompt, encode at the boundary.

Claude Code. Paste the primer into CLAUDE.md (or a skill) — every session and subagent can then read ESON. Two hook points:

  • Subagent returns: instruct agents in .claude/agents/*.md to return their final result as ESON — the orchestrator pays −28% on every handoff it reads.
  • Tool output: pipe bulky JSON through the CLI before it enters context: gh api ... | node bin/eson.js encode --number.

Claude Agent SDK. Append the primer to the system prompt; feed encoded payloads as the message:

import { query } from "@anthropic-ai/claude-agent-sdk";
const ESON_PRIMER = "..."; // the base primer, verbatim from PRIMER.md

for await (const msg of query({
  prompt: encode({ findings }, { number: true }),
  options: { systemPrompt: { type: "preset", preset: "claude_code", append: ESON_PRIMER } },
})) { /* ... */ }

Anthropic API (raw). Put the primer in a cached system block — the cache is what amortizes it:

client.messages.create(
    model="claude-sonnet-5",
    system=[{"type": "text", "text": ESON_PRIMER,
             "cache_control": {"type": "ephemeral"}}],
    messages=[{"role": "user", "content": eson.encode({"findings": findings}, number=True)}],
)

OpenAI (Codex / GPT). Same pattern: primer in AGENTS.md for Codex, or in the system/instructions field for the API; encode tool outputs and inter-agent messages. GPT models scored identically to Claude on the comprehension suite — including the n-field fix for positional access.

MCP servers. Return ESON as the tool result's text content for record-heavy tools (search hits, scan findings, query rows); the client agent carries the primer. Keep the negotiation rule: only emit ESON when the client declared eson/1, else compact JSON (NEGOTIATION.md).

Custom harnesses (LangGraph, CrewAI, hand-rolled). Encode on every model-bound edge, decode in program code at router/aggregator nodes (tryDecode → compact-JSON fallback). Aggregate counts and filters in the node, never in the model — pass results, not raw rows to count.

Repo layout

  • SPEC.md — normative spec, v1.1 (!eson/1 wire format)
  • PROFILE.md — the Honey Wire Profile: rules W1–W8 for any agent pipe
  • NEGOTIATION.md — encoding negotiation and fallback
  • PRIMER.md — canonical model primer; part of the wire contract
  • vectors/vectors.json — conformance vectors; an implementation conforms iff it passes them (vectors/generate.js regenerates)
  • js/, py/ — reference implementations + tests
  • bench/ — deterministic token benchmarks with committed results: FORMATS.md (npm run bench:formats), PRIMER-COST.md (npm run bench:primer), and the live end-to-end WIRE.md (node bench/wire.mjs, needs ANTHROPIC_API_KEY: primer + dispatch line → 100% valid ESON, −24% wire tokens at 100% recovery); the model-comprehension methodology lives in the honey benchmark suite

Provenance

Extracted from Honey, where ESON is the Lever-3 (agent-to-agent) format, validated against JSON/columnar/TOON on token efficiency and model comprehension across Claude and GPT model families.

MIT © Green-PT

About

Honey ESON (Efficient Structured Object Notation) — compact, lossless wire encoding for agent-to-agent structured payloads. Spec, JS + Python reference implementations, conformance vectors, LLM primer, benchmarks. MIT.

Topics

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages