Skip to content

feat: send the model server instructions on how to escalate between tools - #186

Open
karaposu wants to merge 2 commits into
brightdata:mainfrom
karaposu:fix/server-instructions
Open

karaposu wants to merge 2 commits into
brightdata:mainfrom
karaposu:fix/server-instructions

Conversation

@karaposu

@karaposu karaposu commented Sep 15, 2026 •

Copy link
Copy Markdown

Why

MCP lets a server send a short instructions string in the initialize result. OpenAI's Apps SDK documentation states that ChatGPT and Codex use it alongside tool metadata, asks that the most important details sit in the first 512 characters, and asks that it not repeat tool descriptions. This server sends none. The guidance that would help most, which tool to reach for first and what a slow collection job means, exists only in the web_scraping_strategy prompt, an opt-in MCP feature ChatGPT does not support.

What changes

Two commits.

  1. Shared test helper. Identical to the first commit of fix: complete and correct every tool's MCP annotations for plugin review #184; drops on rebase once that lands.

  2. instructions.js, built per session from the tools that session actually registers, so no clause names a tool the model cannot see. Clauses carry a priority and are emitted high-first.

    • ladder (high): escalate only as far as the page needs. A web_data_* tool when one matches the URL, else scrape_as_markdown, and scraping_browser_* only when the page needs JavaScript or interaction. In the default five-tool mode: search_engine to find pages and scrape_as_markdown to read them.
    • job_timing (high, only when a job-starting tool is registered): these start billed collection jobs; most answer in seconds; one that is unusually slow is more often a bad target than a busy server, so report it rather than retrying. With the marketplace tools present it adds that query_dataset's snapshot id can be redeemed with collect_dataset.
    • usage_limit (normal): if a tool reports the free-tier limit, follow the instructions in that message rather than retrying.

    Sizes: pro mode 493 characters with the high-priority clauses ending at 384; with marketplace tools 593, ending at 484; default mode 219. All inside the 512 window.

    discover is deliberately excluded from every clause. Its API is retired: POST /discover answers 410 Gone ("Discover API is no longer available") for every query, so it starts nothing and must not be recommended. It is still registered as one of the five default-mode tools. feat!: remove the discover tool and Discover API integration #164 removes it, and that 410 is the reason feat!: remove the discover tool and Discover API integration #164 should land before any submission.

    The timing clause is measured rather than assumed. Against the deployed server, web_data_* tools answered in 4 to 16 seconds across five calls, and a request for a delisted product took 55 seconds before coming back empty. An earlier figure of 30 to 120 seconds came from the marketplace /datasets/filter flow, which web_data_* does not use.

Verification

npm test: 32 pass. The tests pin the rules by clause id rather than wording; assert that every high-priority clause falls inside the first 512 characters in every mode; assert that no mode is told about a tool family it does not register; assert that nothing recommends discover; and assert that the instructions and the web_scraping_strategy prompt agree on tier order.

Not verified

  • That ChatGPT applies the instructions. OpenAI documents it; it has not been observed in a session. One distinctive sentence and one question in developer mode would settle it.
  • Whether the hosted server, a separate build, forwards instructions once released.

Groundwork for validating what this server advertises to MCP clients.
Checking the whole tool catalog from a test means the spawned server has
to register exactly the tools it would in production, and today every
test that starts one re-implements the spawn and spreads the whole of
process.env into the child. A stray TOOLS or GROUPS in a developer's
shell therefore silently shrinks the catalog under the test.

test/helpers/spawn_server.js owns that contract instead:

  - an explicit inherit list. The MCP SDK already supplies HOME, PATH,
    SHELL, TERM and USER, so this carries the network plumbing a
    developer behind a proxy depends on, and nothing else.
  - network: 'none', which points the child at a refused proxy for tests
    that must never reach the real API.
  - captured, drained stderr, so a test can assert on what the server
    reported. Draining is not optional: a piped stderr nobody reads
    blocks the child once the pipe buffer fills, and this server logs
    every tool call.
  - a deployed-server variant that sends the token in an Authorization
    header rather than in a URL, and scrubs the exact value from
    anything it throws, so a transport error cannot put a credential in
    a terminal or a CI log.

The npm script now names test/*.test.js. Bare `node --test` treats every
.js file under test/ as a test, which would run this helper as an empty
passing test. The pattern stays single-level deliberately: under sh
without globstar, test/**/*.test.js matches only files below test/ and
would silently drop the entire suite.
Why this exists
---------------

An MCP server may return an `instructions` string when a client connects.
OpenAI documents that ChatGPT and Codex use it alongside tool metadata, and
asks that the most important details sit in the first 512 characters. This
server has never sent one.

The guidance does exist, and it is good: two MCP prompts explain the order to
try tools in. But prompts are opt-in -- a client must list them and a user
must invoke one -- and OpenAI's plugin platform has no prompt concept at all.
So in ChatGPT that guidance is invisible: a request like "compare prices on
Amazon" reaches a model holding 74 tool descriptions and no rule about which
to try first.

What the text says, and what it deliberately does not
-----------------------------------------------------

It carries only what no single tool description can:

  - the escalation ladder between tools;
  - what to do when a collection job outlives the client's request timeout;
  - that the free-tier limit message is an instruction to follow, not an
    error to retry.

Everything else stays where it already is. Each web_data_* description
already explains that it can be a cache lookup and more reliable than
scraping -- the specific, honest version of a claim a global rule about 74
heterogeneous tools could only approximate, and where a maintainer will look
when one tool's trade-off changes.

The text makes no cost claim. "Prefer the cheapest tool" would have been
wrong in both directions: a web_data_* call starts a billed job taking tens
of seconds, while scrape_as_markdown answers in a few, and nobody has
measured credits against latency. The rule is about escalating only as far as
a page requires.

The job clause matters most and was the easiest to get wrong. An earlier
draft told the model these jobs take 30-120 seconds and would probably be
cancelled by the client first. Measured against the deployed server on
2026-09-10, the web_data_* tools answered in 4-16 seconds across five calls;
none came close to the 60-second cancellation. The 30-120 figure had been
measured on the marketplace /datasets/filter flow, which no web_data_* tool
uses.

What the measurements did show is more useful: a request for a delisted
product took 55 seconds before coming back empty. Slowness signals a bad
target, not a busy server. So the clause says that, and names the cheap
recovery where one exists -- redeeming query_dataset's snapshot id with
collect_dataset -- rather than warning about a timeout that does not
routinely happen.

Derived from what the session actually registers
------------------------------------------------

Tools are gated by PRO_MODE, TOOLS and GROUPS, and the default mode
registers five tools with no browser and no dataset family. A rule naming a
tool the model cannot see is worse than no rule, because the model looks for
something that is not there. So clauses declare the capability they need and
the text is built from the same gate addTool applies, using tool_groups.js's
catalogue as the source of names -- verified to hold exactly the 74 tools pro
mode registers.

The result differs by mode. Pro mode gets the three-tier ladder. The default
mode gets the ordering that actually applies to it: search_engine to find
pages, scrape_as_markdown to read them.

One tool is deliberately absent from all of it. `discover` is registered in
the default mode -- one of only five tools there -- but its API has been
retired: POST /discover answers 410 "Discover API is no longer available" for
every query, verified 2026-09-10. It therefore starts no job and cannot be
recommended, so it shapes no clause. Because it was the default mode's only
job-starting tool, that mode now correctly gets no job clause at all. The
tool itself should be removed from the server; upstream PR brightdata#164 already
proposes exactly that, and this is the evidence for why it should land before
any submission.

Placement, not just length
--------------------------

"Most important details in the first 512 characters" is an ordering
requirement, not a size cap: text can sit under any total and still bury its
best line. Clauses carry an explicit priority, the builder emits them
high-first, and the tests assert that the high-priority clauses land inside
the window rather than that the whole string is short. Measured: 480
characters in pro mode with the high-priority rules ending at 384, and 593
with the marketplace tools present, ending at 484.

That constraint did real work. The first draft of the two high-priority
clauses ran to 561 characters together and pushed past the window; the test
caught it and the prose was tightened.

Tests
-----

  test/instructions-builder.test.js  the rules, pinned by clause id rather
                                     than by prose, so a deleted rule fails
                                     and a rewording does not. Covers all
                                     four modes including the marketplace
                                     tools that arrive with brightdata#177.
  test/instructions.test.js          what a client observes through
                                     initialize: presence, the priority
                                     window, which clauses each mode gets,
                                     and that no mode is told about tools it
                                     does not register.

The existing web_scraping_strategy prompt is untouched. A test asserts the
two agree on tier order rather than forcing them to share a string: they
serve different readers and should read differently.

tools/list is byte-identical to before -- verified, 74 tools either way.
This is transport metadata and it moves no tool.

What is NOT verified
--------------------

  - That the text reaches the model in ChatGPT. OpenAI documents that it is
    used; this has not been observed. ChatGPT developer mode can settle it
    with a canary sentence and no submission, and that is the next step.
  - Whether the hosted deployment forwards this option. It reports a
    different name and version from this package, so it likely constructs
    its own server; if so, someone with access to that wrapper has to wire
    it before any of this reaches ChatGPT.
  - Which mode a ChatGPT session is served, which decides whether the
    three-tier ladder or the default-mode ordering is the one that matters.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant