Problem
The godot-ai attach bridge deliberately has no post-dispatch read deadline. A bridge-side timeout could turn a successfully completed mutation into an ambiguous result, and retrying it could duplicate the mutation.
This preserves the core safety invariant, but leaves one recovery gap: if the downstream Streamable HTTP response stream is lost while the backend process remains alive with the same instance_id, the bridge can wait indefinitely. Its identity monitor continues to see a healthy backend, while the MCP SDK may have no response stream from which to receive the result.
This follow-up comes from #822 and the PR A implementation.
What PR A established
Phase 0 testing found:
- FastMCP 3.0 and the current 3.4 release both expose an
event_store argument on FastMCP.http_app().
- With the current FastMCP 3.4 stack, a TCP cut-and-resume test successfully reconnected and recovered the response when event IDs were present.
- The equivalent end-to-end flow hangs on the supported FastMCP 3.0 floor, even though the reconnecting GET request occurs.
- Enabling the event store only on newer dependency versions would give the attach bridge version-dependent recovery semantics.
For that reason, PR A did not enable an event store.
Two distinct failure cases
These should remain separate in the implementation and documentation:
-
The response stream is lost while the server continues processing.
This should be recoverable through MCP event IDs and response-stream resumption.
-
The server-side request task dies without producing a response.
An event store cannot recover a response that was never produced. Without introducing an unsafe post-dispatch deadline, recovery remains upstream cancellation. Cancellation must clean up the bridge’s local tasks and must never replay the call.
Proposed work
- Reduce the FastMCP 3.0 hang to a minimal reproducible case and determine whether the cause is in FastMCP, the MCP SDK, or the bridge’s transport construction.
- Implement a compatibility shim, upstream fix, or explicitly justified dependency-floor change so resumption behaves consistently across every supported FastMCP version.
- Add a backend-level, in-memory event store with bounded retention:
- TTL bounded;
- size bounded;
- shared by the backend application;
- no persistent state across backend restarts.
- Wire it through every Streamable HTTP construction path:
- normal CLI startup in
src/godot_ai/__init__.py;
- reload/ASGI startup in
src/godot_ai/asgi.py.
- Keep
read=None for dispatched operations. Do not implement bridge-side read deadlines or replay after ambiguous dispatch.
- Preserve and document the server-task-death limitation.
Required tests
- A real TCP proxy severs the downstream response stream after a
tools/call has been dispatched.
- The backend completes the operation exactly once.
- The client reconnects using the event ID and receives the original result.
- The bridge does not invoke the tool a second time.
- The test passes at both the FastMCP dependency floor and the current pinned/maximum-supported stack.
- Both normal CLI startup and ASGI/reload startup use the event store.
- Cancelling a call whose server task will never produce a response cleans up all bridge tasks and does not replay the operation.
Acceptance criteria
- Same-instance network stream loss is transparently resumable across every supported dependency version.
- No mutation is executed more than once.
- No post-dispatch read deadline is introduced.
- A request task that dies without producing a response remains cancellation-recoverable and is explicitly documented as distinct from stream resumption.
Problem
The
godot-ai attachbridge deliberately has no post-dispatch read deadline. A bridge-side timeout could turn a successfully completed mutation into an ambiguous result, and retrying it could duplicate the mutation.This preserves the core safety invariant, but leaves one recovery gap: if the downstream Streamable HTTP response stream is lost while the backend process remains alive with the same
instance_id, the bridge can wait indefinitely. Its identity monitor continues to see a healthy backend, while the MCP SDK may have no response stream from which to receive the result.This follow-up comes from #822 and the PR A implementation.
What PR A established
Phase 0 testing found:
event_storeargument onFastMCP.http_app().For that reason, PR A did not enable an event store.
Two distinct failure cases
These should remain separate in the implementation and documentation:
The response stream is lost while the server continues processing.
This should be recoverable through MCP event IDs and response-stream resumption.
The server-side request task dies without producing a response.
An event store cannot recover a response that was never produced. Without introducing an unsafe post-dispatch deadline, recovery remains upstream cancellation. Cancellation must clean up the bridge’s local tasks and must never replay the call.
Proposed work
src/godot_ai/__init__.py;src/godot_ai/asgi.py.read=Nonefor dispatched operations. Do not implement bridge-side read deadlines or replay after ambiguous dispatch.Required tests
tools/callhas been dispatched.Acceptance criteria