fix(sessions): reuse one Vertex AI client per event loop in the session and memory bank services - #7354
fix(sessions): reuse one Vertex AI client per event loop in the session and memory bank services#7354vishal-bulbule wants to merge 1 commit into
Conversation
…on and memory bank services VertexAiSessionService and VertexAiMemoryBankService built a new vertexai.Client for every call. Each client keeps resources that are not released after the call (on Agent Engine, the mTLS SSL contexts), so memory grows with request count. Closing more of each client does not help; building fewer clients does. Build the client once per event loop with per_loop_value, the helper Gemini.api_client already uses, so a loop reuses its client and every other loop gets its own. The session service no longer wraps each call in async with, which would close the shared client after one use. A cached client made a used service fail deepcopy and pickle on its cache lock. _PerLoopCache now copies and pickles as an empty cache, so the copy builds its own client. This also fixes the same failure for a Gemini instance whose api_client has been read. Fixes google#7353
|
Agent Engine results for this change. With a client certificate (the case I could not test): @gorankl ran the same per-loop client cache on Agent Engine with a client cert present, google-adk 2.10.0 and google-genai 2.25.0, applied to the session service only (comment on #7353). Over 500 requests at 2 in flight:
Without a client certificate, on my own Agent Engine runtime in us-central1 (
Without a certificate the discarded contexts are collected and RSS grew at the same rate on both builds, so that runtime does not show the leak; it does show the client churn going from three clients per request to one. |
Link to Issue or Description of Change
1. Link to an existing issue (if applicable):
Problem:
VertexAiSessionService._get_api_client()builds a newvertexai.Clientfor every session call, andVertexAiMemoryBankService._get_api_client()does the same for every memory call. Each client allocates resources (on Agent Engine, the mTLS SSL contexts from #7353) that are not released after the call, so memory grows with request count.Measured locally with google-genai 2.25.0 and google-cloud-aiplatform 1.165.1, no client certificate, RSS after
gc.collect()over 1000 calls:async with ... .aio(current)Closing more of each client does not help; building fewer clients does.
Solution:
Build the client once per event loop with the existing
per_loop_valuehelper inutils/_event_loop_cache.py, the same helperGemini.api_clientuses. An async client belongs to the loop that opened it, so a single client for the service's lifetime would break the sync runner entry points and thread-pool servers, which run calls on different loops. Per-loop caching reuses the client within a loop and gives every other loop its own._get_api_client()now returns the cached client; the construction moved to_build_api_client()._api_client_http_options_override()still applies.async with, since that would close the shared client after the first call.Gemini.api_clienthas, so a long-lived server loop keeps one client for the process._get_api_client()to return a new client per call no longer has it closed byasync with;_build_api_client()is the method to override now.Caching a client in the service's
__dict__would make a used service failcopy.deepcopyandpicklewithcannot pickle '_thread.lock' object, which works on main._PerLoopCache.__reduce__now copies and pickles as an empty cache, so the copy builds its own client on first read. This also fixes the same failure for aGeminiinstance whoseapi_clienthas been read, which has been there since that property moved to per-loop caching.Testing Plan
Unit Tests:
New tests:
_PerLoopCache: a used owner deep-copies and pickles without its cached values (2 tests, both fail on main)Two existing session tests assumed a client per call and were updated: the pagination mock now stays open until it is closed, like the real client, and still fails if the client is closed mid-iteration; the remote-failure append test patches
_get_api_clientwith a client instead of a context manager.The 3 failures are not from this change:
test_import_loading.py::test_entry_point_loads_only_allowlisted_packages[agent]fails the same way on main (f44d512) in my environment, andtest_live_tool_shutdown.py::test_handoff_stops_streaming_tools_before_the_transfer_delayandtest_mcp_session_manager.py::TestMCPSessionManager::test_is_session_disconnected_without_streamsare timing tests that failed only under the parallel run and pass on their own, on this branch and on main.integrations/livekitis skipped because it hangs in my environment on main too.pre-commit runon the changed files is clean. mypy reports no new errors (one existing error atvertex_ai_memory_bank_service.py:1072is on main too).Manual End-to-End (E2E) Tests:
On the real services with a mocked client constructor: 10 calls in one loop build 1 client, a second loop builds its own, and 4 threads with their own loops build 4. After use,
VertexAiSessionService,VertexAiMemoryBankServiceandGeminiall deep-copy and pickle.On Agent Engine with a client certificate, the same per-loop client cache kept live
SSLContexts flat (17 to 25 over 500 requests, against about 950 after 100 requests on the stock service), as tested by the reporter of #7353; details and my own Agent Engine run without a certificate are in this comment.Checklist