Skip to content

fix(sessions): reuse one Vertex AI client per event loop in the session and memory bank services - #7354

Open
vishal-bulbule wants to merge 1 commit into
google:mainfrom
vishal-bulbule:fix/vertex-session-client-reuse
Open

vishal-bulbule wants to merge 1 commit into
google:mainfrom
vishal-bulbule:fix/vertex-session-client-reuse

Conversation

@vishal-bulbule

@vishal-bulbule vishal-bulbule commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Link to Issue or Description of Change

1. Link to an existing issue (if applicable):

Problem:

VertexAiSessionService._get_api_client() builds a new vertexai.Client for every session call, and VertexAiMemoryBankService._get_api_client() does the same for every memory call. Each client allocates resources (on Agent Engine, the mTLS SSL contexts from #7353) that are not released after the call, so memory grows with request count.

Measured locally with google-genai 2.25.0 and google-cloud-aiplatform 1.165.1, no client certificate, RSS after gc.collect() over 1000 calls:

Client strategy RSS growth
New client per call, async with ... .aio (current) +47 to +53 MB (about 50 KB per call)
New client per call, sync httpx client also closed +31 to +61 MB
One client reused for all calls +0 MB
One client per event loop, new loop on every call +2 to +6 MB

Closing more of each client does not help; building fewer clients does.

Solution:

Build the client once per event loop with the existing per_loop_value helper in utils/_event_loop_cache.py, the same helper Gemini.api_client uses. An async client belongs to the loop that opened it, so a single client for the service's lifetime would break the sync runner entry points and thread-pool servers, which run calls on different loops. Per-loop caching reuses the client within a loop and gives every other loop its own.

  • _get_api_client() now returns the cached client; the construction moved to _build_api_client(). _api_client_http_options_override() still applies.
  • The session service no longer wraps each call in async with, since that would close the shared client after the first call.
  • The memory bank docstring said the client had to be built per request for event loop reasons. The per-loop cache covers that case.
  • Clients are no longer closed after each call. A loop's client is released when that loop is collected, the same lifetime Gemini.api_client has, so a long-lived server loop keeps one client for the process.
  • A subclass that overrides _get_api_client() to return a new client per call no longer has it closed by async with; _build_api_client() is the method to override now.

Caching a client in the service's __dict__ would make a used service fail copy.deepcopy and pickle with cannot pickle '_thread.lock' object, which works on main. _PerLoopCache.__reduce__ now copies and pickles as an empty cache, so the copy builds its own client on first read. This also fixes the same failure for a Gemini instance whose api_client has been read, which has been there since that property moved to per-loop caching.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally.

New tests:

  • session and memory bank services: one loop reuses one client; a second loop builds its own (4 tests, all fail on main)
  • _PerLoopCache: a used owner deep-copies and pickles without its cached values (2 tests, both fail on main)

Two existing session tests assumed a client per call and were updated: the pagination mock now stays open until it is closed, like the real client, and still fails if the client is closed mid-iteration; the remote-failure append test patches _get_api_client with a client instead of a context manager.

pytest tests/unittests/sessions tests/unittests/memory tests/unittests/utils/test_event_loop_cache.py
674 passed, 1 xfailed

pytest tests/unittests -n auto --ignore=tests/unittests/integrations/livekit
3 failed, 16525 passed, 84 skipped, 25 xfailed, 2 xpassed

The 3 failures are not from this change: test_import_loading.py::test_entry_point_loads_only_allowlisted_packages[agent] fails the same way on main (f44d512) in my environment, and test_live_tool_shutdown.py::test_handoff_stops_streaming_tools_before_the_transfer_delay and test_mcp_session_manager.py::TestMCPSessionManager::test_is_session_disconnected_without_streams are timing tests that failed only under the parallel run and pass on their own, on this branch and on main. integrations/livekit is skipped because it hangs in my environment on main too.

pre-commit run on the changed files is clean. mypy reports no new errors (one existing error at vertex_ai_memory_bank_service.py:1072 is on main too).

Manual End-to-End (E2E) Tests:

On the real services with a mocked client constructor: 10 calls in one loop build 1 client, a second loop builds its own, and 4 threads with their own loops build 4. After use, VertexAiSessionService, VertexAiMemoryBankService and Gemini all deep-copy and pickle.

On Agent Engine with a client certificate, the same per-loop client cache kept live SSLContexts flat (17 to 25 over 500 requests, against about 950 after 100 requests on the stock service), as tested by the reporter of #7353; details and my own Agent Engine run without a certificate are in this comment.

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

…on and memory bank services

VertexAiSessionService and VertexAiMemoryBankService built a new
vertexai.Client for every call. Each client keeps resources that are not
released after the call (on Agent Engine, the mTLS SSL contexts), so
memory grows with request count. Closing more of each client does not
help; building fewer clients does.

Build the client once per event loop with per_loop_value, the helper
Gemini.api_client already uses, so a loop reuses its client and every
other loop gets its own. The session service no longer wraps each call
in async with, which would close the shared client after one use.

A cached client made a used service fail deepcopy and pickle on its
cache lock. _PerLoopCache now copies and pickles as an empty cache, so
the copy builds its own client. This also fixes the same failure for a
Gemini instance whose api_client has been read.

Fixes google#7353
@vishal-bulbule

Copy link
Copy Markdown
Contributor Author

Agent Engine results for this change.

With a client certificate (the case I could not test): @gorankl ran the same per-loop client cache on Agent Engine with a client cert present, google-adk 2.10.0 and google-genai 2.25.0, applied to the session service only (comment on #7353). Over 500 requests at 2 in flight:

Live SSLContexts RSS
stock VertexAiSessionService about 950 after about 100 requests about +7 MB per request
per-loop client cache 17 to 25, flat from request 100 to 500 +58 MB over the run, stable while idle

Without a client certificate, on my own Agent Engine runtime in us-central1 (should_use_client_cert() is False there), two engines that differ only in the ADK wheel, main f44d512 and this branch, called through AdkApp.stream_query:

SSLContexts created per request alive (max)
main f44d512 9.0 in each of the 4 workers 3
this PR 3.0 in each of the 4 workers 3

Without a certificate the discarded contexts are collected and RSS grew at the same rate on both builds, so that runtime does not show the leak; it does show the client churn going from three clients per request to one. AdkApp.stream_query runs each request in a new event loop, which is why it is one client per request there rather than one per process.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

VertexAiSessionService creates a new vertexai.Client per call; SSL contexts accumulate on Agent Engine

2 participants