Summary
GoogleModelSettings.google_cached_content is documented but unusable as
implemented. When the setting is provided, pydantic-AI still includes
system_instruction, tools, and tool_config in the outgoing
GenerateContentConfig — the Vertex API rejects that combination with
400 INVALID_ARGUMENT:
Tool config, tools and system instruction should not be set in the
request when using cached content.
The setting therefore can't drive a successful request without a local
workaround (subclassing GoogleModel and stripping the three fields
post-build).
Reproduction
Pre-create a Vertex cachedContents resource that carries a
system_instruction and at least one tools declaration (any model
that supports caching; minimum ~4096 tokens of content per Vertex's
floor). Then:
from pydantic_ai import Agent
from pydantic_ai.models.google import GoogleModel, GoogleModelSettings
from pydantic_ai.providers.google import GoogleProvider
agent = Agent[None, str](
GoogleModel("gemini-2.5-pro", provider=GoogleProvider(...)),
)
@agent.tool_plain
def echo(text: str) -> str:
return text
result = agent.run_sync(
"say hi",
model_settings=GoogleModelSettings(
google_cached_content="projects/<p>/locations/global/cachedContents/<id>",
),
)
# google.genai.errors.ClientError: 400 INVALID_ARGUMENT.
# Tool config, tools and system instruction should not be set in the
# request when using cached content.
The same shape fails without a tool registered on the agent if the
underlying messages produce a non-None system_instruction.
Root cause
In pydantic_ai/models/google.py::GoogleModel._build_content_and_config,
the request config is built with all of system_instruction, tools,
tool_config, AND cached_content populated:
config = GenerateContentConfigDict(
http_options=http_options,
system_instruction=system_instruction,
...
cached_content=model_settings.get('google_cached_content'),
tools=cast(ToolListUnionDict, tools),
tool_config=tool_config,
...
)
Per the Vertex contract, when cached_content is set those three fields
must be absent — the cache resource owns them.
Expected behavior
When model_settings.google_cached_content is set, the outgoing config
should omit system_instruction, tools, and tool_config. Roughly:
cached_content = model_settings.get('google_cached_content')
config = GenerateContentConfigDict(
http_options=http_options,
cached_content=cached_content,
temperature=model_settings.get('temperature'),
# ... other non-cache-owned fields ...
)
if not cached_content:
config['system_instruction'] = system_instruction
config['tools'] = cast(ToolListUnionDict, tools)
config['tool_config'] = tool_config
A regression test that exercises a real agent.run_sync against a Vertex
endpoint with a pre-created cache would have caught this.
Local workaround
Subclass GoogleModel, override _build_content_and_config to call
super() and then pop system_instruction, tools, tool_config
from the returned config when cached_content is set.
Context
Summary
GoogleModelSettings.google_cached_contentis documented but unusable asimplemented. When the setting is provided, pydantic-AI still includes
system_instruction,tools, andtool_configin the outgoingGenerateContentConfig— the Vertex API rejects that combination with400 INVALID_ARGUMENT:The setting therefore can't drive a successful request without a local
workaround (subclassing
GoogleModeland stripping the three fieldspost-build).
Reproduction
Pre-create a Vertex
cachedContentsresource that carries asystem_instructionand at least onetoolsdeclaration (any modelthat supports caching; minimum ~4096 tokens of content per Vertex's
floor). Then:
The same shape fails without a tool registered on the agent if the
underlying messages produce a non-None
system_instruction.Root cause
In
pydantic_ai/models/google.py::GoogleModel._build_content_and_config,the request config is built with all of
system_instruction,tools,tool_config, ANDcached_contentpopulated:Per the Vertex contract, when
cached_contentis set those three fieldsmust be absent — the cache resource owns them.
Expected behavior
When
model_settings.google_cached_contentis set, the outgoing configshould omit
system_instruction,tools, andtool_config. Roughly:A regression test that exercises a real
agent.run_syncagainst a Vertexendpoint with a pre-created cache would have caught this.
Local workaround
Subclass
GoogleModel, override_build_content_and_configto callsuper()and then popsystem_instruction,tools,tool_configfrom the returned config when
cached_contentis set.Context
GoogleModelSettings.google_cached_contentto passcached_content#2832 (merged 2025-09-08); zero follow-upcomments suggest the end-to-end happy path wasn't exercised.
cachedContentsdocs:https://cloud.google.com/vertex-ai/generative-ai/docs/context-cache/context-cache-overview
the constraint the same way.