TokenCountingHandler don't work for OpenAIAgent Function Call #19303
Replies: 3 comments 6 replies
|
TokenCountingHandler only tracks tokens for LLM and Embedding events that are emitted through the CallbackManager system. When you use QueryEngineTool, it interacts with an LLM and emits these events, so TokenCountingHandler works as expected. However, FunctionTools usually just wrap Python functions, which don't necessarily call an LLM or emit any LLM events—so TokenCountingHandler doesn't see anything to count and remains empty. If your FunctionTool does internally call an LLM (and uses a properly configured CallbackManager), token usage will be tracked; otherwise, it won't be counted by TokenCountingHandler see code reference. If you need to track token usage for FunctionTools that do invoke LLMs, you can implement custom logic to access the raw OpenAI response and extract token usage, as suggested by maintainers see discussion. For pure Python or non-LLM tools, there simply aren't any tokens to count from the LLM's perspective, so TokenCountingHandler won't show any values. To reply, just mention @dosu. Docs are dead. Just use Dosu. |
|
Thanks for highlighting this issue. Accurate token counting during function calls is essential for monitoring usage and controlling API costs. This is especially valuable for food-related AI applications, where recipe generators, restaurant menu assistants, and nutrition planners rely on precise token tracking to deliver efficient and cost-effective user experiences. |
|
The important distinction is where the callback manager is attached.
So if the counter stays at zero on the function-tool path, I would first check that the handler is attached to the agent's own LLM, not only to the query engine or index: from llama_index.core.callbacks import CallbackManager, TokenCountingHandler
token_counter = TokenCountingHandler()
callback_manager = CallbackManager([token_counter])
llm = build_the_llm_used_by_the_agent()
llm.callback_manager = callback_manager
agent = build_the_agent_with_tools(
tools=[my_function_tool],
llm=llm,
)
response = agent.chat("...")
print(token_counter.prompt_llm_token_count)
print(token_counter.completion_llm_token_count)
print(token_counter.total_llm_token_count)If the function tool calls another model internally, that inner model also needs the same callback manager, or its own counter. If the function tool is pure Python and does not call a model or an embedder, there are no additional model tokens for For cost accounting, I would track tool-call counts/latency separately from model tokens. They are different units, and trying to force Python tool execution into |
Uh oh!
There was an error while loading. Please reload this page.
Hi,
The TokenCountingHandler works fine when I do a question to a OpenAIAgent that has a QueryEngineTool inside and use it about indexed documents. But if I do a question to the same OpenAIAgent and it use only the FunctionTools, the TokenCountingHandler don't have any value. How should I use the TokenCountingHandler in this context? Thank you
All reactions