How to Evaluate Agent Tool Use in Production
Measure agent tool selection accuracy, parameter correctness, and workflow success. A practical framework for evaluating agent tool use in production.
By Pulkit Verma, Founder & CEO, WeaveAI
Research and drafting assisted by WeaveAI Cite.
Agents fail most often at the tool boundary—picking the wrong function, hallucinating parameters, or misinterpreting results. A structured evaluation framework catches these issues during development and monitors them in production.
What Makes Tool Use Evaluation Different from General LLM Evaluation
Tool use introduces structured interfaces and side effects that text generation alone does not. An agent can produce fluent text but still break a workflow by calling delete_record instead of update_record, or by passing a string where an integer is required.
General LLM evaluation focuses on semantic similarity, coherence, and factual accuracy. Tool use evaluation adds three layers: did the agent choose the right function from the available set, did it construct valid arguments, and did the execution produce the correct downstream effect? Each layer can fail independently.
Side effects make rollback expensive. A poorly evaluated text summarizer produces bad summaries; a poorly evaluated agent might send duplicate emails, charge a customer twice, or overwrite production data. Evaluation must catch errors before tools execute, not after.
Step-by-Step Framework for Evaluating Agent Tool Use
1. Define Ground Truth for Tool Selection
Build a test set of user intents paired with the correct tool and parameters. Each example should represent a real workflow your agent will encounter. Include edge cases where two tools seem plausible but only one is correct.
For a customer support agent, one example might be: intent "refund the last order", correct tool issue_refund, correct parameter order_id extracted from conversation context. Another: "check order status" maps to get_order_details, not issue_refund.
Label at least 50 examples per tool in your agent's toolkit. Fewer than that and you will miss the boundary cases where the model hesitates between similar functions.
2. Measure Selection Accuracy
Run each test case through your agent and record which tool it selects. Selection accuracy is the percentage of cases where the chosen tool matches ground truth.
Break down accuracy by tool. An agent might reliably pick search_knowledge_base but confuse create_ticket and update_ticket. Per-tool metrics show you where prompt engineering or few-shot examples need to focus.
Track confusion matrices to see which tools the agent conflates. If cancel_subscription and pause_subscription are frequently swapped, you need clearer tool descriptions or a disambiguation step in your prompt.
3. Validate Parameter Correctness
Selection accuracy alone is insufficient. The agent must also construct valid arguments. Check two dimensions: schema compliance (does the parameter match the expected type and format) and semantic correctness (is the value appropriate for the user's intent).
Schema validation is mechanical—did the agent pass a string where a string is required, an integer where an integer is expected, and does the value satisfy any constraints like length or range? Most tool frameworks handle this automatically and return an error if validation fails.
Semantic correctness is harder. The agent might pass a syntactically valid order_id that belongs to the wrong customer, or a date_range that spans the wrong month. Compare extracted parameters against your ground truth labels. Parameter correctness is the percentage of tool calls where all arguments match expected values.
4. Measure Execution Success Rate
Even correct tool selection and valid parameters can fail if the tool returns an error or the agent misinterprets the result. Execution success rate tracks whether the tool call completes and produces the intended outcome.
Log every tool invocation with its return status. Success means the tool executed without error and the agent used the result appropriately in its next action. Failure includes exceptions, timeouts, null results the agent did not handle, and cases where the agent ignored a valid response.
Calculate success rate per tool and per workflow. A tool might work in isolation but fail when called as part of a multi-step sequence.
5. Evaluate End-to-End Workflow Completion
Tool-level metrics do not guarantee the agent solves the user's problem. Workflow completion rate measures whether the agent reaches the correct terminal state—ticket created, refund issued, question answered—across multi-turn interactions.
Define success criteria for each workflow type. For "user requests a refund," success means the agent called issue_refund with the correct order ID and confirmed the action to the user. Partial completion (agent selected the right tool but failed to confirm) counts as failure.
Run end-to-end tests on realistic scenarios with multiple turns and branching logic. Track where workflows break: wrong tool early in the sequence, correct tool but incorrect parameters, or correct execution but poor result handling.
6. Monitor Failure Modes in Production
Evaluation does not stop at deployment. Instrument your agent to log every tool call with context: user input, selected tool, parameters, execution result, and final agent response. Sample and review logs daily.
Common failure modes include: tool selection degrades when user phrasing drifts from training examples, parameter extraction breaks on unexpected input formats, tools return edge-case errors the agent was not trained to handle, and the agent hallucinates tool names not in its toolkit.
Set up alerts for sudden drops in selection accuracy or spikes in execution errors. A 10% drop in success rate over 24 hours usually means a prompt regression, a broken tool endpoint, or a shift in user behavior.
Comparison of Tool Use Evaluation Approaches
| Approach | Best For | Limitation |
|---|---|---|
| Unit testing per tool | Catching schema violations and basic selection errors | Does not evaluate multi-step reasoning or context-dependent parameter extraction |
| Synthetic test sets with labeled ground truth | Systematic measurement of selection and parameter accuracy across known scenarios | Requires manual labeling effort; may not cover real-world edge cases |
| Human review of production logs | Surfacing novel failure modes and ambiguous cases | Does not scale; introduces lag between failure and detection |
| Automated replay of production traces | Regression testing after prompt or model changes | Only catches errors that recur; misses new failure modes |
| A/B testing tool call success rates | Comparing prompt variants or model versions in live traffic | Requires sufficient volume; some failures have delayed or invisible effects |
No single approach is sufficient. Effective evaluation layers unit tests for fast feedback, labeled test sets for systematic coverage, and production monitoring for real-world drift.
How to Build a Labeled Evaluation Set
Start by logging real user interactions and sampling diverse cases. Prioritize examples where the agent succeeded and where it failed—both teach you about boundary conditions.
For each example, label the correct tool, the correct parameters, and the expected outcome. If multiple tools could plausibly satisfy the intent, document why one is preferred. These annotations become your prompt refinement guide.
Aim for balanced coverage: include common cases (80% of traffic) and edge cases (the remaining 20% that cause 80% of failures). Update the set as you add tools or discover new failure modes.
Use a simple annotation format. A JSON file with user_input, correct_tool, correct_parameters, and expected_outcome fields is enough. Version it alongside your agent code so evaluation evolves with the system.
What to Do When Tool Use Accuracy Is Low
If selection accuracy falls below 90%, your agent does not reliably choose the right function. Start by improving tool descriptions—make them more distinct and include examples of when to use each tool versus similar alternatives.
Add few-shot examples to your prompt showing correct tool selection for ambiguous cases. If the agent confuses update_ticket and close_ticket, include an example of each with clear reasoning.
If parameter correctness is low, the agent struggles with extraction or context tracking. Simplify parameter schemas where possible—fewer required fields mean fewer opportunities for error. Add validation hints to tool descriptions, like "order_id is a 10-digit integer found in the conversation history."
When execution success rate drops, investigate whether tools are returning unexpected errors or the agent is mishandling valid responses. Add error handling instructions to your prompt and test tool behavior on edge-case inputs.
Integrating Evaluation into Your Development Workflow
Run your labeled test set on every prompt change and before every deployment. Treat tool use evaluation like unit tests—a commit that drops selection accuracy by more than 2 percentage points should not merge.
Track metrics over time in a dashboard. Plot selection accuracy, parameter correctness, and workflow completion rate per tool and per week. Trends reveal whether changes improve the agent or introduce regressions.
Automate evaluation in CI/CD. A script that runs test cases, compares results to ground truth, and fails the build if accuracy drops below a threshold prevents broken agents from reaching production.
Review failed cases weekly. Categorize failures by type—wrong tool, bad parameter, execution error—and prioritize fixes based on frequency and user impact. Feed the most common failures back into your test set.
Frequently Asked Questions
What metrics matter most when evaluating agent tool use?
Selection accuracy (percentage of cases where the agent picks the correct tool), parameter correctness (percentage of tool calls with valid and semantically appropriate arguments), and workflow completion rate (percentage of multi-step interactions that reach the intended outcome) form the core metric set. Selection accuracy catches reasoning errors, parameter correctness catches extraction and context failures, and workflow completion measures end-to-end reliability. Track all three; optimizing one at the expense of others produces agents that pass unit tests but fail in production.
How large should my tool use evaluation set be?
A useful evaluation set contains at least 50 examples per tool in your agent's toolkit, balanced between common cases and edge cases. For an agent with five tools, that means at least 250 labeled examples. Fewer than 30 examples per tool will miss boundary conditions where the model hesitates between similar functions. More than 100 per tool offers diminishing returns unless you are fine-tuning a model specifically for tool selection. Update the set continuously as you discover new failure modes in production logs.
How do I evaluate tool use when tools have side effects?
Separate evaluation into pre-execution validation and post-execution review. Before the tool runs, check selection accuracy and parameter correctness against ground truth—this catches most errors without triggering side effects. For execution success, use sandbox environments or mock tool implementations during testing. In production, log every tool call with full context and sample a subset for human review. Set up rollback procedures for high-risk tools like payment or data deletion functions, and require human-in-the-loop confirmation for those actions until your agent proves reliable.
Build Agents That Work After the Demo
Evaluating agent tool use is not a one-time gate—it is a feedback loop that runs from development through production. The agents that stay reliable are the ones instrumented to surface failures early and updated as real-world usage teaches you where the model breaks.
WeaveAI builds AI workflow agents with structured evaluation built in from day one. If you are deploying agents that call tools in production and need a framework that catches failures before users do, we can help.
Frequently asked questions
What metrics matter most when evaluating agent tool use?
Selection accuracy (percentage of cases where the agent picks the correct tool), parameter correctness (percentage of tool calls with valid and semantically appropriate arguments), and workflow completion rate (percentage of multi-step interactions that reach the intended outcome) form the core metric set. Selection accuracy catches reasoning errors, parameter correctness catches extraction and context failures, and workflow completion measures end-to-end reliability. Track all three; optimizing one at the expense of others produces agents that pass unit tests but fail in production.
How large should my tool use evaluation set be?
A useful evaluation set contains at least 50 examples per tool in your agent's toolkit, balanced between common cases and edge cases. For an agent with five tools, that means at least 250 labeled examples. Fewer than 30 examples per tool will miss boundary conditions where the model hesitates between similar functions. More than 100 per tool offers diminishing returns unless you are fine-tuning a model specifically for tool selection. Update the set continuously as you discover new failure modes in production logs.
How do I evaluate tool use when tools have side effects?
Separate evaluation into pre-execution validation and post-execution review. Before the tool runs, check selection accuracy and parameter correctness against ground truth—this catches most errors without triggering side effects. For execution success, use sandbox environments or mock tool implementations during testing. In production, log every tool call with full context and sample a subset for human review. Set up rollback procedures for high-risk tools like payment or data deletion functions, and require human-in-the-loop confirmation for those actions until your agent proves reliable.
WeaveAI Cite
Get your business named in AI answers.
Cite finds the questions people ask AI about what you do, then writes and publishes the articles that answer them, on autopilot.
Weekly digest
New articles, once a week
What we published on agent readiness, retrieval and evals, in one email on Mondays. Nothing in weeks with nothing to send.