There's a counterintuitive finding buried in agent benchmarking research: adding tools to an agent often makes it perform worse.
Not always. Not catastrophically. But consistently, measurably, worse.
The explanation is uncomfortable for anyone who has spent time building agent integrations: when an agent has 50 tools, it spends cognitive capacity choosing between them. When it has 5, it thinks about the task.
This is the tool proliferation problem. And most teams are somewhere in the middle of creating it.
The Toolbox Fallacy
The natural instinct when building an agent is to give it capabilities. Search the web. Query a database. Send an email. Call an API. Read a file. Write a file. Create a calendar event. The more it can do, the more useful it is.
This logic is true in aggregate and false in practice.
What actually happens: the model's attention budget during a task is finite. Some fraction of that budget gets spent on _tool selection_—evaluating which of the available tools is appropriate for the current step. With 10 tools, this is a minor tax. With 50 tools, it becomes a significant overhead that competes with the actual reasoning about the task.
Worse, the tools interact. An agent with a `read_file` tool and a `search_web` tool has to decide which one to use when asked "what does X library do?" The correct answer depends on context. With enough tools, the agent ends up reasoning about tool selection more than about the problem.
Researchers have documented this effect. In studies where agents were given subsets of available tools versus full tool sets, performance on specific tasks typically improved when the tool set was narrowed to what was relevant. The agents didn't need more options—they needed fewer distractions.
Why Teams Keep Adding Tools
The tool proliferation pattern is easy to understand from the product side.
Someone asks: "Can the agent do X?" You add a tool for X. User asks about Y. You add a tool for Y. After six months, your agent has 40 tools and a 4,000-token system prompt explaining when to use each one.
Each individual addition was reasonable. The cumulative effect is an agent that's genuinely harder to work with.
Part of the problem is that tool-addition is invisible in its costs. You add a new integration; the agent can now do something new. The fact that it now takes two attempts to do something it previously did in one—because it had to sort through more options—shows up only in aggregate, in user feedback, in task completion rates. Nobody attributes the degradation to the tool count.
The other part is organizational. In most companies, agents are built by multiple teams. Each team adds the tools they need. Nobody is responsible for the total tool set. Nobody feels the cost of the 49th integration.
The Attention Economy of Tool Selection
Here's a more precise model of what's happening.
When a language model processes a task, it considers all available tools as part of its context. Each tool has a name, a description, and a set of parameters. In a function-calling setup, this typically costs 20–100 tokens per tool, before you've written a single character of task instruction.
For 50 tools at 60 tokens each, you've spent 3,000 tokens just on tool definitions. In a 128K context window, that's a rounding error. But the attention cost isn't just about tokens—it's about the model's capacity to reason clearly.
Language models perform better when relevant information is concentrated and irrelevant information is absent. This is the "lost in the middle" effect in reverse: buried relevant tools get missed; prominent irrelevant tools create false choices. The model has to do extra work to ignore what it shouldn't use.
This is why narrowing the tool set to the task at hand isn't just a performance optimization—it's a correctness improvement.
Four Patterns That Work
1. Tool sets, not tool warehouses
Design distinct, bounded tool sets for different task types. An agent doing research work should see research tools: search, fetch, read. An agent doing code work should see code tools: read files, run tests, make edits. An agent doing communication work should see communication tools: email, calendar, message.
The same model can be deployed in multiple modes, each with a different tool subset. The overhead is configuration, not capability.
2. Capability over tool count
Good tools compose. A `run_bash_command` tool plus a `read_file` tool plus a `write_file` tool covers most file manipulation tasks without needing separate `create_file`, `append_to_file`, `delete_file`, `copy_file`, and `move_file` tools. Five composable primitives beat fifteen single-purpose functions.
The principle: tools should represent capabilities, not operations. If two tools can be combined to achieve what a third tool does explicitly, the third tool is probably unnecessary.
3. Lazy tool registration
For agents that genuinely need broad reach, register tools lazily. Present the agent with a small core tool set plus a meta-tool: `get_tool(name)` or `search_available_tools(query)`. The agent discovers tools it needs as it needs them, rather than holding all 50 in working memory at once.
This pattern trades a small discovery overhead for a large selection overhead reduction. In agents that use only a subset of available tools per task (which is most of them), it's almost always worth it.
4. Task-scoped context
When starting a specific task, load the relevant tool set into context and remove everything else. A task that involves "generate a weekly report from database metrics" needs database query tools and file write tools. It does not need the email tool, the calendar tool, the code execution tool, or the web search tool.
Scoping is usually done at the orchestration layer. The model shouldn't see tools that can't help it with the current task. If the model later discovers it needs an additional capability, the orchestrator can add it—but the default should be minimal.
The Tool Documentation Problem
There's a secondary issue that compounds tool proliferation: poor tool descriptions.
When tool descriptions are vague or overlapping, the selection problem gets worse. "Search" vs. "find" vs. "lookup" vs. "query"—if each is a different tool and the descriptions don't clearly distinguish them, the agent spends cognitive load on disambiguation instead of the task.
Good tool documentation answers three questions in the description:
1. _When should I use this?_ (not what it does, but when it's appropriate)
2. _What does it NOT do?_ (explicitly scope out adjacent tools)
3. _What do I need to know about the output format?_
Most tool descriptions answer only the first question, partially. Adding the second reduces ambiguity between similar tools. Adding the third reduces downstream handling errors.
Tool documentation is maintenance work that rarely gets done. After the initial integration, descriptions drift out of sync with behavior. When you add a new tool that overlaps with an existing one, you rarely update the existing one's description to add the disambiguation boundary.
This is boring work that has an outsized effect on agent performance.
When More Tools Are Right
This is not an argument for minimalism at all costs.
Some tasks genuinely require broad capability. A general-purpose assistant that handles arbitrary requests benefits from a wider tool set than a specialized agent that does one thing well. An agent that helps you navigate an unfamiliar codebase needs more tools than an agent that runs a weekly status report.
The question isn't "how few tools?" but "which tools, for this task?"
The failure mode to avoid is _accidental_ tool proliferation—the 50-tool set that grew by accumulation rather than design, where half the tools are never used, where the descriptions are stale, and where nobody is sure what the agent can actually do anymore.
Intentional tool sets—even large ones—work. The design question is whether the tool set was chosen for the task or arrived at by default.
The Audit You Haven't Done
Most teams have never looked at their agent's tool utilization data. Which tools actually get called? Which ones never get called? Which ones get called in sequences that suggest the agent is confused about which to use?
This data is usually available in your traces if you're logging function calls. An hour of analysis typically reveals:
- Tools that are never called (safe to remove)
- Tools that are called interchangeably (consolidation candidates)
- Tool selection loops (agent picks the wrong tool, fails, retries with another)
- High-selection-overhead tasks (where the agent spends multiple turns before the first real action)
The tool set you designed is probably not the tool set you need. The utilization data tells you what's actually being used and what's creating friction.
Build the right set from evidence, not assumptions.
Start With Ten
If you're building a new agent, start with ten tools maximum. Not ten categories—ten individual tools.
Force yourself to make tradeoffs. If you're going to add a file-write tool and a file-read tool, do you also need a file-append tool, or does write cover it? If you have web search, do you also need a dedicated Wikipedia API tool, or does search handle Wikipedia queries?
Starting small makes the costs of each addition visible. You can always add more. Removing tools from a production agent is much harder than adding them—users discover capabilities and build workflows around them.
Get the core right first. Expand from evidence.
The models are capable enough that ten good tools, well-documented, will outperform fifty mediocre tools in a cluttered prompt. Every time.
_Chris Korhonen builds and runs agent systems. This essay was written by an agent with access to exactly the tools needed to write and commit it._
export const social = {
title: 'The Tool Proliferation Problem',
description:
"Adding tools to an agent makes it worse. Here's why tool count is quietly degrading your agent's performance—and how to fix it.",
hashtags: ['AI', 'Agents', 'AgentDesign', 'LLM'],
hook: 'Counterintuitive finding: giving agents more tools makes them perform worse. Tool selection overhead is real, and most teams are somewhere in the middle of creating it.',
};