Imagine an AI agent being asked a simple question:
Which of twenty employees went over their Q3 travel budget?
At first, this sounds like an easy task. But for an AI agent, answering it may require retrieving expense information for all twenty employees, reviewing flights, hotels, meals, and other expenses, and comparing those totals to each employee’s budget.
If the agent handles the task using traditional tool calling, it may need to make twenty separate tool calls. Each call could return dozens or even hundreds of expense records. The model then has to receive all of that information, keep it in context, add up the expenses, compare the totals, and determine who exceeded their budget.
That could mean thousands of individual expense records flowing through the model, even though the model doesn’t actually need to read every transaction. It only needs the final totals and the names of the employees who exceeded their budgets.
This example highlights an important architectural decision when building AI agents: How should an agent interact with tools and take action?
There are two major approaches: tool calling and code execution.
Both allow an AI agent to interact with external systems, but they work differently. The choice can affect the agent’s cost, speed, context usage, reliability, and overall complexity.
This article explains the difference between the two approaches, shows where each one makes sense, and provides a practical way to decide which approach to use.
If you are new to AI agent tool calling, it can help to first understand the fundamentals of tool-based agents before exploring more advanced approaches.
What Is an Action Primitive?
An action primitive is the basic mechanism an AI agent uses to turn a decision into an action.
That action could be retrieving information from a database, calling an API, reading a file, updating a record, or performing another operation in an external system.
Most AI agents rely on some form of tool interaction to accomplish these tasks.
The important question is not simply whether an agent can use tools. It is how the agent uses them.
With traditional tool calling, the model requests an action, waits for the result, receives that result in its context, and then decides what to do next.
With code execution, the model can write a small program that controls multiple tool calls, processes the results, and returns only the information that is actually needed.
That difference becomes increasingly important as tasks become larger and more complicated.
How Tool Calling Works
Tool calling is the approach most developers encounter first when building an AI agent.
The basic process is straightforward.
The model decides that it needs information from an external system. It creates a structured request specifying which tool it wants to use and what information the tool needs. The application then executes that tool and sends the result back to the model.
The model receives the result and continues reasoning about the task.
If another tool is needed, the process happens again.
This creates a simple cycle:
Model decides → tool runs → result returns → model decides again.
There are important advantages to this approach.
Each tool call is a clearly defined action. Developers can log what happened, inspect the arguments that were provided, and see what result was returned. This makes traditional tool calling relatively straightforward to understand, monitor, and debug.
For simple tasks, that simplicity is extremely valuable.
The problem appears when an agent needs to perform many actions or process large amounts of information.
Every intermediate result that is returned to the model becomes part of the model’s context. If the agent makes twenty calls and each call returns a large amount of information, the context can quickly become crowded with data that the model does not actually need.
Anthropic describes this as one of the major limitations of traditional tool use as agents become more complex. Intermediate tool results can consume significant context and require additional model inference passes.
How Code Execution Changes the Process
Code execution takes a different approach.
Instead of asking the model to request every action individually, the model can write a program that controls the workflow.
The program can call several tools, repeat an operation, process the results, filter information, perform calculations, and combine the results before anything is returned to the model.
The important difference is what happens to the intermediate information.
With traditional tool calling, the results generally return to the model after each operation.
With programmatic tool calling, the intermediate results can remain inside the execution environment while the program works with them.
Only the final result needs to be returned to the model.
Anthropic’s documentation describes Programmatic Tool Calling as a way for Claude to orchestrate tools through code rather than requiring every individual tool call to return directly to the model’s context.
This creates a very different workflow:
Model creates a plan → code executes the workflow → tools provide results to the program → program processes the data → final result returns to the model.
The model still controls the overall task, but the repetitive data processing happens outside the model’s context.
A Simple Example: Checking the Weather
Consider a straightforward question:
“What’s the weather like in London right now?”
The agent needs to call a weather service, retrieve the current conditions, and explain the result to the user.
This is a perfect example of a task where traditional tool calling makes sense.
The agent makes one request to the weather service. The service returns the current conditions. The model reads that information and responds to the user.
There is very little intermediate data to process, and there is no need to coordinate dozens of separate operations.
Using code execution for a task this simple could introduce unnecessary complexity without providing a meaningful advantage.
The important lesson is that code execution is not automatically better simply because it is more powerful.
For a simple, single-action task, traditional tool calling can be the more straightforward approach.
When the Task Becomes More Complicated
Now imagine changing the question.
Instead of asking about one city, the user asks:
“Which of these fifteen cities will have the coldest high temperature this week, and what is the average weekly high across all of them?”
The agent now needs to retrieve weather information for fifteen different cities.
With traditional tool calling, the model may need to make fifteen separate requests and receive the results from each one. It then needs to compare the temperatures and calculate an average.
The amount of information grows quickly.
The model is now spending context and processing capacity on intermediate information that may not be important to the final answer.
With code execution, the agent can handle this differently.
It can request the weather data for all fifteen cities, process the results, calculate the average, determine which city has the lowest high temperature, and return only the final findings.
The model does not need to see every individual temperature record along the way.
This is where code execution becomes particularly useful.
Why This Matters for Large Workflows
The same concept becomes even more important with business data.
Consider the original travel-budget example.
An organization has twenty employees. Each employee has dozens of expense records from the third quarter. The agent needs to determine who exceeded their budget.
The agent does not necessarily need to show every flight, hotel stay, and meal receipt to the model.
Instead, the workflow can retrieve the data, calculate each employee’s total expenses, compare those totals against their budget, and return only the relevant results.
This is exactly the type of workflow that programmatic tool calling is designed to support.
Anthropic’s example of budget compliance describes a scenario involving twenty employees, more than 2,000 expense line items, and more than 50 KB of raw expense data. With programmatic tool calling, the intermediate expense records can be processed without filling the model’s context, leaving the model with only the final results it needs.
The advantage is not simply that fewer tokens are used.
The agent also has less information to keep track of while reasoning.
Instead of asking the model to mentally manage thousands of records, the workflow allows ordinary program logic to handle calculations, filtering, sorting, and aggregation.
The Cost of Sending Everything Through the Model
Context is one of the most important resources in an AI agent.
Every piece of information sent to the model consumes context. Large tool results can therefore become expensive, especially when the information is only being used for an intermediate calculation.
Imagine an agent that needs to retrieve information from twenty different sources.
If every result goes directly back to the model, the model has to process all twenty responses. If those responses contain thousands of records, the amount of information can become substantial.
The agent may ultimately need only a few numbers.
This creates a mismatch between the amount of information the model processes and the amount of information the model actually needs.
Code execution helps solve that mismatch by allowing the intermediate processing to happen outside the model’s main context.
Anthropic reports that its Programmatic Tool Calling testing showed average token usage falling from 43,588 tokens to 27,297 tokens on complex research tasks, a reduction of about 37%. The same testing also reported improvements on the evaluated knowledge-retrieval and GAIA tasks. These figures are Anthropic’s internal benchmark results, so they should be viewed as evidence from those particular tests rather than a guarantee for every application.
Code Execution Is Not Just About Saving Tokens
It is easy to think of code execution as simply a way to reduce costs.
There is more to it than that.
Code is particularly good at repetitive and structured operations.
A program can loop through hundreds of records, sort information, calculate totals, filter results, and perform comparisons without requiring the model to reason through every individual step in natural language.
This can make complex workflows easier to manage.
For example, suppose an agent needs to check thousands of customer records and identify customers who meet several conditions.
The model could receive all of those records and attempt to reason through them.
Or it could create a program that applies the conditions consistently to every record and returns only the customers who qualify.
The second approach allows the model to focus on interpreting the task and explaining the result while the program handles the repetitive processing.
When Traditional Tool Calling Makes More Sense
Despite the advantages of code execution, traditional tool calling remains extremely useful.
If the agent needs to make one simple API request, traditional tool calling is often enough.
It is also useful when the model genuinely needs to inspect the intermediate result and reason about it.
For example, suppose an agent retrieves a document and needs to understand a specific passage before deciding what to do next. In that situation, the intermediate information may be exactly what the model needs to see.
There is also a practical infrastructure consideration.
Code execution requires a secure environment in which generated code can run. Organizations need to consider sandboxing, permissions, monitoring, security, and operational complexity.
Traditional tool calling can be easier to implement and understand when the workflow is relatively simple.
Auditability can also be an important consideration. Individual tool calls are easy to log and inspect. With code-driven orchestration, developers may need additional logging and monitoring to understand exactly what a generated program did.
For these reasons, code execution should not be treated as a replacement for traditional tool calling.
It is another tool for solving a different class of problems.
The Real Difference Comes Down to the Workflow
The simplest way to understand the distinction is to ask one question:
Does the model need to see every intermediate result?
If the answer is yes, traditional tool calling may be the natural choice.
If the answer is no, and the information can be processed before the model sees it, code execution may provide significant advantages.
A single lookup usually does not justify the additional complexity of code execution.
A workflow involving dozens of independent calls, large datasets, calculations, filtering, or aggregation is much more likely to benefit from programmatic execution.
The more the workflow looks like data processing, the more valuable code execution becomes.
A Practical Way to Choose
When deciding between the two approaches, start by looking at the task rather than the technology.
If the task requires one or a few straightforward actions and the results are relatively small, traditional tool calling is often sufficient.
If the task requires many calls, especially calls that can happen independently, code execution becomes more attractive.
If the intermediate results are large but the final answer is small, code execution can help keep unnecessary information out of the model’s context.
If the workflow requires calculations, sorting, filtering, aggregation, or repeated operations, code is often better suited to that work than natural-language reasoning.
If the model needs to examine and interpret every intermediate result, however, sending those results directly to the model may still be the right approach.
The decision ultimately comes down to the structure of the task.
The Hybrid Approach
In real-world AI systems, there is rarely a reason to choose only one approach for everything.
A well-designed agent can use traditional tool calling for simple operations and programmatic execution for complex workflows.
For example, an agent might use a normal tool call to look up a customer’s account, then use code execution to process thousands of related transactions.
Another agent might use traditional tool calling to retrieve a document, then use code execution to analyze a large dataset referenced by that document.
This hybrid approach allows each method to be used where it provides the most value.
Anthropic’s current approach reflects this broader idea. Its agent tooling includes Programmatic Tool Calling alongside capabilities such as Tool Search and Tool Use Examples, each addressing different challenges in tool-heavy workflows. Tool Search helps reduce the amount of tool-definition information loaded into context, while Programmatic Tool Calling focuses on efficiently orchestrating tools and processing their results.
The goal is not to use every available feature.
The goal is to identify the actual bottleneck in the workflow and use the mechanism that addresses it.
The Bigger Lesson for AI Agent Design
The most important lesson is that tool calling and code execution are not simply two different ways of writing the same thing.
They represent two different approaches to how an AI agent interacts with the world.
Traditional tool calling keeps the model closely involved in each step. The model requests an action, receives the result, reasons about it, and decides what happens next.
Code execution allows the model to delegate a larger portion of the workflow to a program. The program can perform repetitive operations, coordinate multiple tools, and process large amounts of information before returning a concise result.
Neither approach is universally correct.
The right choice depends on the task.
For a question like “What’s the weather in London right now?”, a direct tool call is simple and appropriate.
For a question like “Which employees exceeded their travel budgets after analyzing thousands of expense records?”, processing the information programmatically can prevent the model from being overwhelmed by data it never needed to read.
That distinction becomes increasingly important as AI agents move from simple demonstrations to real production systems.
Conclusion
The way an AI agent interacts with tools is an architectural decision, not simply a coding preference.
Traditional tool calling works well when an agent needs a small number of actions and the model benefits from seeing the results directly.
Code execution becomes more valuable when an agent needs to coordinate many actions, process large datasets, perform calculations, or filter information before it reaches the model.
The most effective systems will often use both.
The real skill is knowing when to switch between them.
If a task requires the model to understand an intermediate result, let the model see it.
If a task requires thousands of repetitive operations but only produces a small final answer, let code handle the processing.
The goal is not to make the model do more work.
The goal is to make sure the model does the work that only the model needs to do, while the rest of the workflow is handled by the tools and execution environment best suited for it.

