Tool Calling in LLM Agents: A Practical Definition
Tool calling allows an LLM to request an external function, API, database query, search, or other operation through structured data. The surrounding application decides whether to execute the request, supplies the result to the model, and controls permissions, validation, and logging.
How tool calling works#
A typical tool-calling workflow has six stages:
-
Tool definitions are provided to the model.
The application sends the model a list of available tools, including each tool’s name, purpose, input schema, and sometimes usage guidance. -
The model evaluates the task.
Based on the user’s request and the available tools, the model decides whether it can respond directly or should request one or more tool calls. -
The model emits a structured request.
Instead of returning only prose, it produces a tool name and arguments, commonly represented as JSON. The model is proposing an operation, not executing it. -
The application validates the request.
The host system checks that the tool exists, arguments match the expected schema, values are within safe limits, and the user or service identity is authorized. -
The application executes the tool.
The tool may call an internal API, retrieve records, run a search, create a ticket, send a message, or perform another operation. High-impact actions may require approval. -
The result is returned to the model.
The application supplies the tool output as structured context. The model can then explain the result, request another tool, or finish the interaction.
This creates a loop rather than a single request-and-response exchange. An agent may search for an order, inspect its status, call a shipping API, and summarize the result across several tool calls.
Analyst’s Take: The application boundary is the control point in this workflow. The model can choose an inappropriate tool, produce invalid values, or be influenced by malicious tool output, so validation, authorization, execution, and logging cannot be delegated to the model.
Tool calling does not make an LLM deterministic or trustworthy by itself. Models can select an inappropriate tool, misunderstand user intent, produce invalid values, or be influenced by malicious content returned by a tool. Reliability and security must be implemented around the model.
Technical notes
A simplified tool definition and model request might look like this:
{
"name": "lookup_order",
"description": "Retrieve order status for an authenticated customer",
"parameters": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"pattern": "^ORD-[0-9]{8}$"
}
},
"required": ["order_id"],
"additionalProperties": false
}
}
The model may respond with a request such as:
{
"tool_name": "lookup_order",
"arguments": {
"order_id": "ORD-12345678"
}
}
A secure application should not blindly deserialize and execute this output. A simplified control flow is:
request = parse_model_output(response)
if request.tool_name not in allowed_tools:
raise SecurityError("Tool is not available")
arguments = validate_schema(
request.tool_name,
request.arguments
)
authorize(
user=current_user,
tool=request.tool_name,
arguments=arguments
)
result = execute_tool(
request.tool_name,
arguments,
timeout_seconds=10
)
audit_log(
user=current_user.id,
tool=request.tool_name,
arguments=redact_sensitive(arguments),
result_summary=summarize(result)
)
Production implementations should also apply least-privilege credentials, rate limits, timeouts, output-size limits, replay protections where appropriate, and explicit confirmation for destructive or externally visible actions. Credentials and secrets should remain outside model-visible context; organizations may also use a managed password tool such as 1Password to protect administrator access to connected systems.
When you’ll encounter tool calling#
You will encounter tool calling whenever an LLM application needs capabilities beyond generating text. Common examples include:
- Enterprise assistants: Searching internal documents, checking account data, opening support tickets, or retrieving HR information.
- Customer service automation: Looking up orders, checking eligibility, issuing refunds, and updating cases.
- Developer tools: Reading repositories, querying issue trackers, running tests, and generating pull requests.
- Security operations: Enriching indicators, querying SIEM data, creating incidents, or requesting containment actions. Teams may evaluate these workflows alongside established endpoint controls; see this guide to the difference between EDR and XDR.
- Productivity agents: Managing calendars, drafting messages, searching files, and updating project systems.
- Data and research workflows: Querying databases, browsing approved sources, transforming files, and producing reports.
- Infrastructure operations: Inspecting deployments, checking service health, or proposing infrastructure changes.
Risk depends on what the tools can do. A read-only documentation search is generally lower risk than a tool that can delete cloud resources, transfer money, change access controls, or send messages to customers.
Security teams should treat tool definitions as part of an application’s attack surface. Prompt injection can appear in user input, retrieved documents, web pages, email, tickets, or tool results. An attacker may attempt to persuade the model to disclose secrets, ignore restrictions, or invoke a privileged tool. The application must enforce authorization independently of the model’s instructions.
Useful operational questions include:
- Which tools are available to each user, tenant, and workflow?
- Which tools are read-only, and which change state?
- Are arguments validated against strict schemas?
- Does the tool run with a narrowly scoped identity?
- Are sensitive values removed from model-visible output?
- Is human approval required for high-impact operations?
- Can security staff reconstruct who requested and approved each action?
- Do retention and audit controls support applicable requirements, such as those discussed in this PCI DSS overview?
Related terms
Often used interchangeably with tool calling. “Function” commonly describes a callable operation, while “tool” can include functions, APIs, searches, code interpreters, and other capabilities.
An application pattern in which a model plans or performs multiple steps, often using tools, memory, or iterative feedback.
A broader workflow that uses an LLM to select actions or coordinate tasks. It may be deterministic, model-driven, or a combination of both.
A pattern that retrieves relevant information and gives it to the model. Retrieval can be implemented as a tool, but RAG does not necessarily involve taking external actions.
A protocol pattern for connecting AI applications to external tools and data sources. It standardizes parts of tool discovery and interaction but does not replace authorization or application security.
Model output constrained to a specified format or schema. Tool calls are a specialized form of structured output intended to request an operation.
A control requiring a person to review or approve selected model actions before execution.
An attempt to manipulate model behavior through instructions embedded in user input or untrusted content. It is especially relevant when an agent can call privileged tools.