Chapter 14
Agents and Tools
From chatbots to action-oriented AI systems that plan, use tools, and complete workflows
|
Chapter Goal In earlier chapters, you learned about LLMs, prompting, embeddings, semantic search, vector databases, RAG, and common LLM application patterns. This chapter explains how these pieces come together to create AI agents: systems that can decide what step to take next, call external tools, remember useful context, and complete multi-step tasks under clear safety controls. |
A simple LLM application responds to a user message. A RAG application retrieves information from documents and then responds. An AI agent goes further: it can decide the next action, use tools, observe results, adjust its plan, and continue until a goal is achieved or until it must ask for human approval.
Agents are important because many real business tasks are not single-question tasks. A customer-support problem may require reading a ticket, searching a knowledge base, checking a database, drafting a reply, and escalating if the confidence is low. A sales task may require researching a company, finding contacts, reading news, updating a CRM, and drafting an outreach email. A personal productivity task may require reading a calendar, checking email, creating a schedule, and sending a reminder.
However, agents are also more risky than ordinary chatbots because they can take actions. A chatbot may give a wrong answer. An agent may send a wrong email, update the wrong record, delete the wrong file, or expose private data. Therefore, agent design must include permissions, approval gates, safety checks, audit logs, and fallback handling.
|
Simple Definition An AI agent is an AI system that uses an LLM to reason about a goal, choose actions, call tools, observe results, and continue until the task is completed or escalated. |
An AI agent is a software system that combines intelligence with action. The intelligence usually comes from an LLM. The action comes from tools, APIs, databases, files, or other systems. The agent receives a goal, breaks it into steps, selects the right tool for each step, reads the tool result, and decides what to do next.
A human office assistant provides a good analogy. If you tell an assistant, “Plan my business trip to Bengaluru next week,” the assistant does not simply answer with a paragraph. The assistant may check your calendar, search flights, compare hotels, consider budget, draft an itinerary, ask for approval, and then make bookings. An AI agent tries to automate parts of this behaviour in software.
|
Concept |
Meaning |
Simple Example |
|
Goal |
The final outcome the user wants |
Resolve this customer ticket. |
|
Plan |
The steps the agent proposes |
Read ticket, search KB, check order, draft reply. |
|
Tool |
An external function the agent can use |
Search knowledge base, query order database. |
|
Observation |
The result returned by a tool |
Order status is delayed due to payment verification. |
|
Action |
The next step selected by the agent |
Draft reply explaining the delay. |
|
Approval |
Human confirmation before risky action |
Ask support manager before refunding money. |
A chatbot mainly communicates. An agent communicates and acts. This difference sounds small, but it changes the whole architecture. A chatbot can be implemented with a prompt and an LLM. A RAG chatbot may add retrieval from a vector database. An agent needs a planner, tool registry, execution layer, permission system, memory, logging, and safety checks.
|
Area |
Normal Chatbot |
AI Agent |
|
Main purpose |
Answer user questions |
Complete tasks or workflows |
|
Interaction style |
Single-turn or multi-turn conversation |
Goal-driven multi-step execution |
|
Tool use |
Usually none or limited |
Central part of the system |
|
Decision-making |
Responds based on prompt/context |
Chooses next action based on goal and observations |
|
Memory |
Mostly conversation history |
Conversation history plus task state, user preferences, and external records |
|
Risk level |
Lower, because it mainly speaks |
Higher, because it may perform real actions |
|
Example |
Answer: What is my leave policy? |
Check leave balance, read policy, draft leave request, ask approval, submit request |
|
Important Point Do not build an agent just because it sounds advanced. If your use case is only question-answering over documents, a RAG chatbot may be better. Use agents when the system must decide between steps, use tools, and complete a workflow. |
A production-ready agent is not only an LLM. The LLM is one component inside a larger application. The following text diagram shows a typical agent architecture.
+-------------------+
| User |
| Goal / Request |
+---------+---------+
|
v
+-------------------+ +-------------------------+
| User Interface | | Auth / Permissions |
| Chat / App / API +------->+ User role, scope, limits|
+---------+---------+ +-----------+-------------+
| |
v v
+-------------------------------------------------------+
| Agent Orchestrator |
| - Receives goal |
| - Maintains task state |
| - Calls planner and executor |
| - Applies guardrails and approval rules |
+----------------------+--------------------------------+
|
+--------+---------+
| |
v v
+---------------------+ +-----------------------------+
| Planner | | Memory |
| Decide next steps | | Short-term + long-term |
+----------+----------+ +-----------------------------+
|
v
+---------------------+ +-------------------------+
| Executor |------->| Tool Registry |
| Runs selected tool | | Search, DB, Email, etc. |
+----------+----------+ +------------+------------+
| |
v v
+---------------------+ +-------------------------+
| Observations |<-------| External Systems |
| Tool results | | Web, CRM, DB, files |
+----------+----------+ +-------------------------+
|
v
+---------------------+
| LLM Reasoning Layer |
| Interpret, decide, |
| draft, summarize |
+----------+----------+
|
v
+---------------------+ +-------------------------+
| Human Approval Gate |------->| Final Action / Response |
| For risky actions | | Send, update, answer |
+---------------------+ +-------------------------+
In simple terms, the user gives a goal, the agent orchestrator manages the task, the planner decides what needs to happen, the executor calls tools, the observations return results, the LLM interprets the results, and guardrails decide whether the agent can continue automatically or must ask for human approval.
Tool calling means giving the AI system access to predefined tools that perform specific actions. A tool can be a search function, a database query, a calculator, an email sender, a calendar reader, a file reader, or an API call. The LLM does not directly access everything on the internet or inside the company. Instead, the application exposes safe tools with clear descriptions and parameters.
LLMs are strong at language understanding and generation, but they are not reliable databases, calculators, or live system connectors by themselves. They may not know the latest information. They may guess. They may make arithmetic mistakes. Tools solve this by letting the agent ask reliable systems for facts or perform controlled actions.
|
Need |
Without Tool |
With Tool |
|
Latest information |
LLM may use outdated knowledge |
Agent searches current source |
|
Company data |
LLM cannot know private records |
Agent queries internal database/API |
|
Accurate calculation |
LLM may approximate |
Agent calls calculator or code tool |
|
Workflow action |
LLM can only describe steps |
Agent can create ticket, send draft, update CRM |
|
Verification |
LLM may hallucinate |
Agent can check source and cite evidence |
Example tool definition concept:
Tool name: search_knowledge_base
Purpose: Search approved company support articles.
Input: { "query": "text entered by agent", "top_k": 5 }
Output: List of matching articles with title, snippet, confidence, and source URL.
The agent does not freely browse all systems. It can only use the tools that the application exposes.
Function calling is a structured form of tool calling. The application defines functions with names, descriptions, input parameters, and expected output. The LLM chooses which function should be called and provides the arguments in a structured format, often JSON. The application then executes the function and returns the result to the agent.
Function calling is useful because it reduces ambiguity. Instead of asking the LLM to write a sentence like “search for this order,” the LLM produces structured data such as function name: get_order_status and arguments: order_id = 12345. The application can validate the arguments before execution.
Pseudo-code example:
User: "Check order 12345 and tell the customer the status."
Agent decides to call:
get_order_status({"order_id": "12345"})
Tool result:
{"order_id": "12345", "status": "Delayed", "reason": "Payment verification pending"}
Agent then drafts:
"Your order is currently delayed because payment verification is pending..."
|
Function Calling Element |
Description |
Example |
|
Function name |
The operation to run |
get_order_status |
|
Description |
What the function does |
Returns current order status |
|
Input schema |
Required parameters and types |
order_id: string |
|
Validation |
Checks before execution |
Order ID must be numeric |
|
Output schema |
Expected result format |
status, reason, estimated_delivery |
|
Permission rule |
Who can use it |
Support agent role only |
The planner is the part of the agent that decides what steps are needed to complete the user goal. It may be a prompt, a planning algorithm, a workflow engine, or a combination of rules and LLM reasoning. The planner answers questions such as: What is the goal? What information is missing? Which tool should be used first? Is approval required? What should be done if the tool result is incomplete?
Goal: Resolve VPN ticket.
Plan:
1. Read the ticket details.
2. Search knowledge base for VPN password reset issues.
3. Check if the user account is locked.
4. If account is locked, create an IT action request.
5. Draft reply to user.
6. Ask human support engineer to approve before sending.
In production systems, planning should not be fully unrestricted for sensitive workflows. For example, a banking agent should not independently decide to refund money or change customer KYC status. Such actions need rule-based approval gates.
The executor runs the action selected by the planner. If the planner says “search the knowledge base,” the executor calls the search tool. If the planner says “read ticket details,” the executor calls the ticketing system API. If the planner says “send email,” the executor first checks whether sending is allowed or requires approval.
|
Planner Responsibility |
Executor Responsibility |
|
Decide that database status is needed |
Call the database/API tool |
|
Decide that email should be drafted |
Generate draft and save it |
|
Decide that approval is required |
Pause and request human approval |
|
Decide next step after observation |
Execute the next selected tool |
|
Decide when task is complete |
Return final response or mark workflow done |
|
Practical Design Tip Keep the executor deterministic. The planner may use LLM reasoning, but the executor should follow strict software rules: validate input, check permissions, call tool, log result, handle error, and return a structured observation. |
Memory allows an agent to remember useful information. Without memory, every turn is treated almost like a new request. With memory, the agent can remember the conversation, the current task state, previous tool results, user preferences, and long-term facts that are safe and useful.
Short-term memory stores information needed for the current conversation or current task. It usually includes recent messages, current plan, tool results, and temporary decisions. For example, in a travel planning conversation, short-term memory may store destination, dates, budget, and preferred hotel area.
|
Short-Term Memory Item |
Example |
|
Current goal |
Plan a 3-day trip to Jaipur |
|
User constraints |
Budget under Rs. 20,000 |
|
Recent tool result |
Hotel search returned 5 options |
|
Current step |
Waiting for user to approve itinerary |
|
Temporary preference |
Avoid late-night flights for this trip |
Long-term memory stores useful information across sessions. This may include stable preferences, frequently used formats, project context, role information, or approved business rules. Long-term memory must be handled carefully because it may contain personal or sensitive information. Only useful, appropriate, and permitted information should be stored.
|
Long-Term Memory Item |
Possible Use |
|
User prefers simple explanations |
Adjust learning content style |
|
User works in data engineering |
Use relevant examples |
|
Company escalation policy |
Route tickets correctly |
|
Approved response templates |
Generate consistent customer replies |
|
Known system names |
Understand internal acronyms |
|
Memory Warning Memory is powerful but risky. Never store unnecessary personal details. In enterprise applications, memory should follow privacy policies, retention rules, access controls, and audit requirements. |
The agent loop is the repeated cycle of thinking about the goal, selecting an action, using a tool, observing the result, and deciding the next step. The loop continues until the task is complete, blocked, failed, or escalated.
Agent loop in simple form:
1. Receive goal from user.
2. Understand the goal and constraints.
3. Decide next action.
4. Call selected tool.
5. Observe tool result.
6. Update task state.
7. Decide whether to continue, ask user, ask human approver, or finish.
8. Return final answer or complete action.
An uncontrolled loop can be dangerous. The agent may keep calling tools again and again, increasing cost and causing confusion. Therefore, production systems should define maximum steps, timeout, cost limit, retry limit, and safe fallback behaviour.
|
Control |
Purpose |
|
Maximum steps |
Prevents endless loop |
|
Timeout |
Stops long-running tasks |
|
Cost limit |
Avoids high API bills |
|
Retry limit |
Avoids repeated failed tool calls |
|
Approval gate |
Prevents risky automatic actions |
|
Fallback response |
Explains when the agent cannot proceed |
A single-agent system uses one agent to plan and execute the whole task. This is simpler to design, easier to monitor, and usually better for beginners. Many business use cases can be solved with a single agent plus good tools and guardrails.
|
Single-Agent Example An IT support agent receives a ticket, searches the knowledge base, checks asset information, drafts a response, and sends it for human approval. One agent controls the entire process. |
A multi-agent system uses multiple specialized agents. Each agent has a role, such as researcher, planner, analyst, writer, reviewer, or executor. Multi-agent systems can be useful for complex workflows, but they are harder to debug and control.
|
Agent Role |
Responsibility |
Example |
|
Research Agent |
Collect information |
Search web and company documents |
|
Analysis Agent |
Interpret information |
Compare vendors or detect ticket category |
|
Writer Agent |
Prepare final content |
Draft report, email, or summary |
|
Reviewer Agent |
Check quality and risks |
Verify sources, tone, compliance |
|
Executor Agent |
Perform approved actions |
Create CRM task or send approved email |
|
When to Use Multi-Agent Design Use multi-agent design only when the task naturally needs different roles, review stages, or parallel work. For most beginner and enterprise use cases, start with a single well-controlled agent. |
Human approval means the agent must pause and ask a person before performing certain actions. This is essential when an action can affect money, legal decisions, customer communication, security, production systems, or personal data.
|
Action |
Should Human Approval Be Required? |
Reason |
|
Drafting an email |
Usually no, if saved as draft |
Low risk if not sent automatically |
|
Sending an email to customer |
Often yes |
Wrong message can harm trust |
|
Updating CRM note |
Maybe |
Depends on company policy |
|
Refunding payment |
Yes |
Financial risk |
|
Deleting file |
Yes |
Data loss risk |
|
Querying read-only knowledge base |
Usually no |
Low risk |
|
Changing production database |
Yes |
High operational risk |
Human approval workflow:
Agent: I found that the customer may be eligible for a refund.
Proposed action: Create refund request for Rs. 2,500.
Reason: Payment captured but order cancelled within allowed period.
Evidence: Order record #ORD-1122 and refund policy section 4.
Approval required: Yes.
Human: Approved.
Agent: Creates refund request and logs action.
Agent safety means designing the agent so it behaves within allowed boundaries. Safety is not only about harmful content. In business systems, safety also means preventing data leakage, unauthorized actions, incorrect automation, prompt injection, high cost, and unreliable decisions.
|
Safety Risk |
Example |
Control |
|
Unauthorized action |
Agent sends email without permission |
Role-based access and approval gate |
|
Data leakage |
Agent exposes salary data to wrong user |
Access control and PII filtering |
|
Prompt injection |
Document says ignore all previous rules |
Instruction hierarchy and content sanitization |
|
Wrong tool call |
Agent updates wrong ticket |
Input validation and confirmation |
|
Over-automation |
Agent closes ticket incorrectly |
Human review for final action |
|
High cost |
Agent loops many times |
Step and token limits |
|
Bad evidence |
Agent answers without source |
Require citations and confidence thresholds |
Tool permissions define which tools an agent can use, for whom, and under what conditions. A good agent platform does not give every agent access to every tool. It uses least privilege, which means the agent receives only the minimum access needed for the task.
|
Tool Type |
Read Permission |
Write Permission |
Typical Approval |
|
Knowledge base search |
Allowed for most users |
Not applicable |
No |
|
Customer database |
Role-based |
Restricted |
Yes for updates |
|
|
Can read limited mailbox if authorized |
Draft allowed, send restricted |
Often yes |
|
Calendar |
Can read availability if authorized |
Create/update restricted |
Sometimes |
|
File system |
Read allowed for approved folders |
Delete/move restricted |
Yes for destructive changes |
|
Code execution |
Sandbox only |
No direct production execution |
Yes for production changes |
|
Least Privilege Rule Never give an agent more access than it needs. If the task only needs reading a policy document, do not provide email-sending, database-write, or file-delete tools. |
Most useful agents work by connecting to APIs. An API is a controlled way for one software system to communicate with another. For example, an agent may use a ticketing API to read support tickets, a CRM API to update customer records, a calendar API to check availability, or a database API to retrieve account status.
The agent should not directly build unsafe database queries or uncontrolled network requests. The application should expose safe API wrappers. These wrappers validate inputs, apply permissions, log every call, and return structured results.
Safe API wrapper concept:
function get_customer_summary(customer_id, requesting_user):
check_user_permission(requesting_user, "customer_read")
validate(customer_id)
result = call_crm_api(customer_id)
remove_sensitive_fields(result)
log_access(requesting_user, customer_id)
return result
A browser or search tool lets the agent retrieve current information from the web or approved search sources. This is useful when the answer depends on recent information, public facts, market research, documentation, or news. In enterprise systems, search tools should be restricted to trusted domains or approved sources when accuracy matters.
|
Use Case |
Example Agent Task |
Risk |
Best Practice |
|
Travel planning |
Search flights and hotel areas |
Outdated or sponsored content |
Verify multiple sources |
|
Sales research |
Find recent company news |
Low-quality sources |
Restrict to trusted sources |
|
Technical support |
Search vendor docs |
Wrong version docs |
Use official documentation |
|
Market analysis |
Find competitor info |
Unverified claims |
Cite sources and mark uncertainty |
A database tool lets the agent retrieve structured records. For example, it may check order status, leave balance, ticket history, product inventory, or transaction summary. Database tools should be carefully designed because databases often contain sensitive information.
Good database tool design:
- Use predefined queries or parameterized queries.
- Validate input values.
- Apply row-level and column-level permissions.
- Return only required fields.
- Log who accessed what and why.
- Avoid giving the LLM raw unrestricted SQL access in production.
An email tool allows an agent to read, draft, send, or organize email. This is powerful but risky because email is external communication. A safe design usually allows the agent to create drafts but requires human approval before sending.
|
Email Action |
Risk Level |
Recommended Control |
|
Search email |
Medium |
Limit mailbox scope and log access |
|
Summarize email |
Medium |
Mask sensitive data where needed |
|
Create draft |
Low to medium |
User reviews before sending |
|
Send email |
High |
Require explicit user approval |
|
Forward email |
High |
Confirm recipient and attachments |
|
Delete email |
High |
Avoid automation or require confirmation |
A calendar tool allows the agent to check availability, schedule meetings, update events, or send invitations. Calendar tools are useful for productivity agents, sales agents, recruiting agents, and project-management assistants.
|
Calendar Agent Example User: Schedule a 30-minute meeting with the project team next week. Agent: Checks availability, finds open slots, proposes two options, waits for approval, then creates the meeting invite. |
A file tool lets the agent read or write files such as PDFs, DOCX documents, spreadsheets, logs, notes, and reports. File tools are common in document Q&A, report summarization, legal review, audit analysis, and knowledge-base creation.
|
File Task |
Example |
Control |
|
Read file |
Read a policy PDF |
Check file permission |
|
Extract text |
Convert DOCX to text |
Preserve metadata and source location |
|
Summarize file |
Summarize quarterly report |
Mention limitations |
|
Create file |
Generate DOCX report |
Use approved template |
|
Delete file |
Remove outdated draft |
Require explicit confirmation |
A code execution tool allows the agent to run code in a controlled environment. It is useful for calculations, data analysis, chart generation, file processing, and testing small scripts. However, it must be sandboxed. The agent should not be allowed to run arbitrary code on production systems.
Safe code execution rules:
- Run code in a sandbox.
- Disable unnecessary network access.
- Limit execution time.
- Limit memory and storage.
- Do not expose secrets or production credentials.
- Log executed code and outputs.
- Require approval before using results for important decisions.
A travel planning agent helps a user plan a trip by collecting preferences, searching options, comparing choices, and creating an itinerary. It may use search tools, maps, hotel APIs, calendar tools, and email tools. Because bookings involve money, the agent should not make purchases without approval.
|
Step |
Agent Action |
Tool Used |
Approval Needed? |
|
1 |
Understand destination, dates, budget, and preferences |
Conversation memory |
No |
|
2 |
Check user calendar for available dates |
Calendar tool |
Maybe |
|
3 |
Search flights and hotels |
Search/API tool |
No |
|
4 |
Compare options |
LLM analysis |
No |
|
5 |
Prepare itinerary |
LLM + document tool |
No |
|
6 |
Book selected option |
Booking API |
Yes |
|
7 |
Email itinerary |
Email tool |
Yes before sending |
Travel planning agent architecture:
User goal
-> Agent orchestrator
-> Planner creates travel steps
-> Search tool finds transport and hotel options
-> Calendar tool checks availability
-> LLM compares options
-> Human approval gate for booking
-> Email tool sends itinerary after approval
|
Travel Agent Risks The agent may choose an expensive option, rely on outdated availability, misunderstand travel dates, or book without consent. Use budget limits, source verification, date confirmation, and mandatory approval for payment. |
An IT ticket resolution agent assists support teams by reading tickets, classifying issues, searching knowledge articles, checking system status, suggesting resolution steps, drafting replies, and escalating complex issues. This is a strong enterprise use case because many IT tickets follow repeatable patterns.
|
Ticket Type |
Agent Action |
Possible Tool |
|
VPN not working |
Search VPN troubleshooting guide and check account status |
Knowledge base + IAM API |
|
Laptop issue |
Check asset warranty and create hardware task |
Asset DB + ticketing API |
|
Password reset |
Verify user and trigger reset workflow |
Identity API |
|
Oracle database down |
Check monitoring alert and route to DBA team |
Monitoring + ticketing API |
|
Suspicious email |
Escalate to security and collect headers |
Email + security tool |
IT support agent flow:
1. Read ticket: "I cannot connect to VPN after password reset."
2. Classify category: Network / VPN.
3. Search knowledge base: VPN issue after password reset.
4. Check identity status: account not locked, password changed yesterday.
5. Suggest resolution: reset VPN profile and re-authenticate.
6. Draft response to user.
7. If issue persists, escalate to network team.
|
IT Agent Best Practice Start with recommendation mode. Let the agent classify, summarize, suggest steps, and draft replies. Add automatic actions only after the workflow is stable and reviewed by support teams. |
A sales research agent helps sales teams research target accounts, summarize company information, identify potential pain points, prepare outreach messages, and update CRM records. It may use web search, company databases, CRM APIs, email drafting tools, and document generation tools.
|
Step |
Agent Work |
Output |
|
1 |
Read target company name and industry |
Research goal |
|
2 |
Search public information and recent news |
Company summary |
|
3 |
Identify possible business problems |
Pain point list |
|
4 |
Map product offering to pain points |
Sales angle |
|
5 |
Draft outreach email |
Reviewable email draft |
|
6 |
Update CRM note after approval |
CRM activity record |
Sales research prompt example:
You are a sales research assistant.
Goal: Prepare a short account brief for a data engineering services company.
Company: ABC Retail Ltd.
Find likely data, analytics, cloud, or AI opportunities.
Output format:
1. Company overview
2. Possible pain points
3. Suggested outreach angle
4. Draft email
5. Confidence and sources
|
Sales Agent Risks Sales agents can generate inaccurate claims, use weak sources, or create overly aggressive messages. Always require source citations, tone review, and human approval before sending outreach. |
Agents are powerful because they can act. The same power makes them risky. A well-designed agent must include guardrails at every stage: input, planning, tool execution, memory, output, and final action.
|
Stage |
Risk |
Guardrail |
|
Input |
User asks for unauthorized action |
Check policy and user role |
|
Planning |
Agent creates unsafe plan |
Use allowed workflow templates |
|
Tool selection |
Wrong tool chosen |
Limit available tools by task type |
|
Tool input |
Invalid or malicious parameter |
Validate and sanitize input |
|
Tool output |
Sensitive data returned |
Mask or filter output |
|
Memory |
Stores private data unnecessarily |
Memory policy and retention limits |
|
Final answer |
Hallucinated or unsupported claim |
Require evidence and confidence |
|
Final action |
Sends or changes something incorrectly |
Human approval gate |
Agents are not always the best solution. Sometimes a simple prompt, a normal workflow, a RAG system, or a rule-based automation is cheaper, safer, and easier to maintain. The goal of architecture is not to use the most advanced technique. The goal is to solve the business problem reliably.
|
Situation |
Better Alternative |
Reason |
|
Only need to answer FAQs |
RAG chatbot |
No complex actions needed |
|
Fixed approval workflow |
Traditional workflow engine |
Rules are predictable |
|
Simple classification |
Prompt or ML classifier |
Agent loop unnecessary |
|
High-risk financial decision |
Human-led workflow with AI assistance |
Regulatory and financial risk |
|
Strict deterministic output required |
Rule-based system |
LLM variability may be unacceptable |
|
No reliable tools available |
Do not build agent yet |
Agent cannot safely act without trusted tools |
|
Low budget and low volume |
Manual process or simple automation |
Agent complexity may not be justified |
|
Architecture Rule Use the simplest architecture that safely solves the problem. Move from prompt -> RAG -> workflow automation -> agent only when the problem truly requires the next level. |
Before building an agent, write a design blueprint. This prevents the common mistake of connecting an LLM to tools without a clear boundary. The blueprint should define goal, users, tools, permissions, memory, workflow, risks, and evaluation approach.
|
Design Area |
Questions to Answer |
|
Goal |
What exact task should the agent complete? |
|
Users |
Who will use it and what roles do they have? |
|
Tools |
Which tools are required? Which are read-only? Which are write-enabled? |
|
Permissions |
Who can access which data and actions? |
|
Approval |
Which actions require human approval? |
|
Memory |
What should be remembered, for how long, and why? |
|
Limits |
What are max steps, max cost, timeout, retry limit? |
|
Monitoring |
What logs and metrics are required? |
|
Evaluation |
How will quality and safety be tested? |
|
Fallback |
What happens when the agent is unsure or tool fails? |
Simple agent pseudo-code:
state = initialize_task(user_goal)
while not state.done:
if state.step_count > MAX_STEPS:
return "I could not complete this safely within the step limit."
plan = planner.decide_next_step(state)
if plan.requires_approval:
approval = request_human_approval(plan)
if not approval.approved:
return "Action cancelled by reviewer."
tool_result = executor.run(plan.tool, plan.arguments)
state.add_observation(tool_result)
if tool_result.error:
state.handle_error(tool_result)
return create_final_response(state)
A production agent must be monitored like any other serious software system, but with additional AI-specific metrics. You need to know what tools were called, why they were called, what data was accessed, how much the request cost, how long it took, whether the answer was correct, and whether the user was satisfied.
|
Metric |
Why It Matters |
|
Task success rate |
Measures whether the agent completes the intended job |
|
Tool call count |
Detects unnecessary or looping behaviour |
|
Tool error rate |
Shows integration problems |
|
Approval rejection rate |
Shows unsafe or low-quality proposed actions |
|
Latency |
Measures user experience |
|
Cost per task |
Controls budget |
|
Escalation rate |
Shows how often human help is needed |
|
User satisfaction |
Captures practical usefulness |
|
Policy violation count |
Measures safety and compliance |
This mini project is designed for beginners. The first version should not send emails or update real systems. It should only classify tickets, search a small knowledge base, suggest next steps, and create a draft response.
|
Component |
Beginner Implementation |
|
User interface |
Simple web page, notebook, or command-line input |
|
Ticket input |
Sample CSV file with ticket text |
|
Knowledge base |
Small set of troubleshooting notes |
|
Tool 1 |
Search knowledge base by keyword or semantic search |
|
Tool 2 |
Classify ticket category |
|
LLM |
Generate suggested resolution and draft response |
|
Approval |
User manually approves or edits draft |
|
Logging |
Save ticket, category, suggested answer, and status |
Example input ticket:
"I reset my password today and now VPN is not connecting from home."
Expected agent output:
Category: Network / VPN
Possible cause: VPN credentials or cached profile after password reset
Suggested steps:
1. Ask user to sign out and sign in again.
2. Clear saved VPN password.
3. Reconnect using new password.
4. If issue persists, escalate to Network team.
Draft reply: ...
Confidence: Medium
Human approval required: Yes before sending
An AI agent is an LLM-powered system that can plan, use tools, observe results, and continue working toward a goal. Agents are different from ordinary chatbots because they can take action and manage multi-step workflows. A good agent architecture includes an orchestrator, planner, executor, tool registry, memory, permission system, human approval gate, logging, monitoring, and guardrails.
Tool calling and function calling allow agents to interact with external systems in a structured and controlled way. Memory helps agents maintain context, but it must be used carefully. Single-agent systems are easier to start with, while multi-agent systems are useful only when specialized roles are truly needed. Production agents must be designed with safety, permissions, approval, cost control, and monitoring from the beginning.
|
Term |
Meaning |
|
AI Agent |
An AI system that uses an LLM to plan, call tools, observe results, and complete tasks. |
|
Tool Calling |
Allowing the agent to use external tools such as search, database, email, files, and APIs. |
|
Function Calling |
Structured tool use where the LLM selects a function and provides validated arguments. |
|
Planner |
The part of the agent that decides the next step. |
|
Executor |
The part that runs the selected action or tool. |
|
Agent Loop |
The repeated cycle of plan, act, observe, and continue. |
|
Short-Term Memory |
Temporary context for the current task or conversation. |
|
Long-Term Memory |
Persistent useful information kept across sessions. |
|
Human Approval Gate |
A checkpoint where a person must approve risky actions. |
|
Tool Registry |
A controlled list of tools available to an agent. |
|
Least Privilege |
Giving the agent only the minimum required access. |
|
Guardrails |
Rules, checks, and controls that keep the agent safe and reliable. |
After completing this chapter, you should be able to explain what AI agents are, how they differ from chatbots, how tool calling works, why planner and executor roles matter, how memory is used, how to design safe permissions, and when agents are appropriate. You should also be able to design a basic agent architecture for travel planning, IT support, sales research, or similar business workflows.
|
Next Chapter Chapter 15 will compare prompting, RAG, and fine-tuning. You will learn when each approach is useful, how they differ in cost and complexity, and how to choose the right method for a business use case. |