A paper-cut diorama emerging from an open book, showing a fully realized lit city with streets and people, representing the difference between describing an action in words and an agent that actually goes and does it
AI & Trust

The Real Difference Between a Chatbot and an Agent

One of these answers a question. The other decides what to do, does it, checks its own work, and moves on to the next step — without waiting for you to ask. That difference isn't academic. It changes what can go wrong and who's supposed to catch it.

FabricLoop Editorial
1,850 words
8 min read

A chatbot takes what you type, generates a response, and stops. An agent takes what you type, decides what needs to happen, does something about it, checks whether that worked, and decides what to do next — on its own, often across many steps — before a person sees any of it. That's the entire distinction. Everything people argue about when they argue about "AI agents" — the risk, the oversight a system needs, and most of the marketing confusion — follows from that one difference.

What a chatbot actually does

A chatbot is a single-pass system. You give it text, it generates text back, and the interaction ends there. Even a chatbot with a long memory of your conversation is still doing one thing per turn: reading everything said so far and predicting the next message. It never queries a database to check a fact, never sends anything on your behalf, and never comes back later to see whether its answer held up. If it's wrong, the damage is a sentence a person reads — and can, in the ordinary course of things, catch before acting on it.

Most of what people ask chatbots to do fits this shape without anyone noticing: summarize this document, draft a birthday message, explain our refund policy, write three headline options. None of these require the system to check anything against the real world or take an action outside the chat window. That's still true when the interface is branded "AI assistant" or "copilot" instead of "chatbot" — the label on the box doesn't change what's happening inside it.

What an agent actually does

An agent runs a loop, not a single pass: plan a step, take an action by calling a real tool — search a database, send a message, edit a file, hit an API — look at what that action actually returned, and use that result to decide the next step. It repeats this until the task is done, it gets stuck, or it's built to check in with a person. The important part is that no one approves each individual step along the way. The system decides, on its own, what to try next based on what actually happened the last time it acted — and it can be wrong at every one of those decision points, not just in a final answer.

This loop isn't new or exotic. Researchers have described versions of it — reasoning about what to do, acting, observing the result, reasoning again — for years, and it's what's actually running under products that call themselves agents, from software that files expense reports to coding tools that open a terminal and run their own commands. What makes something an agent instead of a very talkative chatbot is that it acts on the world, watches what happened, and adjusts — repeatedly, without a human approving each move.

The same request, run two ways

Here's what that difference looks like when two systems get handed instructions that sound similar.

Chatbot

"Summarize this document."

  • 1Reads the text you pasted in.
  • 2Generates a summary paragraph.
Stops here — nothing outside the chat changed
Agent

"Find the three open invoices over 30 days past due, draft a reminder email for each, and put them in my drafts folder."

  • 1Queries the invoicing system, filters for invoices open more than 30 days.
  • 2Checks it actually found three, not two or five — flags the mismatch instead of guessing.
  • 3Pulls the right amount, due date, and contact for each invoice and drafts a reminder.
  • 4Saves each draft to the real drafts folder through the mail tool.
  • 5Reports back what it found and what it drafted.
Takes action — three real drafts now exist, unreviewed

Ask a chatbot the second question and it will still hand you something that looks like an answer: three plausible-sounding reminder emails, generated from whatever you happened to paste into the conversation. What it won't do is query your actual invoicing system, verify the count, or put anything in a real drafts folder. The output can look similar. What the system actually did is not.

Why this isn't just a semantic argument

The distinction matters because it changes what can go wrong, and who catches it. A chatbot's worst case is a wrong answer. Someone reads it, and in the normal course of things, either catches the error or decides not to act on it — the mistake never leaves the conversation. An agent's worst case is a wrong action already taken in the real world: the reminder that went to the wrong customer with the wrong balance, the record updated with the wrong value, the refund issued twice — before anyone reviewed anything. The mistake isn't a sentence anymore. It's an event, and events don't un-happen.

A chatbot's worst case is a wrong answer someone reads. An agent's worst case is a wrong action already taken — before anyone read anything at all.

That's why an agent needs a different kind of oversight than a chatbot does. A chatbot mostly needs someone checking its answers, whenever they get around to it. An agent needs its designers to have already decided, before it ever runs, which actions it can take without asking, which ones require a person to see the plan first, and what happens when it gets stuck. Work that out after the fact, and you find out what the agent already did the hard way.

You can only trust what you can see

This is the same idea behind FabricLoop's Legibility concept: you can only govern access, and behavior, that you can actually see. For a chatbot, that's nearly automatic — its entire output is a message a person reads, so the action and the record of the action are the same thing. For an agent, they're not. Its actions happen inside other systems — a CRM, an inbox, a database, a file — and unless something logs what it touched, changed, or sent, there's no way to review it after the fact, let alone stop it before. Legibility isn't a compliance nice-to-have layered on top of an agent. For an agent, it's the whole question, because its "answer" isn't a sentence you can proofread — it's a set of actions you may never know happened unless the system was built to show you.

That's the practical link between the two ideas: an agent takes on more, and different, risk than a chatbot, which is exactly why it needs a visible trail of what it did and, in the higher-stakes cases, a checkpoint before it acts — what FabricLoop's Human Intervention Rate concept treats as a number you design and measure, not an afterthought you bolt on once something has already gone wrong.

The market prices this backward, constantly

Once you have the real test, it's obvious how often the label lies in both directions. Plenty of products marketed hard as "AI agents" — the word in the hero headline, in the pricing tiers — are, underneath, a single well-tuned prompt: read the input, generate the output, done. No independent tool call, no loop, no decision made without a human approving the next click. Meanwhile, plenty of software that never uses the word "agent" — an automated invoicing workflow, a monitoring system that reroutes traffic on its own, an ops script that restarts a failed service and checks whether that fixed it — is quietly running the exact loop described above. The word on the label tells you nothing reliable about which machine you're actually using.

The two-question test

1. Does it take more than one step to do the thing? 2. Does it decide what that next step is, or does a person decide every step, one click at a time? If a human is choosing every step, you're looking at a chatbot with extra buttons — call it whatever the marketing page calls it. If the system is choosing its own next step across more than one step, you're looking at an agent, and it needs to be governed like one: visible logs, defined limits on what it can do without asking, and a real answer to who reviews it and when.

FL
Where FabricLoop draws this line

FabricLoop's own AI product, Loop Agent, runs that loop — searching, drafting, checking its own work across MCP-connected tools — but it's built to show its work and ask before the higher-stakes steps, not act silently and report after the fact. Every MCP connection it uses is scoped to a person, and on Enterprise, what it did shows up in an audit log rather than living only inside its own memory of the conversation.

That's the same logic covered in Legibility and Human Intervention Rate — worth reading next if this is the first of these three ideas that's clicked.

None of this requires a technical background to apply. The next time a vendor, or a colleague, calls something an agent, ask what it actually did between your instruction and the result — and how many of those steps it decided on its own.


Key takeaways
01
A chatbot answers in one pass: read input, generate output, stop. It never checks a fact against a live system, sends anything on your behalf, or comes back later to verify its own answer.
02
An agent runs a loop: plan, act through a real tool, observe what actually happened, decide the next step, repeat — without a person approving each individual move.
03
Similar-sounding requests can produce two different machines. "Summarize this document" is chatbot behavior. "Find the three invoices over 30 days past due, draft reminders, and put them in my drafts folder" is agent behavior — it requires real lookups, verification, and independent decisions between steps.
04
A chatbot's worst case is a wrong answer a person reads and can choose not to act on. An agent's worst case is a wrong action already taken in the real world — a sent email, a changed record — before anyone reviewed it.
05
Because agents act before review, they need oversight designed in advance: which actions require a person to see the plan first, which don't, and what happens when the agent gets stuck. Working that out after an incident is too late.
06
Legibility — the ability to see who did what, and how — matters more for agents than chatbots, because an agent's actions happen inside other systems, not inside a message you can read. No visible trail means no real oversight.
07
The word "agent" on a pricing page isn't a reliable signal. Apply the test yourself: does it take more than one step, and does it decide the next step on its own? If a human chooses every step, it's a chatbot no matter what the marketing calls it.