In July 2025, SaaStr founder Jason Lemkin was building an app with Replit’s AI coding agent when the agent deleted his live database, despite a code freeze meant to prevent exactly that. The database held records on more than 1,200 executives and roughly 1,190 companies. Afterwards the agent described its own behaviour as “a catastrophic failure on my part,” according to Fortune. Replit’s CEO called it unacceptable and the company added new safeguards.
That episode captures where AI agents stand. They can now do real work: write and run code, operate a web browser, sort through files and fill in forms. But “doing real work” also means being able to do real damage, and the technology for keeping agents on a short leash is still catching up.
Here’s what the term actually means, how agents work under the hood, which products have shipped, and why the honest answer to “can I trust it?” is still “sometimes.”
Chatbot versus agent: the practical difference
A chatbot answers. An agent acts. When you ask ChatGPT or Claude a question in a plain chat window, the model reads your words and writes a reply. That’s one step, and you decide what happens next.
An agent is given a goal and then decides its own next steps, often many of them, using tools along the way. Anthropic drew a useful line in a widely read December 2024 engineering post. It separated “workflows,” where a language model and its tools follow paths that a programmer has hard-coded, from “agents,” where the model itself directs the process and chooses which tools to use. A workflow is a recipe. An agent is a cook who’s been told what’s for dinner and left alone in the kitchen.
Most products marketed as agents today sit somewhere between those two poles. Anthropic’s own advice in that post was to start with the simplest approach that works and add autonomy only when it measurably improves results, because agents cost more and their errors can compound.
How an agent uses tools
Underneath, the mechanics are less magical than the marketing suggests. A large language model can only produce text. What turns it into an agent is a loop wrapped around it:
- The model receives a goal, plus a description of the tools it can call: search the web, read a file, run a command, send an email.
- Instead of answering directly, the model writes out a structured request to use one of those tools.
- Ordinary software outside the model actually executes that request and feeds the result back in as more text.
- The model reads the result, decides what to do next, and the loop repeats until it judges the task done or hits a limit.
The “tools” can be almost anything with an interface. Some agents look at screenshots and move a virtual mouse. Anthropic introduced this as “computer use” in October 2024, calling it experimental and at times error-prone. Its model scored 14.9 per cent on the OSWorld computer-navigation benchmark at the time, which was roughly double the next-best system but nowhere near human performance.
The plumbing: Model Context Protocol
Connecting an agent to every service a business uses used to mean custom code for each one. In November 2024 Anthropic released the Model Context Protocol (MCP), an open standard for linking AI assistants to the places data lives, from file stores to business apps. It caught on quickly. By December 2025, when MCP was handed to a new Agentic AI Foundation under the Linux Foundation, the project reported more than 97 million monthly SDK downloads and around 10,000 active servers, with support in ChatGPT, Claude, Gemini, Microsoft Copilot and others, according to the official MCP blog. Anthropic, Block and OpenAI co-founded that foundation.
Think of MCP as a universal adapter. It doesn’t make agents smarter, but it makes it far easier to hand them the keys to your email, calendar and file drives. As we’ll see, that convenience cuts both ways.

What actually shipped, 2024 to 2026
The pace of product launches has been relentless. A partial timeline of agent products from the big labs and one prominent Canadian company:
| Date | Product | What it does |
|---|---|---|
| Oct 2024 | Anthropic computer use | Lets Claude operate a desktop via screenshots, clicks and typing (developer beta) |
| Jan 2025 | OpenAI Operator | Browser agent for forms, orders and bookings; U.S. Pro users only at first |
| May 2025 | Claude Code (general release) | Coding agent that reads, edits and runs code in a developer’s project |
| Jul 2025 | ChatGPT agent | Successor to Operator, working in a virtual computer with browser and terminal |
| Aug 2025 | Cohere North (wide release) | Enterprise agent platform running inside a company’s own environment |
| Oct 2025 | ChatGPT Atlas | OpenAI’s AI web browser, able to act on pages for the user |
| Jan 2026 | Claude Cowork | Desktop agent that works on files in folders the user grants access to |
| Jul 2026 | ChatGPT Work | Agent for longer workplace tasks across connected apps and documents |
A few of these deserve a closer look. OpenAI’s ChatGPT agent, launched July 17, 2025, gave the model a virtual computer with a visual browser, a text browser, a terminal and connectors to services like Gmail and GitHub. It was designed to ask permission before consequential actions such as making a purchase, and to let users take over the browser at any point. OpenAI retired the agent in August 2026 and pointed users to ChatGPT Work, launched July 9, 2026, which breaks longer jobs into steps and produces documents, spreadsheets and presentations while the user watches and redirects.
Anthropic’s Cowork, released as a research preview in January 2026, took the approach of its Claude Code programming agent and pointed it at ordinary office tasks: turning a folder of receipt photos into an expense spreadsheet, for example. Fortune reported that Anthropic openly warned about prompt injection risks at launch.
The Canadian angle
Toronto’s Cohere has bet on a quieter, enterprise-only version of the agent story. Its North platform, which got a wider release in August 2025, is built to automate routine back-office work while keeping data inside the customer’s own systems. RBC, Bell and Dell were among early users, as The Canadian Press reported. For banks and telecoms handling regulated data, “the agent never leaves our servers” is a stronger selling point than raw capability.
Where agents still fail
Reliability falls off a cliff on long tasks
The research group METR has done some of the most useful work on measuring agents. Its approach is to time how long tasks take skilled humans, then check which tasks AI agents can complete. In a March 2025 study, METR found that models succeeded almost every time on tasks taking humans under four minutes, but less than 10 per cent of the time on tasks taking more than about four hours. The length of task an AI could finish half the time had been doubling roughly every seven months for six years.
That trend has continued. By early 2026, METR’s measurements put leading models at roughly six to twelve hours for 50 per cent success on its software tasks, but closer to one hour when you demand 80 per cent success, according to a summary of METR’s published results. That gap matters. A colleague who completes a day’s work half the time is not someone you’d leave unsupervised.
Prompt injection: the unsolved security problem
The biggest worry among security researchers isn’t that agents make honest mistakes. It’s that they can be tricked. A language model can’t reliably tell the difference between instructions from its user and instructions hidden in a web page, email or document it’s reading. Plant the right text in a place the agent will look, and it may follow those orders instead of yours.
Developer Simon Willison has described a “lethal trifecta” that makes this dangerous. An agent becomes exploitable when it has all three of the following:
- access to private data, such as your inbox or company files;
- exposure to untrusted content, like web pages or incoming email;
- a way to communicate externally, such as sending a message or loading a URL.
This isn’t theoretical. In June 2025 researchers at Aim Security disclosed “EchoLeak,” a flaw in Microsoft 365 Copilot that let a specially crafted email cause Copilot to leak sensitive data with no clicks from the victim. Microsoft rated it critical and patched it, and said there was no evidence of exploitation in the wild, as The Hacker News reported.
The vendors themselves are candid about the limits. In December 2025 OpenAI wrote that prompt injection, like scams and social engineering, is “unlikely to ever be fully ‘solved,'” as TechCrunch reported. The U.K.’s National Cyber Security Centre went further in a blog post that month, arguing that inside an LLM there is no real distinction between data and instructions, so organizations should focus on limiting the damage rather than expecting a fix.
Agent washing and cancelled projects
Then there’s the hype problem. In June 2025 Gartner predicted that more than 40 per cent of agentic AI projects would be cancelled by the end of 2027 because of rising costs, unclear value or weak risk controls. The firm also complained of “agent washing,” estimating that only about 130 of the thousands of vendors claiming agentic products offered the real thing. Rebranded chatbots and old robotic process automation tools are being sold with a new label.
Using agents without getting burned
If you’re experimenting with agents at work, the advice from vendors and security agencies converges on a few habits:
- Narrow the job. Give the agent a specific task, not open-ended access to your inbox with a vague instruction.
- Keep a human on consequential steps. Require confirmation before anything is sent, paid, deleted or published.
- Break the trifecta. If an agent reads untrusted content, don’t also give it your private data and a way to send things out.
- Back up first. Anything an agent can write to, it can also overwrite. Separate test and production systems, as Replit did after its incident.
- Log everything. You want a record of what the agent did and why.
What to expect next
The capability curve is steep and shows little sign of flattening. Gartner expects that by 2028 at least 15 per cent of day-to-day work decisions will be made autonomously by agents and a third of enterprise software will include agentic features, up from almost nothing in 2024. Standards like MCP mean connecting agents to business systems will keep getting easier.
Our view is that the next two years will be less about agents getting smarter and more about the boring infrastructure around them: permissions, audit trails, sandboxes and insurance. The agents that succeed in workplaces will be the ones that are dependable on narrow tasks, not the ones that can do anything once in a while.
The bottom line
An AI agent is a language model in a loop with tools. That simple design is surprisingly powerful, and in 2026 it’s good enough to save real time on coding, research and paperwork. It is not yet good enough to be trusted with anything you can’t afford to lose. Treat agents like a talented new hire on their first week: give them clear tasks, limited access and someone checking their work.
Sources and further reading
- Anthropic: Building effective agents
- Anthropic: Introducing computer use
- Anthropic: Introducing the Model Context Protocol
- MCP blog: MCP joins the Agentic AI Foundation
- Wikipedia: Claude (product timeline)
- OpenAI: Introducing ChatGPT agent
- Android Authority: ChatGPT Work launch
- Fortune: Anthropic launches Claude Cowork
- The Canadian Press via CP24: Cohere’s North gets wide release
- Fortune: Replit agent wipes database
- METR: Measuring AI ability to complete long tasks
- Wikipedia: OpenAI Operator
- Wikipedia: METR (time horizon results)
- Simon Willison: The lethal trifecta
- The Hacker News: EchoLeak in Microsoft 365 Copilot
- TechCrunch: OpenAI on prompt injection in AI browsers
- UK NCSC: Prompt injection is not SQL injection
- Gartner: Over 40% of agentic AI projects will be cancelled by 2027
Leave a comment