What Can AI Agents Actually Do Today? From Coding to Business Workflows

AI agents can already write code, search across files, use software, produce reports, and move parts of a business process forward.

But that sentence can easily sound more impressive than reality.

The harder question is:

Which kinds of work can an agent complete reliably enough that you would actually delegate the job?

That is where the picture becomes more useful.

AI agents are strongest today when the work is digital, the next result is visible, success can be checked, and mistakes can be recovered from.

Quick Answer: What Can AI Agents Actually Do Today?

As of October 2026, useful agents can already handle meaningful work in several areas:

  • software development,
  • research and knowledge work,
  • web and desktop software through computer use—AI seeing a screen and operating a mouse and keyboard,
  • and bounded business workflows.

A bounded workflow is simply a process with a defined goal, known tools, clear permissions, and limits on what the system may change.

The important word is not autonomous.

It is bounded.

Agents are much more useful when they know where to work, what tools they can use, what counts as success, and when to stop or ask for help.

Original Asset 1: The Four Conditions of Agent-Friendly Work

Instead of asking whether an industry is “ready for agents,” ask whether the task has four properties.

Condition Plain-language question Why it helps an agent
Digital Can the agent reach the files, software, or tools it needs? The work can actually be acted on.
Observable Can the agent see what happened after each step? A failed action becomes feedback instead of a hidden mistake.
Verifiable Can the system check whether the job is really done? It can distinguish “I tried” from “I finished.”
Recoverable Can a mistake be retried, reversed, or sent to a person? The cost of an imperfect decision stays contained.

This explains why coding has become such a strong agent use case—and why a vague, high-stakes business decision with no clear feedback is still much harder.

1. Coding Agents Can Now Do Real Multi-Step Development Work

A codebase is the collection of source files, tests, configuration, and documentation that make up a software project.

That environment is unusually friendly to agents.

The agent can read files, edit code, run commands, execute tests, inspect errors, and try again.

Read → Edit → Test → Inspect → Fix → Test again

The key is feedback.

A failing test tells the agent something concrete. Version control and checkpoints make many changes reversible.

OpenAI reports that Codex is increasingly used for longer delegated tasks, and in May 2026 more than 70% of individual Codex users asked it to complete at least one task estimated to require more than an hour of human work.[1]

Anthropic has also shown Claude Code being used for large modernization projects, including legacy-code migrations, while its current tooling supports subagents, background tasks, project rules, and integrations with development systems.[2][3]

That does not mean “give the agent a repository and walk away.”

Large projects still depend heavily on good tests, clear project instructions, review, and a way to recover from bad changes.

2. Agents Are Moving From Coding Into Knowledge Work

Knowledge work means work whose main input is information—documents, data, research, analysis, writing, planning, or decisions.

Here, agents are increasingly doing more than drafting one answer.

They can gather information from several places, analyze files, run code, create a spreadsheet or report, revise the work after feedback, and produce a finished deliverable.

OpenAI reported in June 2026 that knowledge workers were using Codex for reports, spreadsheets, presentations, contracts, research, data analysis, and lightweight internal tools.[4]

The newer shift is toward longer-running infrastructure.

In September 2026, OpenAI introduced the Agents API, a managed service for cloud agents that can keep working over long periods, use files and code, preserve intermediate results, and coordinate additional agents.[5]

A harness is the software around the AI model that manages things such as tools, context, saved progress, retries, and execution.

That term matters because long-running work is not only a model problem.

The surrounding harness helps the agent remember what it has done and continue after interruptions.

3. Computer-Use Agents Can Work Through Software Interfaces

Many older business systems do not expose an API—a direct software-to-software connection that lets one program reliably ask another program to do something.

Humans still click through websites, desktop apps, dashboards, and forms.

Computer-use agents try to bridge that gap by looking at the graphical interface and operating it much like a person would.

Google integrated computer use directly into Gemini 3.5 Flash in June 2026 for browser, mobile, and desktop environments.[6]

Microsoft's Copilot Studio computer-use feature is also generally available and can operate websites and Windows applications using a virtual mouse and keyboard.[7]

This expands the reachable work.

It does not make graphical interfaces as reliable as APIs.

A page can load slowly. A dialog can appear unexpectedly. A button can move. An agent can think an action succeeded when the screen is actually in a different state.

That is why browser and desktop automation often needs explicit verification after important steps.

4. Business Workflows Are Becoming a Hybrid of Rules and Agent Judgment

A workflow is a repeatable sequence of steps used to move a business task from start to finish.

Examples include:

  • handling an invoice,
  • routing a customer request,
  • checking an order,
  • preparing an onboarding package,
  • or moving a document through review and approval.

The practical pattern is not usually “let the agent invent the whole process.”

It is more often:

Deterministic workflow → Agent judgment where needed → Deterministic workflow

Deterministic here means the step follows a fixed rule: if the same input arrives, the system should take the same defined action.

Microsoft made its newer Copilot Studio harness generally available in August 2026 for more complex business processes with many steps, multiple information sources, and ambiguous decision points.[8]

Its current platform also lets workflow steps call agents for reasoning while keeping approvals, routing, and known process logic in the surrounding workflow.[9]

This hybrid design is important.

Routine steps do not need an AI to rethink them every time.

The agent is most valuable where the process reaches a point that needs interpretation, flexible research, or a decision among several valid next steps.

5. Long-Running Agents Are Becoming More Practical—but “Long-Running” Needs a Definition

A long-running task is a job that takes many steps or continues across a meaningful period rather than finishing in one short interaction.

That might mean an hour-long software task, an overnight research job, or a workflow that pauses while waiting for new information.

The infrastructure for this has improved quickly.

OpenAI's Agents API is explicitly designed for agents that can continue for days.[5]

Anthropic's research on real Claude Code usage found that among the longest-running sessions, the time Claude worked before stopping had nearly doubled over a three-month period, from under 25 minutes to over 45 minutes.[10]

But duration alone is not the achievement.

A system that runs for eight hours while repeating itself is not more useful than one that completes a task in twenty minutes.

The real question is whether it can maintain state—the record of what has happened, what remains, and what constraints still apply—and keep making measurable progress.

Original Asset 2: The Delegation Ladder

Agent capability is easier to understand as a ladder of responsibility.

Suggest → Act → Recover → Complete

Suggest

The AI drafts or recommends. A person still performs the action.

Act

The AI uses a tool, edits a file, runs code, or operates software.

Recover

The AI notices a failure, reads the result, changes its approach, and tries again.

Complete

The AI maintains state across several steps, verifies the outcome, and either finishes or escalates an exception.

An exception is simply a case that falls outside the normal path and needs different handling.

This ladder is more useful than calling every product “autonomous.”

Two systems may use the same model, yet one only suggests and the other can recover from failures and finish a bounded job.

What Still Breaks Today?

Reader discussions about agents repeatedly converge on one issue: reliability becomes harder as the task gets longer and the environment gets messier.

Several failure modes matter.

1. The goal is underspecified

If “done” is vague, the agent has no reliable stopping condition.

2. The environment changes underneath it

Web pages, files, permissions, data, APIs, and software sessions can change while the agent is working.

3. One bad step propagates

A wrong assumption early in a long task can contaminate later steps.

4. The agent cannot verify success

A click happened, but the order was not submitted. A file was created, but in the wrong folder. An email draft exists, but it was never sent.

5. Human rescue becomes the hidden cost

If people constantly have to watch, correct, restart, or interpret the agent's failures, the automation may be less valuable than the demo suggests.

This is why agent evaluation increasingly looks at the whole trajectory—tool calls, state changes, errors, recovery, and final result—not just the final sentence.[11]

Computer Use Shows the Reliability Problem Clearly

Computer use is impressive because it lets agents reach software that was previously difficult to automate.

It is also a good example of why access is not the same as reliability.

Developers discussing production browser automation often describe graphical interfaces as slower and more fragile than structured software connections, especially when layouts change or the system cannot prove that a step succeeded.

A stronger pattern is:

Act → Check an explicit post-condition → Continue

A post-condition is a simple statement of what must be true after an action—for example, “the confirmation page is visible” or “the record now has status = approved.”

That small idea turns “click and hope” into a more testable workflow.

Original Asset 3: Measure Cost per Successful Outcome, Not Cost per AI Call

An AI call can be cheap while the workflow is expensive.

Why?

Because the real cost also includes human review, failed attempts, retries, and recovery.

A better mental model is:

Cost per successful outcome = agent/tool cost + human review + exception recovery

This is not an accounting standard.

It is a way to avoid a common mistake: celebrating low token cost while ignoring the people who still have to rescue the workflow.

For a business process, useful metrics are often closer to:

  • time saved per completed workflow,
  • successful completion rate,
  • human minutes per exception,
  • error or rework rate,
  • and total cycle time from request to finished result.

That is a much harder test than “the agent produced a plausible response.”

When Should You Use an Agent Today?

Use an agent when the path can change and the environment provides feedback.

The task is especially promising when:

  1. the work is already digital,
  2. the agent has clear tools or software access,
  3. intermediate results are visible,
  4. success can be checked,
  5. and mistakes can be reversed, retried, or escalated.

Coding often meets all five.

Research and document work can meet many of them.

Business workflows can work well when the process is already understood and the agent is inserted only where flexible judgment helps.

When Is Ordinary Automation Better?

Not every repetitive task needs an agent.

If the process is stable and rules are known, ordinary automation can be cheaper and easier to test.

For example:

Invoice arrives → extract fixed fields → save record → send confirmation

If no judgment is required, a conventional script or workflow may be the better design.

Likewise, an agent is a poor fit when:

  • the goal is vague,
  • success cannot be verified,
  • mistakes are hard to reverse,
  • the underlying data is unreliable,
  • or every exception requires a human anyway.

AI does not repair a broken process merely by being inserted into it.

What Is Still Genuinely Open?

How much messy office work can computer-use agents handle reliably?

Access is improving quickly, but changing interfaces, pop-ups, authentication, and hidden application state still make visual automation harder than clean APIs.

How long can an agent work before human intervention becomes the bottleneck?

Infrastructure can now keep agents running much longer, but sustained useful progress is harder than simply keeping a process alive.

Which business workflows produce durable ROI?

Simple, bounded processes often show value first. Broad “AI employee” claims are much harder to evaluate because review, exceptions, and data quality can dominate economics.

When do multiple agents help rather than add coordination cost?

Specialization and parallel work can help, but shared state, duplicate work, conflicting actions, and review overhead can offset the gain.

Will agents shift from completing tasks to managing ongoing responsibilities?

That is already beginning in long-running and always-on systems, but identity, permissions, monitoring, and accountability become more important as the time horizon expands.

The Bigger Lesson

AI agents can do real work today.

But the useful boundary is not “can the model reason?”

It is whether the whole system can reach the work, see what happened, check success, and recover when something goes wrong.

Coding works well because the environment already supplies files, commands, tests, and reversible changes.

Knowledge work is becoming more practical as agents gain files, code execution, and long-running state.

Computer use expands reach into older software, but reliability still needs careful verification.

Business workflows work best when fixed process logic and flexible agent judgment are combined rather than confused.

Do not ask only, “Can an AI agent do this task?” Ask, “Can it complete the task, prove that it worked, recover from exceptions, and still save enough human effort to matter?”

Key Terms

bounded workflow
A process with a defined goal, known tools, permissions, and limits on what the system may change.

codebase
The files, tests, configuration, and documentation that make up a software project.

computer use
AI operating software through a visual interface by seeing the screen and controlling mouse or keyboard actions.

API
A direct software-to-software connection that lets one program request data or actions from another program.

harness
The software around an AI model that manages context, tools, state, execution, retries, and other runtime behavior.

state
A record of what has happened in a task, what remains unfinished, and what constraints still apply.

exception
A case that falls outside the normal process and needs different handling or human judgment.

post-condition
A fact that should be true after an action if the action really succeeded.

Related Reading

Sources

  1. OpenAI — How Agents Are Transforming Work, June 25, 2026.
  2. Anthropic — Modernizing Legacy Code with Claude Code, September 24, 2026.
  3. Anthropic — Claude Code: Foundations, July 8, 2026.
  4. OpenAI — Codex Is Becoming a Productivity Tool for Everyone, June 2, 2026.
  5. OpenAI — Introducing the Agents API, September 10, 2026.
  6. Google — Introducing Computer Use in Gemini 3.5 Flash, June 24, 2026.
  7. Microsoft Learn — Automate Web and Desktop Apps with Computer Use, checked October 3, 2026.
  8. Microsoft — More Powerful Agents and Workflows for Autonomous Business Processes, August 3, 2026.
  9. Microsoft Learn — Copilot Studio AI at Work Roadmap, checked October 3, 2026.
  10. Anthropic — Measuring AI Agent Autonomy in Practice, February 18, 2026.
  11. Anthropic — Demystifying Evals for AI Agents, January 9, 2026.
Update History
  • October 3, 2026 — Major rebuild with current long-running agent infrastructure, computer-use and workflow updates, reliability limits, and new task-fit and outcome-cost frameworks.
  • August 22, 2026 — First published.