AI agents can already carry out long chains of work, but that does not mean they should run a business on their own. The Atria Dawn Preview project shows a clear pattern: the agent performs more actions, while the human still chooses the goal, the method, and the boundaries of the task.
For a business owner, the lesson is practical. A good AI agent is not an autopilot replacing the owner; it is a working partner. It removes routine, prepares options, tests ideas, and revises output, but the key decisions should stay with the person who understands the business context.
What did the Atria Dawn Preview project show?
A team involving researchers from Fudan University studied its own process of building a new language model called Atria Dawn Preview. The analysis covered more than 700 task logs from 56 participants, together with logs from the agents they used.
The model is based on a mixture-of-experts architecture with 744 billion parameters and is designed for research and engineering tasks. The project used a pipeline where each task is tied to a real execution environment: the model calls tools, produces intermediate results, and is checked against external signals such as tests, metrics, or source evidence.
The team says Atria Dawn Preview leads on five of 16 benchmarks, including web search and cybersecurity, although it does not have an overall edge over competitors. For a business, the ranking is less important than the way people and agents divided the work.
AI was used in 96.5 percent of the reviewed tasks. Over four weeks, the median ratio of agent actions to human inputs rose from 11 to 28.5. The researchers warn against reading this as growing autonomy. One human decision simply triggered more agent steps.
AI proposes, the human chooses
The central finding is that agents became more active in proposing and executing, but they did not become the main decision-makers. For methods and parameters, the most common pattern was “AI proposes, human selects.” It accounted for 55.4 percent of cases.
Overall, humans made 85.5 percent of decisions about methods and parameters, while AI made only 9.2 percent. Humans made the final decision on goals and scope in 93.4 percent of cases. Even among tasks that participants rated as infeasible without AI, the human chose the goal 95.4 percent of the time.
This is an important correction to the usual image of autonomous agents. From the outside, it may look as if an agent that writes, searches, runs tools, and fixes results is leading the project. Inside the workflow, the picture is different: the agent expands the set of options and handles execution, but the human sets the frame.
- AI is useful for options, drafts, checks, and revisions.
- A human is needed to choose the goal, judge the result, and stop unnecessary work.
- The longer the agent’s chain of actions, the more important it is to define its authority in advance.
- Autonomous mode is convenient, but it should not be enabled only to avoid frequent approvals.
What does this mean for a business with its own AI agent?
If a company already has a personal AI agent, or is considering one, the main question is not “how smart is it?” The better question is: which decisions may the agent make on its own, and where must it stop and ask a human?
In the study, 151 of 455 completed AI-assisted tasks were rated by participants as infeasible without AI, roughly a third. These tasks were spread across 27 of the 56 participants, not concentrated among a few power users. AI was not only speeding up existing work; it made some work possible that would not otherwise have been attempted.
That is especially relevant for small businesses. An agent can help sort files, prepare customer replies, collect material for a proposal, spot inconsistencies in data, remember tasks, and improve a draft. But the owner or responsible employee should decide what the task is trying to achieve, what data may be used, and what counts as a correct result.
In practice, launching an agent requires working rules, not just attractive scenarios:
- describe typical tasks the agent may complete without extra approval;
- list the actions that require human confirmation;
- assign a person responsible for goals, access, and output quality;
- keep a history of tasks and decisions so the reasoning path is visible;
- review the rules when the agent receives new files, tools, or areas of work.
This is the right mindset for a personal agent in Telegram with task memory and access to working data. At NekoAgent, we follow the same principle: an agent should reduce the human workload, not replace the process owner where accountability is required.
Why oversight must not become a rubber stamp
The researchers highlight the risk of rubber-stamping. When each decision rests on a long chain of agent work that no human can fully review, oversight becomes fragile. In the worst case, a person sees only the final output and approves it without understanding what happened inside.
In the project, many participants ran agents in autonomous modes to avoid interrupting long runs with constant approvals. In those cases, the boundary of authority was drawn for convenience, not through a deliberate decision about how much power AI should have.
This is a very practical warning for businesses. If an agent works with customer messages, files, spreadsheets, or internal knowledge, convenience should not be the only criterion. The company needs to decide in advance where speed matters more than review, and where a mistake would be too costly.
Good oversight does not require reading every intermediate step. It requires clear checkpoints: the task goal, the data sources, significant changes, and the final action. Then the human is not a micromanager, but the editor of meaning and the owner of the decision.
Where does the human remain irreplaceable?
When tasks ran into difficulties, 76 percent of the 588 tasks with recorded difficulty moved forward through human intervention. In 23 percent of cases, the agent solved the problem on its own. But the type of help matters more than the numbers: the human usually added information rather than doing the agent’s work.
In 35.2 percent of cases, help meant adding context or clarifying requirements. In 34.7 percent, it meant diagnosing the issue and switching methods. Partial edits accounted for 3.2 percent, and full takeovers for only 0.7 percent.
In other words, the bottleneck was not manual labor, but judgment. The human explained what mattered, what was wrong, which constraints had been missed, and why the current path did not fit. After receiving feedback, AI handled the revision itself 75.4 percent of the time when output needed changes.
For a business, this changes the role of the employee working with an agent. The person does not need to compete with AI on execution speed. It is more important to set the task, provide context, notice reasoning errors, and make the final call.
Source: The Decoder: AI agents do more of the work in model development, but humans still make the decisions
Quick answers
Can an AI agent make business decisions on its own?
An agent can execute chains of actions and choose some intermediate steps, but goals, scope, and important parameters should stay with a human. The Atria Dawn study showed that final decisions in such workflows were mostly made by people.
How is an AI agent different from a regular chatbot?
An AI agent does more than reply in text. It can call tools, work with files or data, produce intermediate results, and revise them after feedback, but its boundaries should be set in advance.
What tasks can a small business give to an AI agent?
Reasonable tasks include drafting, searching through data, sorting information, reminders, initial checks, and preparing options. Decisions involving money, customer commitments, or access rights should go through a human.
Why is it risky to simply enable autonomous mode?
Autonomous mode is useful for long tasks, but it can hide the chain of decisions from the human reviewer. Without predefined checkpoints, oversight can turn into a formal approval of the final output.
