1,024 AI agents are not an advantage. A reusable process is.

Hundreds of AI agents in one run make a strong demo. But agent count tells you almost nothing about whether the work will be cheaper, better or easier next week. If every run starts with a new prompt and ends with a manual rescue, you have not created scale. You have only created more activity.
The useful question is not how many agents a platform can start. It is: what has improved when the same process runs for the hundredth time?
1,024 agents measure capacity, not advantage
xAI describes Grok Build workflows that split a large job among parallel agents. A normal run gets a budget of 128 agents, while large jobs can use up to 1,024. That suits work that genuinely decomposes: reviewing many parts of a codebase, sorting a hundred issues or asking independent reviewers to verify the same findings.
The more important part of the launch is not the maximum count. A successful workflow can be saved, shared with a team, accept arguments and run as its own command. Progress is also preserved across pauses and resumptions. That is where a one-off run starts to resemble an operating process.
Source: Workflows in Grok Build — xAI
More agents can increase throughput. On their own, they cannot decide what an accepted result looks like, which sources should be used or whether the run produced something a customer actually needed. An unclear job does not become clearer because 1,024 instances work on it at once.
Three launches, the same direction
Grok Workflows, OpenAI Presence and Manus Plan Mode are different products for different users under different terms. They are not interchangeable. Yet they point in the same direction: value is moving from the individual prompt to a process that can be reviewed, reused and improved.
OpenAI Presence packages policies, standard operating procedures, approved actions, simulations, evaluations and an improvement loop for voice and chat agents. Each deployment starts with a specific job. OpenAI says its own English-language phone support resolves 75% of inbound issues without human assistance and that human handoffs fell by 15 percentage points in ten days. Those are vendor-reported internal results, not an independent industry benchmark. Presence is also offered through limited general availability with OpenAI engineers and selected integrators, not as a self-service product.
Source: Introducing OpenAI Presence — OpenAI
Manus Plan Mode turns the approach into an editable document before execution. The user can change goals, steps and constraints before confirming the work. The feature is manual and available to all users on web and mobile. It does not prove the final result will be correct, but it makes the approach visible before hours are spent in the wrong direction.
Source: Introducing Plan Mode — Manus
What the three examples share is not a particular model. The process gains memory outside the chat window: a saved workflow, a set of evaluations or a plan that can be edited and run again.
Reuse changes the economics
A prompt can save time once. A reusable process can reduce the cost of every case that follows. You see the difference only when you count the whole job, not just the model price or response time.
Consider a service company handling 250 invoice queries per month. Manual preparation takes an average of 18 minutes: find the contract, check the order, summarize the discrepancy and draft a reply. If a stable AI workflow reduces preparation to 6 minutes, it saves 12 minutes per case, or 50 hours per month. This is an illustration, not an outcome promise. The point is that the effect is measurable.
The first version does not have to solve everything. It can retrieve the right material, label the discrepancy and draft a response. Once the team sees which cases still get stuck, the next version can target them. Value then compounds through repetition:
- the same job definition applies to every case,
- the same quality criteria follow every version,
- the same exceptions become visible in measurement,
- and each improvement affects future runs, not only today's prompt.
That accumulated effect turns the process into an asset.
A concrete case: the invoice query that returns every week
A company does not have to start with an agent platform. It can start with one recurring case whose result already has a human definition.
For an invoice query, the job could be to identify the order and contract term the customer refers to, describe the discrepancy, propose the next action and draft a response that a case handler can accept or edit. The run becomes useful when it leaves evidence: which sources it used, what discrepancy it found, what the handler changed and whether the customer had to come back.
After a few weeks, versions can be compared. Were more drafts accepted without substantial rewriting? Did waiting time fall? Did fewer customers return with the same question? If the answers are no, the company does not need more agents. The process needs better inputs or a clearer definition of done.
When many agents actually help
Parallelism is useful when three conditions exist at the same time:
- The work can split without constant coordination. A hundred documents, code files or support cases can be reviewed separately.
- The partial results can be checked. Sources, test cases or criteria let a second agent or a person catch errors.
- The synthesis has a clear target. The final output should be a ranked list, report or decision brief, not just hundreds of answers.
When one of these is missing, more agents can increase cost and text volume without improving the outcome. The capacity is real, but the business advantage is not.
Ask what improves on run 100
When evaluating an AI project, replace “how many agents can we run?” with four more revealing questions:
- What is reused from one run to the next?
- What outcome does the business accept as done?
- Which measurement shows that version two is better than version one?
- Who owns improvement when customer behavior, rules or systems change?
If those answers exist, you have the start of a process that can grow. Buying more parallel capacity only makes sense after that. If you want to turn a recurring workflow into a measurable operating asset, that is the kind of work we do in Tool Forge.
The real advantage is not starting the most agents. It is no longer having to reinvent the work every Monday.
FAQ
Does a company need hundreds of AI agents to automate a workflow?
Usually not. Start with a recurring case where the goal, inputs, accepted outcome and economic effect can be measured. Parallel agents matter only when the work genuinely splits and the combined result can be evaluated.
What makes an AI workflow reusable?
It requires a stable job definition, explicit inputs and outputs, measurable quality criteria, versioning and an improvement loop. That turns the process into an operating asset rather than a one-off prompt.
When are many parallel AI agents an advantage?
When the job contains many independent parts, such as reviewing many files or cases, and the results can be checked and combined against explicit criteria. More agents do not help when the task, data or quality measure is unclear.
The Forge newsletter
Get new articles in your inbox
Pick the topics you care about. No noise, at most one email a week.
We follow GDPR. Unsubscribe anytime.


