Skip to content
Yusuf Bagus Anggara
All posts

Multi-Agent AI Orchestration as a Daily Work Tool

How I run several AI agents in parallel on real development work — and why the final review always stays with the human.

OCT 18, 2026 · 5 MIN READ · AI, ENGINEERING PRACTICE

Single prompts got boring fast

My first year using AI for development was one agent, one chat, one prompt at a time. It worked. It was also slow in a way that's hard to describe until you've felt the alternative: every task serialized behind the previous one, every context switch costing a full conversation rebuild.

What changed my daily workflow wasn't a smarter model. It was treating agents like a team I can brief and run in parallel — orchestration, not prompting. This portfolio, my client work, my side projects: I now routinely have several agents working at the same time, on different slices, while I stay on review and integration.

I want to be precise about what that means and what it doesn't. No magic. No "AI writes my architecture." A lot of mundane discipline, honestly.

The shape of the workflow

Every task I delegate goes through the same three stages. The stages matter more than any particular tool.

1. Decompose before you delegate

The failure mode of single-prompt AI isn't answer quality. It's task size. A vague "build me a blog feature" produces a vague thousand-line diff you can't review.

Before any agent touches anything, I cut the work into slices that are independently verifiable:

  • a loader with a clear input/output contract
  • a component with an explicit props interface
  • a test file with named cases
  • a migration, isolated from the feature that uses it

Each slice gets a short brief: what exists already, what the interface must be, what "done" means, what is explicitly out of scope. Five minutes of decomposition saves an hour of untangling.

2. Delegate to specialists, not generalists

With several agents running, I assign by domain rather than dumping everything into one context window. One agent on the server-side data layer. One on UI. One writing tests against an interface that another agent is implementing. The contexts stay small, which keeps each agent's focus sharp and its output reviewable.

The parallelism is the obvious benefit. The underrated one: smaller contexts produce more predictable code. An agent that has only the loader brief in front of it doesn't get creative about the UI.

3. Verify like you don't trust it — because you shouldn't yet

This is the stage people skip. Green tests are not the same as correct code, and "the agent said it works" is not a verification method. My gate:

npm run lint && npm run typecheck && npm run test

That's the floor, not the ceiling. Then I read the diff. Every line that will ship, I read. The point of parallel agents is speed of production, not speed of approval. If review becomes the bottleneck, the bottleneck is working as designed.

When multi-agent is worth it — and when it's overkill

Worth it when the work has natural seams: a feature spanning loader, UI, and tests; a refactor with mechanical steps; a migration plus the call sites that consume it. Anywhere parallelism maps onto real boundaries in the codebase.

Overkill when the task is small enough to hold in one head. A five-line fix, a typo, a config change — spawning three agents for that is theater. It also wastes your review budget: context-switching between three trivial diffs costs more than writing the fix yourself.

There's also a category where I don't delegate at all: decisions. API design, data modeling, naming that will outlive the feature. Agents are strong executors of decisions and weak owners of them. I make the call, then delegate the implementation.

Anti-patterns I've been burned by

Letting an agent review its own code as the only review. An agent critiquing its own diff is useful as a second pass, never a first. It has the same blind spots in both roles. Human review stays the gate — that's not Luddism, that's ownership. If I ship it, I'm responsible for it.

Unbounded autonomy. "Fix everything you find along the way" sounds efficient and produces scope creep no one asked for. Every agent brief now has an explicit out-of-scope list. Related work goes in a notes file, not into the diff.

Trusting plausible code. AI-generated code fails in a specific way: it looks right. The imports resolve, the types check, the happy path works. The bug hides in the branch nobody exercised — the empty list, the concurrent request, the retry. Review means hunting those branches specifically, not admiring the structure.

Skipping the decomposition step when you're rushed. The times I regretted multi-agent the most were the times I skipped briefs and pasted a whole ticket into three chats. Parallel garbage is still garbage, and now you have three piles to read.

What stays human

Delegation, review, and integration. I decide what gets built and how the pieces fit. Agents write the first draft of code faster than I can. Whether that draft ships — that's still my signature on it.

Eight years of development taught me that speed without a review gate is just a faster way to be wrong. Multi-agent orchestration made me faster. The checklist made it safe. Keep both.