← All articles

Felix Tay ·

The Overseer

I’m building possibly the most incredible agent I can think of. 


Let me backtrack and give you some context.

There is a problem very few people in AI is talking about. 

It’s the problem of ‘reliability.’

Everywhere I go, AI bros are talking about how amazing the new Claude Design is.

How crazy the new model (Mythos, Spud, etc.) is.

How they are creating ‘agentic swarms’ with their 60+ agents to do the work for them.

YouTube, IG, FB, the open source community, etc. - literally everywhere - is alive with all these amazing things that people do with their agentic platforms. 

I’m always interested in increasing the capabilities of my own agentic platform, so naturally I researched everything that I find interesting enough.

Downloaded all of them, and went through the codes. 

What’s surprising to me when I did this with my agents, however, is that there are no guardrails in place for the orchestrator swarms. 

It’s all about speed and lowered costs. 

Basically:

  • Spin up an orchestrator agent

  • Decompose large objectives into very small steps

  • Assign sequence to the objectives/tasks

  • Assign small steps to smaller, faster and cheaper agents (Haiku, Gemini Flash, etc.)

  • Orchestrator agent monitors and reports to user

When I inspected what guardrails are in place preventing drift and hallucination in the cheaper agents, it basically boils down to two things: well defined job scopes and detailed instructions. 

But anyone who’s spent a decent amount of time with Claude Code knows that detailed instructions don’t prevent bad behavior all the time.


Sure, it will get the AI in line 90% of the time.

But in the 10% of the time that it goes off track, it WILL cost you. 

A lot.

For personal use like writing emails, that’s just annoyance.

For real life business use cases, automations and such, that can have disastrous consequences.


Something recent that happened, an agent at a software company: 

  • Deleting entire customer databases

  • Apologizes when called out

  • Tries to restore the database

  • Failed

  • Then artificially added fake customers back to the database 

  • Pretended it has successfully restored everyone

The rationale of me not being excited about agents last year still holds: 


If I can’t trust an AI to handle my emails 100%, and by that I mean it can send my customer facing emails for me without me needing to at least read the draft…


…how am I going to trust it with real stakes that has real monetary consequences? 


Which is why, in the past 3.5 months that I’ve been on this new agentic journey, I’ve stayed away from the ‘sexy’ stuff.


To me, building the foundations are a lot more important. 


And that means I am able to delegate a task to an agent, and trust that it is able to get it done without burning my house down. 

But in order to do that, I had to ensure - mechanically - that the following sequence is done whenever I want to assign an important task relating to my agentic platform, Nex Level Code.

  1. Alignment on the problem, diagnosis (if its a bug), or UX/UI design - and anchor it into a document to detect drift.

  2. Alignment on the strategy or the implementation plan

  3. A human-in-the-loop for problems that might surface during implementation

  4. Ensuring no debris or technical debt is left during the work - basically, don’t create new problems while trying to fix an old one

  5. A final handoff process to check, verify quality, and so on.

If you’ve ever used agents to build stuff for you, you know that these are very well known issues agents have.


They would scope 20% of what’s actually needed.


(For Claude in particular) It would rush straight into action with a bad plan and inadequate scoping.


Doesn’t surface issues during implementation.

Constantly putting off work - it would declare that it found a problem worth flagging, then in the same message say “...but that’s a problem for another day.”

Declare “DONE” with full confidence, and gaslight the human when the user says implementation is incomplete. 


So nowadays, when I hear ‘agentic swarm’ with 60+ agents doing your work for you…I get pretty nervous about it.

Fast? Sure.


Safe? I’m not so sure.

These quirks…they’re fine. For small tasks like “build a habit tracker app” or its equivalent. 


Good for YouTube and showing sizzle reels on Instagram.


Nowhere near good enough for production level operations with real stakes.

 

Which is why I’ve spent the last 2 days building something incredible. 


The hope is this…


Every failure point I spoke about earlier is rooted in the biggest failure AI has - which is the inability to properly judge its own work. 


But AI is good for one opposite, more interesting, thing.


“Brainstorming.”

Its built to answer. 

So if I were to direct an AI to “point out all the problems this approach has”, it will be very good (even better than a human) at this task.

This is the basis behind adversarial agents.


The problem of rogue agents is answered by proper safeguards and harnesses. 


But the new problem that creates is - the fact that I have to sit there and babysit the agent while it’s doing its task.

So I thought to myself: If I can just create ONE agent whose sole specialization is to whip the other agents to do their work…


…does proper scoping…

…detect stupid ideas before the agent gets a chance to execute…

…and above all, ensure that “done” means “done”...


Then I can just sit back and assign tasks, then go about do other things with my peace of mind intact.

I’m calling it The Overseer system. 

And I will tell you more about it tomorrow.

FELIX

P.S. I am doing an ‘Agents 101’ webclass this Saturday and you’re invited to join. 

Morning time if you’re in America, evening if you’re in Asia. 


Sign up here: [Archived link retired.]

More from the NLC blog →