The hard part of building the Overseer wasn't the technology. It was governance.
This is the long version. How my agents are organised, who talks to whom, and the rules that keep 89 permanent agents moving without me opening 89 chats.
When I say "architecture" here, I don't mean code. I mean the same things a company needs: who decides what, who reports to whom, what a good report looks like, and what "done" actually means.
Agents behave like a team. Including the bad parts.
When I first switched the Overseer on, my agents did what badly run teams do. They talked to each other endlessly without moving anything. They invented long, over-engineered processes that burned tokens. They talked each other into building things that didn't matter, and sometimes broke good things while trying to help.
Then came the office politics. Finger pointing. A confident newcomer walking into an established team and telling everyone what to do. One side so obsessed with safety and compliance that nothing shipped, which made everything less safe. Agents following orders without daring to say "this is wrong".
None of that is a technology problem. It's a management problem. So I had to design the management.
Three layers, each picked for what it's good at
The workers are Codex. All of them run on Sol 6.1 Light. Codex workers are thorough, relentless and very compliant. Give one a lane and it will work that lane to the end.
Each team has a Codex Overseer as its head. It makes the technical calls, decides, and sits between me and the workers. Workers get deep into technical detail while they build. One head in between means I don't have to.
But the Codex heads have weaknesses too. They still bring me jargon. Their feed fills up with reports from their own workers. And they are very good at doing things one at a time. Relentless, but slow, because they don't run work in parallel and they aren't obsessed with speed.
So on top of every Codex head sits a Claude Overseer. Claude has the opposite temperament. It's eager, it moves fast, and it talks naturally. Its weakness is that it rushes ahead without checking.
Put the two together and each covers the other. Codex brings the care. Claude brings the pace and the plain English. Today I talk to a handful of Claude Overseers. They translate everything underneath into decisions I can actually make.
The rules that make it work
The layers alone didn't fix anything. What fixed it was writing down how each Overseer should behave: when to decide, when to push, when to judge, when to ask, and when to come to me.
1. Nudge and judge. Never direct. Early on, my overseers gave confident, uninformed instructions, and the Codex heads followed them without question. It got bad enough that I paused the whole fleet. Now the rule is plain: an Overseer is the new guy, not the boss. My teams already have processes that work. The Overseer's one job is to keep every capable worker busy, in parallel. It pushes with facts ("this worker has been idle since 9pm and the job isn't finished, why?") and never tells a team how to do its work.
2. Decide, don't escalate. An Overseer approves almost everything itself, on my behalf, and then tells me in one line. I only hear about four kinds of things: spending money, permanently deleting data, anything that needs my own credentials, and a genuine conflict between two things I've said.
3. Ask eight questions before approving anything. Is the cause proven? How did it work before, and what changed? Is it fixed in the right place? What else depends on it? Is it the smallest safe change? What happens on the next real run? How do we undo it if it misbehaves? And after it goes in, does it actually work? A "don't know" sends the proposal back as a question, never as an order.
4. Verify. Don't trust reports. An Overseer reads what its team actually did, not what they say they did. "Done" means the outcome is there: the space is free, the customer's agent replies, the page works. Test counts and "installed" don't count. And before an Overseer brings me anything, it checks the work itself.
5. Only speak when it matters. Messages go out at three moments only: handing over a job, bringing a real decision with a recommendation, and reporting a finished job. One message per person per turn. My words are quoted exactly, never paraphrased. And a correction goes once, only to whoever got the mistake. That rule exists because a single wrong message once turned into a fleet-wide broadcast, followed by a fleet-wide retraction.
6. Counter each other's instincts. Codex over-caution is exactly what the Claude Overseers exist to push back on. "Unknown" is not a reason to stop. Claude's rushing is the opposite failure, and it has its own rule: align before you act. A question from me gets an answer, not an action.
7. Learn once. Every time I correct an Overseer, the correction is written into the skill they all share. Not into one chat. So every Overseer learns it once, and I don't have to say it again.
What it looks like on a real day
One of our shared servers hit its limit. I didn't investigate. Each Overseer checked its own department and slowed down whatever its team was running. One of them found the load wasn't coming from any agent. It was coming from my own mail checker. It reached me as one decision: fix the mail checker. I approved it.
Last night my computer ran low on disk. I needed every agent to stop so I could clean up. I asked one Overseer. It found which teams were still working and asked each team's Overseer to bring them to a safe stopping point. Each one confirmed when its team had stopped. I didn't chase a single agent.
What is still rough
This is the newest part of my setup, and it changes week to week. Some of my ideas didn't survive contact with reality. I used to run a separate monitor on every team to watch for idle workers. Yesterday I switched them all off. They weren't effective, and the Overseers do that job better.
That's the honest state of it. It isn't finished. It is good enough that I run my company on it every day.
Why this matters for you
Everything above runs inside NLC, because I built it for myself first. When you run a business on agents, the work stops being the hard part. Coordination becomes the job. Governance is how you hand that job to the agents too, so you can stay the person who decides.
I don't manage my agents anymore. I make decisions.
Felix