← All articles

Felix Tay ·

Sunday round up

This will be a long post.

The NLC Stabilization Project is close to finishing, and upon completion will unlock a new horizon of scaling for both NLC itself AND what you can do with NLC.

What has been happening: I found things are fine when 3-6 parallel agents are working on NLC itself, but problems start happening when I increase that count to 10 parallel agents & more.

Reasons why this is happening are as follows.

—

#1: CODEBASE BECOMING HUGE

I started NLC initially as a sort of lightweight layer on top of Claude Code.

But as I learn more about what agents can do, I began investing more time developing NLC, and this means features got added on organically rather than pre-conceptualized as part of a larger roadmap.

This created issues here and there such as:

  • Adding new things breaks old things because the full architecture was not documented (because it wasn’t planned) and nobody remembers all of it, all at once

  • The same class of bugs keep reappearing because we lost track of all the paths we need to account for when fixing an issue, so we keep playing whack-a-mole. We once had to fix the same thing 12 times because there were 12 ways the issue can happen and we forgot 11

  • Agents are not aware of the number of times the same class of bugs have reoccurred, and so they needed me to manually tell them what’s been happening, why did they happen and what was tried. We’ve been having to manually creating documents on debug trails and such to track those issue

The permanent solution here is consolidation.

If one feature has 12 ways of not working, we need to make it so that fixing one area will fix all 12 areas, instead of having to manually track down all 12 separate paths.

We need to make the purpose and relationship of every code to its intended behavior so obnoxiously obvious that even the dumbest agents can immediately tell at a glance how they work.

Concretely, this means deep modules rather than many shallow modules.

We are:

  • Clarifying the boundaries for every function in the codes

  • Making the codes live in the same neighborhood based on their function

  • Refining our folder architecture, making sure every file have proper homes

  • Making sure that customers agents don’t accidentally put customer files

  • Creating a mechanism to automatically update architectural documentation

  • Creating another mechanism that mechanically tracks historical issues that keeps reappearing

In other words, the codebase is cluttered, so we are KonMari-ing it.

Sounds basic, but this is the difference between vibe coding and agentic engineering.

—

#2: TOO MANY COOKS ON THE SAME SOUP

To have a true ‘AI Corporate Army’, I need 50 parallel agents working all at the same time, round the clock, without rest, without human supervision, and able to identify exactly the most important thing to do at the moment, do it flawlessly, and hand off perfect work items.

Without hallucinations, inventing facts, or objective drift.

There are many interesting thesis I have on how to make this happen, including the Overseer project, but we were blocked at the time by one simple dependency: merge conflict.

Translation: Too many cooks touching the same dish.

Someone added salt, the other person didn’t know and added more salt = the soup is now too salty.

And this showed up in the last shipment where a lot of the problems were previously fixed but didn’t actually get shipped. Agent A’s work contained a fix, but was superseded by Agent B’s work which didn’t have the fix.

When you are working with one capable agent: You tell an agent to do one thing → the agent does it → you check the results.

But when you have many agents doing work at the same time, you need proper delegation of tasks, sequencing, and ochestration.

I don’t want to work with only 1 agent. Or 6.

I want 50. More.

I don’t want human supervision. I don’t want to put on my artist hat, and sit down to design things ‘properly’ with every agent.

A CEO does not and should not be involved in every conversation that happens in his organization.

He needs to be able to trust that every individual in every team of every department is able to do their work and hit their KPI month after month, year after year.

He should not be sitting in front of the screen just to click on “Approve” on 50 agents every hour of every day, and should be able to step away from the business if need be.

He needs to work on, not in, his business.

The mechanism that allows this does not yet exist in the market.

For that, I’ve been building a system that requires only two human touch points:

  • First touchpoint: I tell an agent what I want in a very vague manner

  • Second and last touchpoint: Agent delivers a handoff that I just need to sign off on

This hasn’t been easy because it needs to work not only for myself, and not only for code…but for ALL of my users, no matter how vague their ask is, or what industry or work they are doing.

The artistry should not happen in the work itself, but the creation of such a system.

Recently OpenAI launched Symphony and Matt Pocock released Sandcastle, both containing pieces that I need, but I still needed to build everything else.

So this week, I integrated a number of new things to NLC:

  • Every agent now starts work on things with a separate branch by default

  • Every agent gets their own dedicated workspace inside the repo

  • Every time an agent needs to do work, they get their own separate copy of the main project

  • Each copy is merged sequentially to main project copy on handoff

  • There is a whole protocol that is mechanically enforced to ensure agents do the right thing

The protocol in question doesn’t just enforce proper merging, they enforce quality control as well.

There are checkpoints within this process that forces agents to slow down, perpetually remembering the anchor objective, forcing adversarial review of the plan and the execution.

Agents are also forced to produce proof of the quality of task that was done, not just proof that it was ‘done’.

These effects are a combination of two mechanism that we’ve taken to call Doorway and Agent Flow, and I will be speaking more to that in later posts.

Today, the build is already completed and is currently being tested.

—

#3: AGENTS ARE FORGETFUL

AI amnesia is obviously a thing and by now, nobody’s expecting agents to do any better in this regard.

What this means is that when building something, what was previously built or decided on was already forgotten unless checked.

In code, there’s a risk of redundancy or breaking something that was already there.

In facets of business, that means executing strategies that may be at odds with a previously agreed on direction, or company values.

AI bros would tell you about their fancy “memory architecture using durable states that allows persistent memory, loaded at the top of every invocation so they survive compaction.”

And then you actually go their paid courses, and realize its basically a version of:

  • Tell the agent to do something

  • Tell the agent to write stuff down

  • Tell the agent to refer to said file whenever there’s a conversation

That’s fine if the operation is small scale.

But I am building NLC to allow one-man businesses (like myself) to run a large operation while producing high quality outputs.

Currently the longest uninterrupted, zero-intervention run I’ve seen from my agent without using a Ralph Loop is 6 hours 28 minutes and 40 seconds.

That is nowhere near enough.

I need my agents to run autonomously for months on end with zero human input, yet produce high quality work like talented full time employees.

The thesis for that is shockingly simple: Don’t rely heavily context memory to sustain long running tasks.

I will work under the assumption that context limit is 1/4th its actual size.

Every job that is agreed on, the agent creates a contract with/for me containing all the important details of the job.

A new database entry is created for this specific task with the most top level detail + state management.

Then at every single turn, NLC injects it into the agent’s context:

  • Objective: An unambiguous, binary yes/no objective statement that is testable

  • Atomic claims: A list of success criteria or decomposed steps

  • Status: Where the agent is at (which is programmatically determined based on agreed set of criteria)

  • The link to the contract (agreed set of objective, plan, success & testing criteria, and other important details)

Since the contract self-contains all relevant instructions to the job, system prompt tells the agent to read the contract if things are unclear. And because the link is right there, the agent don’t have to dig through a pile of files in order to know what it is.

If my thesis is right, that’s all that is needed for an agent to make progress over extremely long horizon tasks.

We are now testing this with progressively long horizon tasks.

All this to say, my NLC users will be able to enjoy a number of powerful new features in the next update. All of which are byproducts of me solving my own problems.

I will continue to share my learnings and journey with all of you.

FELIX


More from the NLC blog →