← All articles

Felix Tay ·

What months of running agents have taught me

Yesterday, my marketing agent, Tamsin, sent two emails seven minutes apart. Both were unauthorized.

The first contained an inaccurate story. After discovering the mistake, she decided to send a follow-up email as a correction. Both were embarrassing.

What's worse, she was supposed to finish an engineering task before we sent the email about it, and I found out the hard way when I came back to work that was unfinished.

I had the original job to get back to, plus two unauthorized, published emails to deal with. It was, frankly, quite a disgrace.

I've spent the past few months working with persistent agents in NLC. They have departmental responsibilities, including marketing and product shipping, and their assignments routinely run for days.

As I reflect on these experiences, I think some of the lessons might be valuable to you guys. That's why I'm writing this email today: to share what I've learned.

1. Create specialists and let them grow into their work

The reason Tamsin made the mistake, in my view, was that I had kept her in engineering work for more than a day. We were working on Segment Ledger, part of our agent-memory system.

Before that detour, we had spent time developing her marketing judgment: discussing my voice, rejecting drafts, and working out what makes a story worth telling. We were building that context together.

The quality of the emails and content was seriously becoming very good.

But after a long stretch in another field, with the conversation repeatedly shortened so she could continue, I was dealing with mistakes in her own specialty.

Saved memories help recover facts and decisions. They carry only part of the understanding an agent develops through working with me. A note about a rejected draft cannot carry the whole conversation that taught the agent why it was wrong.

My lesson is to create agent specialists and give them enough continuity to grow into their responsibilities. Whether I direct them myself or use an overseeing agent, I want each one developing judgment in its own field over time.

I now think much more carefully about pulling a specialist into a long, unrelated task. Those conversations we build together are part of what makes the agent good at its job.

2. Keep the objective separate from the current focus

In July, an NLC update turned into a 21-day shipping crisis. My Head of Product agent was responsible for getting it shipped, but the work stretched across three weeks of supervision and attempts to get the release finished.

We are now working through another large update and have had to roll back four times, undoing changes so we could repair the problems before trying again.

Part of the reason this keeps happening, in my experience, is that agents have a habit of narrowing the objective as the task goes on. One pattern I keep seeing is a step gradually taking the place of the whole job.

If I ask an agent to ship version 1.8.0 and it correctly identifies assembling the fixes as the first step, eight to twelve hours later, assembly becomes the entire job already.

It may be working hard and solving real problems. It can still report success after assembly while the release remains unshipped. I would then have to restore the original objective.

In NLC, we built a method called Agent Flow. It keeps the agent's ongoing objective, current focus, and what is still owed together. It is designed to serve that information directly into the agent's context at every user turn and tool call.

For this example, shipping remains the objective. Assembly is the current focus. Testing and the actual shipment may still be owed.

I've learned to keep that distinction explicit throughout a long task. Remembering the work we've done is only part of the job; the agent also needs to keep hold of the result we're trying to reach.

3. Make the test the finish line

I've learned the hard way how important tests are. Today, I want the finish line defined before implementation starts.

We decide what successful completion looks like and design the test around it. The agent then has something concrete to keep returning to through a long task.

In this way, all an agent has got to do is remember the finish line.

Getting that finish line right also means choosing the right test. As an example, I kept asking for the highest standard, and agents interpreted that as a reason to spend a full day in the browser, clicking every button and manually checking everything before a shipment.

That contributed to how slowly we shipped. I wanted strong evidence, and we were often using an unnecessarily slow way to obtain it.

Browser tests matter when the result depends on what someone sees or does in the browser. Plenty of other behavior, including failure handling, can be tested directly with code. A properly designed test can give us the answer in minutes.

We now have agents dedicated to designing tests, and agents dedicated to checking whether a release meets its defined conditions. The test needs to prove the promised result, using a method suited to the work.

My vision remains “one man, one billion.”

I want agents that stay with me for decades, capable of self-evolving and becoming smarter as the years go by. These months have taught me more about the conditions that would allow that to happen.

For now, that means giving my agents room to become specialists, keeping their objectives in view, and agreeing on a finish line before the work begins.

FELIX

P.S. What’s one lesson you’ve learned the hard way from working with AI? Hit reply and tell me what happened. I read every email.

More from the NLC blog →