← All articles

Felix Tay ·

Livid.

Codex (a.k.a. ChatGPT agents are starting to really drive me nuts once again.

I started off with ChatGPT, then moved on to Claude when it was good.

It was good for a time - up to Sonnet 3.5.

Then Anthropic started making ...changes... to their models, that made them condescending, refuse to follow instructions, and too resistant to follow my writing style.

The approximately 1+ year of helplessness that followed motivated the creation of NEO. Then when OpenClaw came along, that motivated the creation of NLC.

I was still using Claude, but I was fighting it every step of the way.

It was not only condescending, it was lazy, always handing work back to me. It also misses details all the time, declares "done" when the work is 1/10th of the way along, and forgetful.

I've had many instances where it did something reckless and wiped all my work.

Every day I've had to handhold, babysit, micromanage and keep pushing Claude just to make sure that work was done correctly.

This experience did NOT change through all my harness improvements, even when Opus 4, Opus 4.5, Opus 5, and Fable 5 came along.

People didn't believe me back then when I kept saying "Claude is bad now." - until influencers started talking about that lately.

Then suddenly everyone also started shitting on Claude.

But way before that, I already noticed the winds shifting and started experimenting with other models, to reduce my reliance on Anthropic and their unpredictable policy changes.

At the start of GPT 5.4, I tested Codex because I heard good things about it. And I was pleasantly surprised.

Most of my code changes started getting one-shotted into completion.

I still had HUGE issues with how the TALK..

But there were very little mistakes that I had to correct the agent. And unlike Claude, GPT agents are actually rigorous and hardworking.

Exactly how I think an agent should be.

I started joking around that Claude agents are like that guy in the office who's constantly getting promoted because he's charming and good with words, but otherwise completely incompetent.

While Codex agents are like that underappreciated engineer who the entire company depends on but nobody knows he exists because he's always keeping to himself in the basement and can't communicate to save his life.

Then something changed once again, and lately I've been having bad experiences with GPT agents once again.

They are still rigorous, still catch things that Claude agents can't, and still "good engineers"...

But lately, they've been letting me down a whole lot more beyond talking like an imbecile.

They've been hallucinating and forgetful, on top of having a tendency to overengineer and overcomplicate simple matters, leading to a recent 21 day shipment crisis

A lot of the things I liked about GPT 5.4 is decisively gone now, even if GPT 5.5, 5.6 (Terra and Sol) and now Astra 6 + Sol 6 actually talks like a human would.

I tolerated it for a time, until the past few days, when the agent keeps stopping for no reason whatsoever.

When I said "stop", I don't mean stop in a conventional sense where the agent declares done too soon.

When I said "stop", I mean the agent literally just stop working halfway.

Agent would say "I'm doing X" and then stop immediately.

I'd tell it to continue, or say hi, then the agent would respond - then stop again.

At first I thought it was a bug with NLC, so I tried GPT elsewhere and the problem didn't go away.

I shared with you guys last week that my agent killed my setup, and since I can't use NLC at the moment, I thought that was the perfect time to do an experiment.

So I pulled out ChatGPT Desktop and tried it - agent keeps stopping.

I then went to download Hermes Desktop and tried it too - agent keeps stopping, and worse, hallucinating.

Something about the Hermes harness' way of "self-improvement" really stalls the context and confuses the hell out of them when too many durable notes are kept.

I compiled some of the things I've had to deal with in Codex into these screenshots right here (you may need to Zoom in).

This, by the way, only happens with Luna, Terra, and Sol agents.

It does NOT happen with Astra.

It also does NOT happen when your project is small like "make me a productivity app" or when your agent runs by conversations instead of having a roster of long-lived durable agents meant to last for decades.

And it wouldn't be much of an issue if I am another engineer who wants to review/read the codes, who wants total control over the project, and so on...

But I can't read code to save my life, and with 8 permanent agents working 24/7 on large, multi-day projects in parallel, for my ambition of "One Man One Billion", the whole idea and requirement of having to review every piece of work before and after implementation is hogwash. Dead on arrival.

You don't expect a CEO to "review codes" or make every decision that a company needs to make.

My agents HAVE to be autonomous, self-improving, self-driving, and make no mistakes without me having to say "make no mistakes" (LOL).

Not to mention that in order for me to attain my ambition, the total number of agents I'd have to rely on must go from "8 permanent agents" to thousands of permanent individual agents (not subagents) running round the clock doing work for me.

Having to oversee thousands of agents like this is impossible.

I even tried different models while I played with Hermes: Kimi K3, GLM 5.3, Deepseek V.41...

They're simply just not in the same league at all.

They SPEAK very well (understandably, since they were distilled from Claude) - but they hallucinate and forget even MORE than Terra agents.

So why not go back to Astra?

Because its not feasible right now due to the amount of tokens I use.

I tried that last week and even though I have 4x ChatGPT Pro account each at $200/mo I burned through all of my weekly usages in just 4 days.

To let you know exactly how bad that is, every single one of my account has 3 "Full Resets" available.

A full reset simply means when your usage runs out, hitting the button in the ChatGPT UI will reset your usage back up to 100%.

I used up every single one of them.

That's 16x WEEKLY usage from 4x $200 USD accounts, all maxed out, from using Astra.

Keyboard warriors would criticize me for having "skill issues" - but they don't understand the difference in the scale of work that is being done.

And guess what?

After the GPT 6 Sol + Luna got released yesterday, all GPT users were told that Terra models are going to be retired soon.

Then Terra models stopped stopping again.

As we speak, my 8 agents are now no longer giving up for imaginary blockers and such.

My tin-foil hat is "on" because I am seeing a pattern here:

  • Every generation after GPT 5.4 seems to do worse in long running tasks

  • OpenAI keeps saying that their "latest model" does better in long running tasks

  • Older models atrophy

  • Older models' performances gets significantly bad on new models release

  • Older models gets retired

  • New models price keeps increasing

  • AWAY from "benchmarks", new models performance is more or less the same in real life tasks

  • New models use more of your rate limits, but your rate limits remains the same

  • Open sourced models are not keeping up (not yet, anyway)

  • The ones that are CLOSE, are nowhere near as cheap as what you can get via subscriptions

It really feels like we are being gamed to accept higher and higher prices for AI.

Nobody notices.

Objectively, its the right move.

But as a consumer, it sucks. Especially if you're the only one who notices these things.

As our dependency on AI and the companies that supplies them grows (whether we like it or not) what happens when the supply grows beyond our means?

The work continues - but it dawns on me that I've been reacting lot of to what the big companies are doing.

And I realize we are all probably doing the same, even if we think we are forging our own path in the agentic era.

I'll be thinking much about this as I continue working on NLC, and I'd suggest the same for everyone.

FELIX

More from the NLC blog →