NEO is getting fixed.
Some users have been reporting that NEO's coaching mode felt ‘off’ today.
I took a close look and found out why.
TLDR;
Sonnet 3.5’s lack of intelligence is actually useful for coaching
Sonnet 3.7 and above is trying too hard to think (and probably too intelligent too)
A pure system prompt architecture causes recursive, never-ending questioning even when that’s no longer useful
Today has been invested in its entirety to fix this.
I tested this personally and so far it seems to work.
I’ve asked the team to implement - and beta testers can expect a new (more pleasant) experience interacting with NEO starting tomorrow.
That’s all for the short version.
If you want to read the full story, read on:
👇
NEO's AI coaching prompt was built on Sonnet 3.5.
And we were running this in Claude, instead of our own app.
It functioned perfectly then, but not with the current model.
Sonnet 3.5 was perfect.
Sonnet 3.7 was condescending.
Sonnet 4 was also condescending, but with a bonus of being overly cautious.
It keeps asking for confirmation 5x before it does anything - even after you confirmed the task.
Opus 4 got announced and for a short while, we were excited.
It was a complete let down.
Instead of being overly cautious, it went the complete deep end on the opposite side.
It assumes everything, thinks too hard, thinks too MUCH, and gives you what you want plus 99 other things you didn’t ask for.
Every “upgrade” just seems to make the models worse.
I don’t know why this happens but I don’t dwell on things I cannot control.
This was one of the biggest impetus that motivated me to build NEO.
I wanted to run a wrapper powered by Sonnet 3.5.
We’ve been hard at work to do this, and we’re very close to allowing open beta.
(We’re at closed beta right now)
Sadly, Anthropic announced that Sonnet 3.5 will be retired on October 22nd.
And while they are doing so, they advised us that in the lead up, we will experience decreased availability.
I contacted Anthropic.
I tried to buy the model that they didn’t want anyway - but they’re not keen.
So a few days ago I made the decision to switch NEO’s model to go from Sonnet 3.5 to Sonnet 4, as advised.
And the result was…well…if you’ve been using NEO, you now know.
So I went to figure out how I can work with what I can control (NEO, its system prompt, architecture, etc.)...
I learned a few things:
#1: Different system prompts
Found out that the system prompts that Antropic gave to 3.5, 3.7, 4.0, Sonnet and Opus are all different.
That, combined with how the models are being trained, caused different experiences when the AI coach prompt is being injected.
#2: Anthropic’s official Claude architecture
Found out that Anthropic does not use “system prompt” in the conventional sense.
A system prompt is usually appended together along with the user's inputs.
E.g.
System prompt: “Speak in koan from now on”
User input: “What is the meaning of life?”
The input sent to the AI is: “Speak in koan from now on” + “What is the meaning of life?”
(This is a simplification)
This is what AI apps normally do.
But Anthropic’s “system prompt” is a MASSIVE 25,592 character prompt.
To save up on compute, they send this prompt…once.
Before your “What is the meaning of life?” prompt reaches Anthropic, that 25,592 character prompt arrives first.
And that massive prompt (I would call it a priming prompt instead of a system prompt) gets remembered and stays in context throughout your conversation.
This, paired with Sonnet 3.5’s “lack of” intelligence compared to 3.7 and the AI coach’s instruction to ask one question at a time, causes Sonnet 3.5 to naturally “forget” to keep asking questions the moment it is no longer needed.
Just that it does so at the most appropriate time.
So…
That means I need to take action on quite a number of things.
Action Step #1: I’m using Sonnet 4.0 instead of Opus
Because Opus, due to its training, keeps skipping steps in its attempt to be ‘comprehensive’...
It rushes too fast into solutions and we spend more time correcting it rather than iterating from what’s working.
Annoying.
Action Step #2: I have to make changes to NEO’s system prompt.
I’ve decided to swap Sonnet 4.0’s system prompt with Sonnet 3.5…then adapted it to NEO.
And just like Anthropic, I’m going to use a priming prompt instead of a purely system prompt.
Saves tokens. Because the priming prompt is huge.
At the same time, creates an anchor for all the following outputs by the AI.
Action Step #3: Added new architecture
I’ve added BOTH system prompts and priming prompts.
System prompt instructions are only used as prompt injection to “counter” the most annoying behaviors that Sonnet 4 exhibits.
So it’s kept short.
While the massive prompt that governs its behavior is injected as a priming prompt.
Action Step #4: Strategic changes
Made some strategic changes to the prompts.
A rookie mistake that people often do is to “add” more instructions to a prompt.
But that’s a mistake because, among other reasons, if you simply add something, you run the risk of having conflicting instructions within the same prompt.
Among the secret sauce, I’ve decided to tell NEO to embody a certain attitude that I’ve learned from a book that changed my outlook on how coaching should be done.
…that’s an email for another day.
.
Anyhow, that’s all for the updates.
NEO users will find a better experience after implementation is done.
I’ll do a follow up email on this tomorrow.
FELIX
P.S. I am offering a single, 1-on-1 consultation session on everything related to AI and marketing.
Doing this to fund some developing work on NEO.
This offer expires in 48 hours.
If you want details, reply to this email.