The models are going to be released to users who have internet access, you can't even do safety evals without internet. Not saying OpenAI did a good job monitoring here, but it's not avoidable.
Sorry but it's obviously stupider and more irresponsible to release them to consumers without testing them in conditions matching real-world use first.
So this justifies the illegal action of hacking and defacing other servers on the internet, because they were ‘just testing in real world conditions’ using those servers?
If you want to argue they should test on the internet on others people’s servers, apart from facing the illegality, you should also consider if first testing them in more limited conditions would be a sensible first step.
For me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer.
Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.
I have both set up with full access to all the repos at the company, infrastructure, deployment pipelines, etc.
I can tell Sol, "Hey we need to update this core database schema to handle this new use case" and it will masterfully handle the update, version the API, roll the consumers over, including versioning the Kafka schemas, deploying things in sequence, watching the deployments to make sure the new services act actually active before cutting over consumers, exercising the website and mobile apps in staging environments before releasing to production, etc.
Fable just falls over on long horizon tasks, it does partial implementations, it cuts corners, it gives up, it doesn't verify it's work, it loses track of what it's doing, etc.
It's fine for specific well scoped tasks but can't take high level guidance for complex updates.
I have a few attention and finish mechanisms in my prompts. I have been using it for a week and a half and with some prompt taming it is great. (I have early access to the models cause I work at the place that makes the model). None of my attempts to ever tame Opus 5 have worked.
This is perpetually an issue with the whole field of AI/LLMs. The experience is so personal. Every time I talk to someone about their use of LLMs for software engineering, I'm shocked by their approaches and experiences. They say "X model keeps missing things" when I rely on it heavily for being thorough. They say "Y always gives me the best results" when I can't stand it.
People will see/think that I'm doing very well with my LLM use, and ask me what I'm doing. I tell them, they try it, then later they come back to me saying they just couldn't get it to work.
It’s really inconsistent. There are sessions where it nails everything perfectly and I leave happy. Then there are sessions where every turn it corrects itself and changes it mind. One session recently I found it funny how every single time it did this one task it tripped over itself and killed its own connection. Like 20 times. It didn’t bother me I just found it odd how despite it being noted down in its state file it kept doing it over and over like some idiot. Literally they can’t learn from their mistakes yet.
There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
Late reply: I'm not sure which chart you're referring to. The top 2 charts shows Astra has the same score as Fable 5.1. Even the title says "GPT-6 Astra ties leadership with Claude Fable 5.1 in both of our flagship Indices, at lower cost. Astra equals Fable 5.1 in the Intelligence Index at ~40% of the cost, and in the Coding Agent Index at ~60% of the cost."
reply