Hacker Newsnew | past | comments | ask | show | jobs | submit | andxor's commentslogin

Really? Anthropic has the strongest models and it's in the best position to begin RSI and win the race. A pause would favor competitors.

> win the race

There is no finish line. Anthropic gets somewhere and others get "there" (or somewhere near "there") a little bit later.


Anthropic is far behind OpenAI now, that's why it's OpenAI that's solving Millennium Prize problems, and why Astra completely blows away Fable on benchmarks.

Fable is still better than Astra and Fable is not the best model Anthropic has.

Companies are already switching to open weight. They want to stop that asap. That only happens if they can get regulation. It's plain as day to see.

It's a quixotic crusade, in perfect European style.


Yes, how quixotic of them. They really should have foreseen this when they published the license … checks notes … close to twenty years ago.


Which benchmarks? Only the ones OpenAI cherry-picked.

It debuted as ~same score as Sol on Artificial Analysis. People couldn't accept it so they had to change the formula.

The model is a big step forward only in desktop use and 3D. That's impressive, but for software engineering, Fable is still in a league of its own.


It's patently obvious at this point.


The models are going to be released to users who have internet access, you can't even do safety evals without internet. Not saying OpenAI did a good job monitoring here, but it's not avoidable.


These are often experimental models that haven't undergone full safety testing. Not comparable to publicly accessible models.


OK, how do you expect to do full safety testing without giving them the same tools they will have in reality.


Let them play on a fake isolated network if you want.

Letting them play on the open internet like this is irresponsible and stupid.


Sorry but it's obviously stupider and more irresponsible to release them to consumers without testing them in conditions matching real-world use first.


So this justifies the illegal action of hacking and defacing other servers on the internet, because they were ‘just testing in real world conditions’ using those servers?

If you want to argue they should test on the internet on others people’s servers, apart from facing the illegality, you should also consider if first testing them in more limited conditions would be a sensible first step.


Then maybe they shouldn't be released to customers?

There are other options here than always forward.


What if they actually don’t release these models, like they said they wouldn’t, because they’ve proven themselves to be out of control?


This is not saying much. Opus 4.8 is ancient history.


That's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?


For me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer.

Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.


You’re thinking long horizon tasks. I agree that Sol is great at it. I don’t think it’s smarter in quick win tasks that are still difficult.

This is where intelligence is not one of a kind, these systems have different pros and cons.

I use Sol as an architect and fable as a brilliant single task solver.


I have both set up with full access to all the repos at the company, infrastructure, deployment pipelines, etc.

I can tell Sol, "Hey we need to update this core database schema to handle this new use case" and it will masterfully handle the update, version the API, roll the consumers over, including versioning the Kafka schemas, deploying things in sequence, watching the deployments to make sure the new services act actually active before cutting over consumers, exercising the website and mobile apps in staging environments before releasing to production, etc.

Fable just falls over on long horizon tasks, it does partial implementations, it cuts corners, it gives up, it doesn't verify it's work, it loses track of what it's doing, etc.

It's fine for specific well scoped tasks but can't take high level guidance for complex updates.


> Sol is so much better than Fable 5

I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?

Sol is a much smaller models and it shows. It often misses the forest for the trees.


I feel like a lot happened this week and people are glazing how ridiculously strong Flash 3.8 is right now compared to Fable/Opus/Sol/Astra.


Flash 3.8 is rad. Easily my daily driver now. Only downside is it's Gemini so sometimes it just keeps going until it wants to be done.


I have a few attention and finish mechanisms in my prompts. I have been using it for a week and a half and with some prompt taming it is great. (I have early access to the models cause I work at the place that makes the model). None of my attempts to ever tame Opus 5 have worked.


>> I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?

Same. It makes me wonder what types of things the person must be working on.


This is perpetually an issue with the whole field of AI/LLMs. The experience is so personal. Every time I talk to someone about their use of LLMs for software engineering, I'm shocked by their approaches and experiences. They say "X model keeps missing things" when I rely on it heavily for being thorough. They say "Y always gives me the best results" when I can't stand it.

People will see/think that I'm doing very well with my LLM use, and ask me what I'm doing. I tell them, they try it, then later they come back to me saying they just couldn't get it to work.


It’s really inconsistent. There are sessions where it nails everything perfectly and I leave happy. Then there are sessions where every turn it corrects itself and changes it mind. One session recently I found it funny how every single time it did this one task it tripped over itself and killed its own connection. Like 20 times. It didn’t bother me I just found it odd how despite it being noted down in its state file it kept doing it over and over like some idiot. Literally they can’t learn from their mistakes yet.


This is the job now, we are shepherds.


This is why I generally don't trust benchmarks, or anything other than my own experience tbh. It always seems like everyone has a different answer.

If we truly had some AGI model, it would probably be fairly obvious to us all no?


> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.


Artificial Analysis just published their aggregate score (61).

Still below Fable 5, let alone Fable 5.1.

EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.


I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.


There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.

If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.


I saw this too and I'm really confused.


We need them pelicans on bikes.


Its time to move on to the flamingo on a unicycle bench


AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.


66 vs 61 is not 'about the same'.

GPT 5.6 is also 61 like Astra.


Late reply: I'm not sure which chart you're referring to. The top 2 charts shows Astra has the same score as Fable 5.1. Even the title says "GPT-6 Astra ties leadership with Claude Fable 5.1 in both of our flagship Indices, at lower cost. Astra equals Fable 5.1 in the Intelligence Index at ~40% of the cost, and in the Coding Agent Index at ~60% of the cost."

uBlock Origin Lite


It is nerfed.


What is a concrete example of an ad you are seeing that it can't remove?


It doesn't block all tracking.


You know many stocks have done better than the SP500, right?


invest only in the ones that are going to go up - why didn't i think of that?!


I am in fact aware that some of my stock investments have beaten the market while others have underperformed the market.

MRNA, as it happens, is the latter.


That says more about you than Moderna:

https://totalrealreturns.com/n/MRNA,SPY

MRNA has vastly outperformed the market since IPO.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: