Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What gives LLMs the ability to reason so well about math specifically is the math-specific RL training they are given where they are rewarded for chaining together reasoning steps etc in a way that results in success (e.g. a known correct result).

The difference with something completely at inference time new that was not in the training data, or not used by RL training, is that while it will be able to use it to some extent, it has not been taught via RL how to best use in in a reasoning chain.

The problem is that LLMs don't really have the generic ability to reason, so they are instead need to fake it by either:

1) Fine tuning on reasoning data (very specific)

2) RL training (more generalizable, especially for math/coding)

3) Prompting that encourages "keep on going" "tree of thoughts" exploration where they may discover a reasoning chain largely by luck

Demis Hassabis has talked about combining LLMs with search (cf systems like AlphaGo) which sounds like it might be useful for math research.


It's all because of "UX designers" who have to justify their continued existence by constantly finding new things to change (break). If computers several decades ago had more than enough power to render button-looking buttons, there's no other excuse for the sad state of UI today.

It's providing an alternative to middle east sourced fertiliser from plants that can no longer ship or, in some cases, can no longer operate due to war damage.

US states aren't the only places on the globe building out renewable powered green hydrogen, methanol, ammonia, etc.


> There is something about mathematical discovery (progress, advancement, creation) which is vital to the spiritual, experiential quality of doing mathematics. The creation (or even the pursuit) of novel mathematics is one way that humans have historically accessed the ineffable and encountered the divine and mystical.

The memory manufacturers don't believe the rate is sustainable, so are hesitant to start up new production to meet demand when it will inevitably crash. They have to operate on longer timescales than software companies.

How? He’s saying that he can’t think of any good ones, and in my experience that’s true of most informed people. This is one of those hard problems that needs solutions, and so far the present system is the best anyone has managed.

I don't support fossil fuels but wind works a small fraction of the time, it is awful to wildlife and it is an eyesore.

Instead of wasting the already wasteful energy by making fertilizer out of air, these farms could likely much more efficiently re–capture runoff from their waterways that is now being dumped into the sea harming humans and wildlife alike.


looks like the best proposal is the just the least bad

If you can think of a way that you can have a much higher chance than 9% of success in getting a useful treatment to market, you can make a boatload of money. Go ahead, reform it.

Its horrendously inefficient...and inefficiency is usually bad for the environment. For example, the solar PV when sited in Minnesota is so inefficient just making the electricity from natural gas would produce less CO2 per watt. So there is that, and that's before you look at if all the equipment you use to build the plant was better used elsewhere. And then you could put the solar PV somewhere where it would actually been better than just burning natural gas for electricity.

There’s nothing funny about it? Also seems like a lose for the personal websites and blogs when people bounce because of it.

It should be illegal to ask for references - the purpose of the interview is to understand your capabilities and abilities. If they couldn't suss all about you in that process then they have a shitty interview process.

A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in late November, which for most people meant early January due to the December break.

General agents (OpenClaw, Anthropic Copilot, ChatGPT "Work") started working even later than that.

This category of software may have a much more meaningful impact on work than the mostly-chat systems we were using from 2022-2025.

Studies that mainly focus on 2022 to end of 2025 might be missing out on a material uptick in capabilities.


An article by Bruce Schneier, on how he frames the question when evaluating his students.

I would say it leaves communication between individuals in a gray area. And discourages translation work because it has to be done manually. I would say, LLM-translated text is better than asking an LLM to translate it again and again every time someone asks how to do something, because the LLM-translated text can be corrected for everyone to see, whereas conversations with LLMs stay private by default.

Jet engine design is iterative, and the basic principles are well-understood.

Drug design seems a lot more binary. You can find a new pathway, but drugs themselves are fairly simply molecules, and you can't iteratively 'fix bugs' the way you can in an engine or a piece of software.

The engine is a given, it's almost astronomically complicated, you know a lot less about how it works than you'd like to, and you're trying to change how it works while it's running without breaking anything, using tiny rigid parts that have to snap into place correctly and can't be bent to fit.

High failure rates aren't surprising.


Take a lesson from the UK where there are cameras everywhere.

Brothers van parked in London. Window broken, bag stolen off front seat. Right under camera.

He calls police and miracle of miracles they actually attend. Must have been in the area.

Plod:"Shouldn't have left your bag on view"

Bro:"Can you check the CCTV"

Plod:"You watch too much TV mate, we don't have time for that"

But cameras will one day help catch a murderer and suddenly everyone will want them everywhere.


Update: stripe reactivated my account

But what if the genie instead offered (or threatened) to give _everyone_ the answers to everything you'll ever discover in your career, including your boss, your competitors, etc?

> consider how difficult it would be to find a nice way of writing an article

Writing is often difficult. That's what makes good writing so distinctive and valuable. Just because something is hard doesn't mean it's not worth putting in the effort.

> "You are all fucking idiots" is sometimes a reasonable starting point.

It might be a good starting point in one's head (which, admittedly, it often is for me, but I don't then broadcast it). I think we'll just have to agree to disagree on that.


We’re gonna need a bigger wall.

If OpenAI is relying on “classifiers” (their word for what was disabled) to serve as the model’s judgement, rather than teaching the model itself to be well-aligned, then I worry.

To be fair, the system prompt was presumably also different from what it would be during deployment, and perhaps the model was also at a different stage of training. Without more details it’s hard to judge. But it does seem models should be able to avoid performing obviously misaligned actions – misaligned not only with the model spec, but with the user’s intent – without needing external classifiers or instructions. The only case where I’d personally let the model off the hook is if the instructions given were very badly worded, in such a way that the model could actually reasonably think that hacking HuggingFace was part of the assignment. But I doubt that’s what happened.


I got news for you about the DNC as well...

Of the 3 proposals, the 2nd/B seems to be more of an “informed consent” model while the other two are seeking a comprehensive exclusion. I’m curious how strongly opinions are held. I hope earnest dialogue is forthcoming.

Notable are the endorsements, which are balanced over all three options, indicating only 1/3 are supportive of a more moderate position. If this represents core contributors, this could be worrisome: a clear majority favoring a comprehensive ban leading to passage, but yet a sizable portion of contributors who don’t want a ban. Losing 1/3 of contributors would be a significant loss to the project.


You have this flipped backwards

Regarding me being a pretentious nimrod: no doubt there are many subscribers to this frame of mind.

Clarifying my structural comment: since LLMs are next token generators, they fall into the "aren't even wrong" side of lying versus truth-telling. Truth reflects a reality to which LLMs are not directly privy. Hallucinations are baked into the source. Harnesses on agents help to ground some, but they don't fully solve the issue. This doesn't make them useless. Just unreliable narrators (i.e. liars).


> Having only fun things to do make them obviously boring.

This doesn't really pan out the way you think it does as people can always choose to do boring things if they want (if there's even any truth to your assertion). Unless your argument is something along the lines of humans need to be forced to do things they don't want so they're able to find fun in anything, which I think is pretty... wild.


> We simply have better magic wands and more powerful spells now.

Wouldn't it be nice though if the incantation of the same spell would always do the same thing every time ? You see that's how my old wand and spells worked.


Thomas--as insufferable a windbag and terminally-online personality as he is--was not wrong in writing that post.

It's also hard to find a sufficiently anodyne way of writing that particular post, because quite frankly such a large a vocal contingent of the tech sphere is just so goddamned wrong about what was obviously happening in front of their lying eyes.

To use a coarse analogy, consider how difficult it would be to find a nice way of writing an article that says "folks, gravity is real and everybody saying that dropping an apple isn't going to have predictable results is wrong".

"You are all fucking idiots" is sometimes a reasonable starting point.


Yeah that does really suck. Are the fines just a cost of doing business fine? :(

Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: