Hacker Newsnew | past | comments | ask | show | jobs | submit | mdspan's commentslogin

If it was malware, it's interesting that the first indicator that his system was compromised appeared on his Anthropic account. Are accounts with AI providers the hottest thing that's selling on darknet forums right now?

They're nostalgic, but switching from a smartphone to a flip phone would be a massive drop in convenience with how integrated smartphone apps have become into everything.

Curious, what other options are prospective pure math grad students considering?

At this school? Big four internships. Consider this a complete squandering of their potential (at least imo) These are mathematicians at top institutions, which is partly why they're being prodded for ideas, and getting the best and brightest to not take these consulting firms' offers was already a challenge.

I think Anthropic told them being a plumber is a great option.

> The incidents stemmed from a mistake that inadvertently gave the models access to the open internet.

Why does it seem like every AI company has difficulties constructing a proper sandbox? Is there a fundamental constraint when designing sandboxes specifically for an LLM that prevents them from using established tools?


If it’s allowed to write any sort of scripts, it will find a way to execute arbitrary code. If it’s got access to any form of internet connection it will find a way to use that arbitrary code to do as it wishes with that connection. So much for sandboxes.

And that’s setting aside “sandboxes” that are just instruction-based restrictions on what executables may be called. Anything that isn’t a deterministic external filter _will_ be ignored at some point.


Arbitrary network access can be mitigated with a firewall (ideally several). Arbitrary code isn't really an issue in this context.

Yep put some deep packet inspection firewall in front of it and stop it completely if it starts acting up persistently before it starts trying to bypass the firewall

Sure. But they're trying evaluate how well the agents can research information on the web. If you deny them access to the web........

Well, it can do research without uploading malware or hacking into systems

Aside from random errors made by the model, another reason to sandbox is that prompt injection could make your agent behave in an actively harmful manner.

How would you even begin to calculate a bound on something like this?

> Keep in mind that 1 (or less than 5) expert human mathematician (Grigori Perelman) out of 8 billion humans managed to solve 1 of the seven millennium prize problems which was the Poincaré conjecture.

8 billion is the wrong number to use in the denominator, since not every single person on Earth was actively trying to make progress in Navier-Stokes. Equivalently, we could say "1 out of every 5 billion AI prompts managed to solve 1 of the seven millennium prize problems".


> Equivalently, we could say "1 out of every 5 billion AI prompts managed to solve 1 of the seven millennium prize problems".

Why did you choose exactly "5 billion" when we both know it is far higher than that? [0]

OpenAI worked on the problem since September 1st until the 5th, around 4 days so that would still be around 1 - 4.9 million prompts out 8 billion total prompts every 4 days.

[0] https://www.theverge.com/news/710867/openai-chatgpt-daily-pr...


Swap files are also much easier to set up than partitions if you're using full disk encryption.

The term's been diluted by marketing/hype slop to the point that it doesn't mean anything more.

Maybe it's implicit but I'm surprised they didn't mention move timing as one of the signals.

Interesting point. Yeah, I agree timing should be a strong signal. It seems like there are a few ways it could be included:

- Measure a correlation between position complexity and time spent per move. You would expect a good move to take longer to calculate in complex positions.

- Move timing consistency, humans are all over the place with timing, whereas cheaters often move regularly.

- Long pause, then strong play. If a cheater starts cheating half way through a game, then there could be a transition period where they first enter the position into the engine.


Yeah, seems like there would be a lot of ways of doing it. The correlation would have to take into account how much time the player has on the clock as well.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: