Hacker Newsnew | past | comments | ask | show | jobs | submit | mlmonkey's commentslogin

IMHO (not a paper writer, but read a lot during my grad school years), the Genie is out of the bottle. The only way forward, as I see it, is using LLMs for reviews also. Basically, filter all submitted papers with an LLM and ask it to summarize it, find the biggest weaknesses and main strong points, etc. that a human can then use to review the paper. Basically, LLM-as-a-reviewer .

Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.


maybe but as a paper writer, the quality of reviews and reviewers have gone down because of LLM-as-a-reviewer too, because LLM reviews seem to regurgitate the limitations section of the paper, and are highly influenced by the way things are phrased in the paper rather than the actual substance.

but I'm hopeful that some middle ground will be found in the future


I dunno man, I use LLM's to review things before I send them in and I often feel like I've been handcuffed to a chair, had a bright light shone in my face and had to account for all the mistakes I've made in my life.

They can be brutal.


Completely agree. I’ve had a Nature paper and a Nature Neuroscience paper published this year and the reviews from Opus were much more thorough and complete than those of the human reviewers. I still benefitted from the human reviewers, but now I consider an LLM review essential.

There should at least be a code of professional conduct where authors state the extent to which LLMs were used. (This would also help not wasting time by asking some “authors” about “their” paper.)

Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.


There is. NLP conferences, like the ACL family (and the ARR) require you to disclose in the paper if you used LLMs for writing and coding, which ones and how. Whether every author is honest is a different matter.

This already happens. It's clear when a reviewer used an LLM, and it's very annoying for the authors that have to respond to what are usually low quality, superficial reviews.

Boy do I have news for you! :-D

"low quality, superficial reviews" have always been around. Reviewing is most often an unpaid, thankless job and many times reviewers barely put in the effort.


There is still a difference. Previously, if the review was low quality and superficial, it was also quite short and easy to answer. Now I receive fully LLM-written reviews of 2-3 pages (or more) of superficial comments asking for tons of additional text and experiments that are almost always outside the scope of the work. Answering such a review requires a lot of work with absolutely no gain.

Similar experience to sibling: a disinterested human isn't going to ask for 30 different detailed things to add to, or change in, the paper that could technically be improved but aren't worth the additional page real estate. An LLM reviewer definitely does this.

To be fair, low quality superficial reviews were also not uncommon before LLMs…

Here is my proposal (or actually our proposal - me and my agents): https://zby.github.io/commonplace/articles/what-an-automated...

I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted

The entire point of an academic paper is to add to the sum of human knowledge. How can an LLM trained on a subset of human knowledge possibly even begin to accurate evaluate such a paper?

I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.


This is very bad logic

1) Humans also are trained on a subset of human knowledge. 2)A lot of papers are just about experimenting something, and then applying simple stats. Eg empirical studies, around 1/3rd of published papers. Like, we tried this drug or did this experiment, from a sample size X here are the results. An expert is needed to maybe comment on the conclusion/hypothesis of the underlying suspected mechanism, but LLMs are still very useful on catching bad statistics or p hacking (so so common)


Schmidthuber has an answer: https://arxiv.org/abs/0812.4360

For a more practical approach you need to use proxies: https://zby.github.io/commonplace/articles/what-an-automated...


Extend this beyond review. The most value we would get is from quality checking existing published papers.

Yes. I want reproducability, open-sourcing, accessibility, correctness, and most of all: usefulness. I don't care how it was written or reviewed, as long as some assurances regarding above things can be made, and I don't see why LLMs would get in the way of that.

What can't be gotten rid of fast enough is the notion that having written something is meaningful on its own. Making something that looks right was a level above total novice: now it's the floor.


Nope, I’ve tried this, it’s awful.

For a start LLMs love LLM generated text, so you are boosting papers people never had any input in.

Secondly, LLMs in my experience are good at small issues, but fail totally at the whole paper being obviously poorly constructed, or clearly fake.


> Let's see if it (or crossbreeds) can be tamed.

Clearly you are not a cat person ... the question is not if it can be tamed, but how quickly it can tame the humans around it ... ;-)


All along I had thought that "AGI", "RSI", etc. were at the model level: but this paper seems to be talking about "agents", etc. I'm not sure having a swarm of agents explore a problem space in parallel via brute force is what "AGI" is about. I'd be happy to be proven wrong.

AGI and RSI are both meaningless terms, meaning whatever you choose them to mean.

RSI is the new sexy. Models are RSI-ing themselves towards the singularity, these folks' agents are RSI-ing themselves towards mastery of their training environments, and my pet cat is RSI-ing himself into the best cat that he can be.


> Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models?

Catch is, even Googlers internally do not have access to top-tier models. (or did not until recently, when apparently Claude was made accessible to the SWEs internally).


Googlers now have internal access to frontier models.

Funny enough, after trying it, I went back to G3.8


Yes, agreed, 3.8 with thinking is quite good.

The goal is to kill us all .... j/k. :-)

I don't understand the point of a "one time" tax. What happens once the money runs out a year later? Another "one time" tax?? Why not implement a permanent tax if you want to permanently solve the problem??

If they are blocking the one time tax, would they allow for a perma one. Don't be naive. No matter what type of tax it is, they will throw money to keep the status quo. $100M is chomp change to him and he can deduct it most likely.

But what is the point of a 1-time tax? It doesn't solve anything! It's just a slush fund for the SEIU!

They pay off some debt that has piled up. Like paying off your credit card debt.

> don't understand the point of a "one time" tax

It's basically to send a message to a group of people that they should fuck off. There is a reasonable gamut of opinions about whether this is valid. But I think it's fair to say Californians are upset about billionaires, and that this is a coherent way to message that opposition. (Similar to the dumbfuck states messaging around wokeness or whatever. A narrowing minority of America is prosperity first.)


This is like Antonelli saying "we're going too fast, everyone slow down"!

Is there an imagemagick implementation that I can just throw into the pipeline?

Agreed. And his doom words have set a 1000 mouths in the Pentagon/Whitehall/August 1st Building/Kremlin salivating with excitement.

Take China, for example. Look at any recent ML conference, and see the fraction of articles majority-authored from Chinese universities and labs. Do you think they'll slow things down anytime soon? I don't think so!

It's a global arms race, and we're just spectators.


Does this matter vs actual capabilities?

Does the Kremlin being excited about a tech mean anything of the tech doesn’t deliver?


Hand over every piece of personal information to Zuck .... um... no thanks, dawg, I'll pass.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: