Hacker Newsnew | past | comments | ask | show | jobs | submit | hypfer's commentslogin

I feel like this is either AI-generated or severely lacks context.

There might be a point in there, but it might also just be paragraphs of text with little structure joined together.

What's the narrative? What does this want to tell me? And why?

Something like an opening header block with 2-5 sentences would go a long way.


I just copy-paste the text and ask AI to summarize for 90% of articles I see on HN. One of the most time-saving uses of AI.

That answers the what, but not the why. The why however I'd argue is a lot more interesting.

Also, you can often derive the what if you know the why and just follow the thread and map it out.


> The MCP Industrial Complex

Was that a real thing? I mean it must've been for it to be mentioned there, but, rephrased: what was the scale of that?

How many individuals were involved in that? 1? 10? 100? 1000? 10000? 100000?


I'm not sure if this prediction will hold true.

We're not seeing the progress in those "frontier models" that we have previously seen. There's certainly still gas left in tank tank, but we're way into the diminishing returns by now.

Cloud inference still beats hardware investments by orders of magnitude of course, but that's only if your data doesn't really matter to you.


We are certainly not in the diminishing returns phase for LLM progress. No sign of that yet.

I’ll grant that for specialized applications like coding agents and mathematics, but even there I suspect that most the real gains are actually taking place in the harness.

But I suspect returns may have already diminished into negative territory for at least some other use cases. One of my least favorite job responsibilities in this brave new era is figuring out how to avoid performance and behavior regressions when an older model were using for some application reaches end of life. It’s getting uncommon for me to look at our benchmark results and say, “Oh, good, it does better on one of the newer models!”


One thing that I really want to know - the better models from today vs a year ago - what has changed. They have already pre-trained on all available public data. Scooping up the last percentage of archaic texts which were never digitized is not going to move the needle.

Is it just that the providers are generating tons of synthetic datasets on coding tasks so that the models get more exposure to the right thing to do? Every time someone points out an LLM stupidity they add some training data to patch over the weakness (trivial to generate "there are two 'l's in llama")?


>suspect that most the real gains are actually taking place in the harness.

Part of the reason harnesses work well is you can run a lot of agents in parallel. That doesn't slow down demand.


That is true, but the eventual realization that more machines doing more coin flips in parallel does not mean "more work gets done" might.

LLMs are amazing tech, but they're terrible without oversight. More agents faster just makes reality collapse on them quicker.

But yeah, you're right, temporarily, this will still push demand. But the topic was about "diminishing returns" as in "tech getting better". Not as in "customer spending".


It's kind of weird because more machines working together does mean more work gets done. Coin flips and weighted coin flips are totally different things. Any biases weights towards reality push you closer to reality when you use them.

New models keep being able to use more and more agents on longer time frames. Your hypothesis doesn't look like what we're measuring.


Who is we?

The people mapping AI capabilities.

Oh cool, so that we includes me! :)

I had actually been thinking more about all the non-LLM functionality that go into the harnesses. I'm not going to name names and I haven't done any rigorous testing, but my general impression is that choice of harness matters more than choice of model. In terms of basic task completion success specifically, not code aesthetics.

A perfect harness will not extract gold from a dumb model. It's a system that builds on each other, though we've not probed that frontier much to have a good intuition on what effects what.

Well I mean if I wanted to be extra pedantic, I would argue that we've been in that phase since LLMs were first introduced.

Before that, we had 0. After that, we had more than 1.

A leap as far as that is hard to recreate.

But that wasn't my point. That's just trolling.

The actual point is that LLMs aren't gaining new capabilities anymore. They just get more reliable at the ones they already have; turning what was a coin flip to some higher probability.

That's (intuitively speaking, not strictly mathematically speaking) kinda the mathematical definition of diminishing returns.


It’s a constant tension in computing that has been around since mainframes and clients… Neither is going to disappear. My general feeling is normal people care more about how thin and light something is than their privacy, so if data center powered LLMs will have a strong future.

Hmm I'm not 100% sure about that, given that edge is very viable, and the geopolitical climate has changed quite significantly.

I agree that datacenters are not going to go away, but I have doubts that the buildup that has happened is really going to pay off for most operators.


they really dont want to hear this bro lol

I can see that by those reddit-style vote swings, but who are "they", exactly?

Who is so emotionally invested into random comment sections being purely positive about their pet.. uuuuuuuh.. tech?

Very weird.


Are there known examples of software where the "software factory pattern" (an incorrect term, given that it's not a pattern but a workflow. Or rather an idea. A hope. A wish.) proved to work long-term?

Preferably ones that I could validate myself instead of just having to take someone's word for it.


It’s all experiments at this point

This feels agentically generated.

The blog, the post here, the (auto?)killed LLM comment.


It actually is though?

Though arguably more of a process and judgement issue than skill.

What makes LLM-generated code a bit special there is that misjudging how to deal with it seems to be what most people do. So the default is broken.

Whereas in prior iterations of "skill issue", the default was working.


> A desktop that loves you back has to know you.

Is this another downstream consequence of the COVID lockdowns? Do people need a hug?

___

Hmm. So from the actual talk description, these seem to be UX people that might have interesting ideas to share but also have completely lost the plot when it comes to the overarching narrative.

Which is unsurprising, given that their summary is clearly written by AI (albeit with human editing, but barely)

https://conf.kde.org/event/11/contributions/314/


Maybe someone should sell "the end is nigh" sign nfts with fun AI-generated designs on them, to be used in your metaverse villa.

"Turn the kitchen to 230°C" was executed with "confidence": 0.9015

lol well i guess you can turn the whole kitchen into an oven with needle :)

But for real usecases you are able to set explicit minimum and maximum values on the output range of numeric arguments, so that you can avoid situations like these. In this case it was hard for us to do that while keeping a broadly appealing demo since celsius and fahrenheit have different "reasonable" output ranges.


I wouldn't list Opencode as "good reputation".

They had their own unbound "harness scans the whole user directory" oopsie and handled concerns about that by introducing code signing.

Which, yes, does have absolutely nothing to do with that issue.

I guess by now it is better, but to me they seem to lack the engineering culture necessary for a "good reputation" stamp.

__

Ref: https://github.com/anomalyco/opencode/issues/14925#issuecomm...

among other issues.


How about the one where if you start a session outside of a Git repository, the "worktree root" is set to /. Bug report closed as "not planned".

FWIW, I don't think that they're being malicious. They instead just seem to have no idea nor do they care.

And the original comment I've replied to proves this strategy right! So from a business standpoint: excellent work.


Glad my arbitrary failure to try them has worked out! For people seeking OS-native harnesses, I can recommend Factory's Droid. I know I'll be returning to it with my head hung low today, after I uninstall ZCode.

It does have a "mission" feature that's stuck in the strange, distant times of 2025 by way overdoing mandatory verification steps, which means they don't support swarms/workflows/crews/fleets yet -- that is, it's all done in sequence. But they have the boring, corporate engineering attitude that I think we're are all craving rn, and generally seem competent.

I can heartily dis-recommend Vix, even though they gamed themselves to the top of at least one ranking site that shall not be named; exactly like the quasi-bad-faith incompetence described with OpenCode above, but without even the "Open-" branding! Though perhaps that word has been so thoroughly burnt as a prefix by Sam Altman & Microsoft's criminal behavior that we should let it go...

Is this how "FLOSS" wins over "OSS"? Not with an ideological bang, but with a marketing issue?


[flagged]


This is kinda beside the point and this whole thread may be wiped when dang wakes up and notices the AI slop article we're commending under, but your reply is thought provoking so I'll attempt a response anyway;

I'm sure you're far more experienced than I with basically every aspect of this discussion, but I'd argue that's given you a blindspot, here. I'll hit some specifics below, but the headline is that you're effectively taking a stand against Eternal September II -- a goal that I hope we can all agree would be quixotically antisocial, given what followed the first one!

  we arguably do not want FLOSS to "win" over OSS
I think(/hope) that fellow FLOSS proponents would passionately disagree. FLOSS isn't a brand of chatroom, nor even merely a community: it's an ethos regarding labor, property, and liberty. Demanding that all users of your software are also activists for your particular take on intellectual property is clearly a doomed undertaking for anything beyond a toy or library, anyway.

Didn't you get into this stuff to change the world? To liberate the oppressed, undereducated, and forgotten with the radical power of the information superhighway? Cause it reads here like you're more motivated by selfishness (not wanting to bother talking to people with less expertise than you) and resentment. On that note...

  They want that. They do it themselves all the time.
Here you equate "non-hacker people" with software engineers you don't agree with, it seems. You're ofc welcome to think companies X Y & Z produce "miserable-ness", but as absurd as it sounds, it sure seems like you've forgotten the fact that some users are not developers. Many, in fact! Over 99%, even!

Less confrontationally; my mom is in her late 60s, and is pretty computer-literate for her age after decades of knowledge work. Surely you'd agree that she's not, like, evil for using OSX, iOS, GMail, Word, etc.? That she didn't chose those things because of a philosophical commitment to defending IP laws, but rather because of structural reasons? Even if she were pro-IP, wouldn't we want to win good, well-meaning people to our side?

  So let them have the "Open" prefix. It's just words, anyway.
I do agree with this still, but as a philosopher I just have to say that everything is just words. It's language games, in fact! Which is why I simply had to reply.

I hope none of the above was rude; I'm trying hard to keep my passion for this topic from pushing me past HN guidelines :)


Cutting things short:

> but I'd argue that's given you a blindspot, here

I'd argue it's the opposite. The idealism there _is_ the blindspot. Not the other way round.

You can't save everyone. And you will die trying.

That's the first thing that gets (or should get) hammered into people's heads when they pick up a career in all things social.

Which isn't to say that we shouldn't dream, but I believe that our dreams should be optimized for maximum gain with minimum pain.


Well I personally think we can find a middle ground between single-handedly saving "everyone" from poverty and addiction and oppression as social workers, and not letting anyone into our exclusive philosophy-of-property clubhouse. I would invite you to join us on this pro-social mission, but you seem perfectly content as-is!

Some people are still on Usenet after all (?), so I suppose it's not a big deal if a few people want to cling to old communities. I hope you don't mind if we use the word for what it was coined for though in the meantime, back in the real world.


To be frank, I'm not really interested in "joining" your thing there, when joining your thing usually means me doing the work while others get to decide on how it should be done and feel good about that it is being done as if it was their own achievement.

That said, spite has served me well so far, so maybe it can also serve you?

This is after all a great opportunity to prove me and my worldview wrong by simply putting in the work and creating what you seem to believe is the correct form of existing.

I can only encourage bringing your ideas into reality. Seriously. That is that whole Foss spirit thing. You don't need to invite anyone (including me) to that to make it happen.

Let's manifest some code and change the world :)


I'm glad i'm not the only one that noticed w/opencode.. I tried it a few times, every single time... it immediately went bad, didn't work, can't do anything simple like simple file edits.. yet somehow, there are thousands on the internet ready to tell me how great it is! Opencode looks good in the terminal but that's really the only good thing about it.

Their reputation is “bad” but not because of privacy concerns. I personally think they’re trustworthy

We use opencode with self hosted llm for privacy reasons. Good, right? Well, no, because opencode by default uses a "free" cloud model to summarize all chats even if a different model was configured as the main one.

I wonder how many opencode users upload their private secrets to the cloud, while thinking they're using a self hosted model.

Btw. I don't think this is malicious, just sloppy.


It uses gpt-5-nano through OpenCode Zen to generate the title unless you override `small_model`. https://opencode.ai/docs/providers/#self-hosted-gitlab

Why?

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: