Hacker Newsnew | past | comments | ask | show | jobs | submit | kanemcgrath's commentslogin

Or Eldin ring naming your horse "Torrent" so that searches for "Eldin Ring Torrent" would bring up your horse instead of pirate downloads

I find it unfortunate that tele-operated robotics could be a legitimate and successful strategy. But everyone is so desperate for AI hype money that I don’t think we will see it for a while.


I really think dwarf planet is a fine classification. Pluto is a planet, just a really small one. Same with Eris and the other planets in the Kuiper Belt. I don’t know why the planet count of our solar system is seemingly such a Religious number. It seems as if it is the focus of why we should define what a planet is.

When I think of what a planet is, I think of a destination. Somewhere that maybe in the far future we could visit one day. I think the definition for a planet should be a body with sufficient gravity that we could reasonably orbit and land on. And we can have as many sub-classifications as we want from nano-planet to Mega-planet.


Googles local gemma models which target roughly the same parameter count range, are known for being a lot better at vision tasks than qwen, no idea if 3.8 has changed that though


Really? Gemma4-31B should be better than Qwen3.8-27B? I'm happy to test that.


I think I am going to buy a second rtx 3060, as 27B has been just outside of my range for to long, and this looks like the parameter count tipping point


Running it on 2x3060 now. Works pretty well but VRAM is tight. 4bit quants. 1x128k context, 8bit KV, MTP on.


whats the tok/s you get on that. I have heard a few claims of around 30-50 with mtp, but for how cheap the setup is I am surprised I don't hear more about 3060 stacks so I assume there has to be some catch.


I used GPT-5.6 Sol high to optimize it, and it claimed it was getting 50. I'm seeing ~40 on my goto smoketest: "Make me a vector add in CUDA".

Funny side note. It successfully one shot the program, but it wasn't able to run it because there literally wasn't enough VRAM left to allocate CUDA memory. Watching it try to debug that was fascinating. I'm pretty sure it would have killed the llama-server (and thus itself) if it hadn't been running in a separate container.


Because of all the svg rendering stuff, I added a draw_svg tool to my harness and it has been really nice to get a quick mock-up of ui changes. And conveniently, the new DeepSeek models are really good at knowing when to use it. So I do look at pelican rendering as a small metric of useful capability.


I tried it for a bit, and It was not really worth its size. It got swept up in all the other AI news recently, but laguna s 2.1 I think is the best ~100B moe model right now


I didn't mention it above, but Laguna S is my other favorite model. I use Qwen a lot more, it's smaller and faster, but I like to switch to Laguna when I feel like I need a "heavy hitter" for certain huge or complex tasks.


What on earth hardwares do you guys have to be able to run 100gb models locally?! That's crazy! I'm here struggling to even get 27b models to run in somewhat usable way


Haha I'm on an Mac Studio with an M1 Ultra, 64gb ram. I bought it when it first came out, it just happens to be good for local LLMs. I have to use a smaller quant of Laguna S though (I think 4-bit? Not at my machine to check), as 8-bit and full size definitely don't fit in the 64gb I have.


Yeah, a good rule of thumb is that the weights take up ~100% of the size of the model, so 100B bytes (8-bit quant) would be, well, 100GB and a 4-bit quant would be half that.


I don't know if that's, well, a rule of thumb, it might be, well, straight multiplication.


Well, yeah, the straight multiplication, er, well, "follows" from the rule of thumb. Hope that helps!


Saying, well a rule of thumb is, well, 100 billion bytes is a 100 gigabytes, is well, not a rule of thumb. It is, well, just the common definition.


Right. The rule of thumb is that the overhead size of the model that's not the weights is so vastly outweighed by the actual number of weights that it can be disregarded. My shorthand for that was to write "the weights take up ~100% of the size of the model". What then "follows", both in the sense that the explanation is written after the rule as well as that it logically follows, is that, well, 100 billion bytes is, you know, 100 GB. I don't see why you're, like, paying so much attention to this?


Right. That's, well, not a rule of thumb. Asking how much of a bottle of water is, well, water, and someone says "well, a good rule of thumb is that it's all water", is just an answer. You don't need an estimate when, well, there is nothing to estimate.


Yeah, well, that's just, like, your opinion, man.

Anyway, more seriously, I hope it's obvious by now that I don't particularly care that my means of communication is so offensive to you. I think you should, like, cry a river, build a bridge, then, well... get over it, you know?

And on the, er, "topic"? A rule of thumb for me does not have to be one for you, even if it's explicitly presented as such a rule. I thought that would be obvious but, well, here we are. Anyway, how's life been treating you?


A rule of thumb is, well, an estimation. 1 = 1 isn't an estimation, it's just, well, the same number.


At this point, what isn't much of an estimation anymore is that you are here to explain things that nobody's asked for, and like to assume people around you have been waiting for your pearls of wisdom. You're unable to realize when this isn't so. You're so full of yourself that you don't notice.

Also, in this particular thread, you started by wrongly making a correction of something that was clearly not an error, nor a wrong use of words, not even a misspelling. The problem is you started to post a correction before realizing that you didn't read right. That happens when one is more eager to boast one's own greatness than one is interested in the topic at hand. The result is that in this thread, you were clearly, as a rule of thumb, well, 100% wrong.


I understand that's you're frustrated from some deeply held software beliefs not holding up to examination, but making things up in a different thread is not a healthy way to work towards acceptance.


Oh, that is a useful rule to know! Thanks!


That's not exactly the math. Theres also vram needed for context. I operate several 72-128 GB machines and the larger the context the slower they go.

And the context takes space +kv cache. KV cache drives usefulness as your context grows, it needs to pull the kv cache.

Simplified, the context has to be run on every turn, so the KV cache supplies the processed tokens, so it just needs the new inpute.


brb, going to see if 2nd hand mac studios are available!


Strix Halo, 128GB RAM. I got a refurbished Corsair AI Workstation for a smoking price ($2100) about two months ago. Lucky timing that it was in stock.


Strix Halo as well. Bought it for $1,800 new on sale and shoved an extra 4tb drive into it. Been amazing for local AI. Maybe not the absolute fastest thing (usually around 30t/s depending on the task) but has been awesome for a local AI box that I can solar power.


Nice. Mind sharing the solar side of your setup?


Couple of rack mount batteries and roughly 5kw of solar panels. Feeds into a subpanel so I can flip it when I want a couple rooms of solar on the house, or hook a generator up if needed. Can't power the entire house, but works well for thinks like computers, lighting, etc. And if I want to expand, just throw on more panels, or realistically, just throw on more batteries to store the juice.


Thanks - that’s what I want to do!


Usually theyre quantized. Also, there was a window where AMD 395+ W/128GB was just a high end $2500 hardware with unified gpu memory.


Yeah, here I am sitting deeply deeply deeply regretting not buying couple CMP 170HX at $200 or $350, knowing I could just flip them ethically at purchase price if nothing came of it... I could have just casually built a 128GB dual A100 local AI monster


I'm working with a lab that has a few Ampere GPUs on infiniband and they are just not compatible with the latest quants and vLLM updates. FP8 is about as low as you can go.


But they're reportedly a soft nerfed GA100 64GB/40GB at $1200, that's not more expensive and certainly can't be slower than a Mac Studio.


quantized + offload

I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s


IQ4 qwen 122b-a10b would mean 61GB total size and 5GB active, so about 5GB of the model loaded into GPURAM plus any generated context, and 61GB of weights loaded into system RAM? I don't know if that math is correct, but does that run well? Wouldn't that only leave 3GB of system RAM?


MoE models can use system memory along with a GPU.


and get high token bandwidth?


Not badly so because MoE models(identifiable by "CoolName-xxxB-AxxB" naming scheme) have bunch of branches in the middle that only one out of all gets non-zero values. Each of branches aka "Experts" as well as top/bottom parts are significantly smaller than the whole, and so CPU emulation of CUDA operations mixed with GPU taking as much as possible become not so out of question, unlike for dense models("CoolName-xxxB" without "-AxxB")


Similar to a spark, which isn't blazing fast but usable.


dgx spark, nvfp4 so I have spare room for KV cache (context)


Using qwen 3.6 27b for local coding as well and downloaded Laguna s 2.1 but haven't had time to give it a full spin yet.

Curious for any more experiences


I agree. 27b dense really did seem like the sweet spot.


My favorite one was the blender addon https://github.com/libsm64/libsm64-blender


Also the steam DRM on a lot of single player games can be removed very easily. https://github.com/atom0s/Steamless

and is a recommended step for engine modding games like skyrim.


One of my favorite things about my custom mechanical keyboard, is being able to remap the entire key set in the firmware with VIA. I have fn+arrow keys for media, fn+space for play/pause fn+end for calculator, and a bunch of random others. It is so useful I could never get another keyboard that doesn’t have a similar functionality.


> It is so useful I could never get another keyboard that doesn’t have a similar functionality.

I do it at the software level (Linux / Xorg): complete remapping, with an "hyper" key modifier etc.

The reason I do it at the software level is that you can pry my Topre switches from my cold dead hands and the HHKB Pro JP I'm using doesn't have, by default, a programmable controller. Now I know some people mod their Topre keyboards to add a programmable controller but I never got to that point.

Doing so in hardware using .xkb files is... Something. I know way more about .xkb files than I should but, thankfully, so far I've just been able to brink my .xkb file to every new Linux version (supporting Xorg, I'm not on Wayland).

I take at some point I'll look more into how to mod my HHKB keyboards with programmable controller.


I put a Hasu controller[1] in my Leopold F660C (Topre switches) and couldn't be happier. Though it voids the warranty, installation was very easy and it allows total control of the keymap via TMK. Looks like there are similar controllers by Hasu for HHKB[2]. If you can manage to order one I highly recommend it!

[1]: https://geekhack.org/index.php?topic=88720.0

[2]: https://geekhack.org/index.php?topic=71517.0


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: