Hacker Newsnew | past | comments | ask | show | jobs | submit | typon's commentslogin

I am using a model that runs on a 100 GB300s, near AGI, all bets are off, $10B training run on a 1GW cluster, and it can't realize that when I told it to "please implement v2 of feature X" that I mean delete v1, not support v1 and v2 together, in a weird Frankenstein's monster of the two. Sorry, but I think my job is quite secure for the forseeable future.


This is a perpetual pet peeve of mine with LLMs. They will always opt to ensure "backwards compatibility" with a codebase built 30 seconds ago.

I always have to explicitly state, "we are making a clean break to v1.0, do not implement any compatibility layers".


This is hilarious and blindingly obvious in hindsight that an LLM would fall for that


If there really was a "simple" solution to Fermat's Last Theorem, Andrew Wiles wouldn't have achieved the important result he did, ending up making connections across disparate fields of math.

The LLMs "sweeping up" easy, or previously missed, results seems like a net negative. It's probably better for humans to struggle and come up with new tools than to just "clean up" low hanging fruit that doesn't add much value to the field.


By early 2028, major players like Intel, AMD, QC will ship accelerators in consumer laptops capable of running ~1T MoE models at ~100 tok/s


Literally the only way this is going to happen is if aliens come to earth and gift us some amazing technology.


Yes that's called Mythos 2 or GPT 6


Unless there are major improvements to how much hardware it takes to run a 1T model, this is deeply unrealistic. First because why release hardware that puts your biggest customers (data centers) out of business. Second because as I understand it the data centers have bought up all the high end chip production capacity for at least the next year and unless the bubble pops that'll continue for a while.


Because for the company that will actually do it, their biggest customers aren’t data centers they are iPhone owners.


First off the math doesn’t math. Datacenters are willing to pay $50k for a single high end GPU. If you have unlimited capacity, yeah sell millions for $100 a pop or $10 a pop or whatever the bom cost of a phone GPU would be - but if you have limited capacity, you’re gonna sell all of that to the customer who is willing to pay the most PER UNIT.

Second off, this doesn’t work from a power consumption standpoint. When I run qwen3.6-35b, a far smaller model than op is suggesting, power usage spikes to 150-200W during inference. To fit a 1T model in the palm of my hand, the amount of processing required doesn’t fit the amount of power available.

Now I’m not saying this will never happen - there are some great leads, e.g. burning models directly on to a chip - but op’s scenario is definitely not happening in two years. Maybe 5, a lot more likely 10, unless of course local ai is made illegal


There is ton of room for improvement "down there".

* Software inference optimizations

* Heavy quantization

* Chips with hardcoded transformer architecture

* Much cheaper HBM

* Much sparser models - 1T total with ~1-10B active params e.g.

* Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.


> * Software inference optimizations

Absolutely. I'd be surprised if they couldn't 2x performance in the next year. Still doesn't make a 1T model fit on your phone.

> * Heavy quantization

I think this is a dead end if you're trying to fit a 1T model into a phone. Makes much more sense to train a model that's designed to be small, than train a model that's smart and then quantize it into stupidity.

> * Chips with hardcoded transformer architecture

Totally, this will probably work great. Now good luck booking fab time any time in the next 2 years.

> * Much cheaper HBM

Totally, this will probably work great. Now good luck booking fab time any time in the next two years.

> * Much sparser models - 1T total with ~1-10B active params e.g.

Fewer active params helps with the speed of token generation, but if the whole model doesn't fit into ram it doesn't solve the issue of having to constantly stream portions of the model from disk to ram.

> * Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.

IMO this is a delusional myth-making idea being sold to us by ai companies. Machines that generate output based on statistical averages won't generate genuinely new ideas. They can help us try out ideas faster, but they're simply not capable of the kind of creativity and understanding required to push a field forward, except incrementally.


We dont need god models (1T+) to do dog work. People use far more powerful models than they ever need to, but they don’t even know it yet. They create fake demand with FOMO.


You are assuming people need the models they use today. The reality is much much smaller models will suffice (i.e. dont use god models for dog work)


I’m not assuming that at all, I’m responding to someone suggesting we’ll be able to run 1T models on phones in 2-ish years.

I absolutely agree that models are going to advance on to “edge” hardware over the next few years by becoming small + specialized.


Sorry to put words in your mouth, totally agree


  > Datacenters are willing to pay $50k for a single high end GPU.
its true for now, because capital is flowing like a torrent, but how long will that last if returns start to be expected (aka the bubble pops)?


Even if the bubble pops and anthropic and openai et al implode - genie doesn’t go back in the bottle. The usefulness of LLMs for coding is proven, and a chip in a datacenter running 24/7 is always going to be more valuable than in a personal device running occasionally.

That doesn’t change until production capacity exceeds the datacenter demand. When that happens, they’ll start selling them down the market until it eventually reaches phones and toasters and whatever. But not in two years.


LLMs for coding is too small of a benefit to justify this investment, the bubble is indeed going to burst. Genie is already on its way back into the bottle.


I agree it’s too small a benefit to justify the investment, and I agree the bubble will pop. I just don’t think that means hardware prices become sane again for quite a while. I think if you half the price of a server GPU because demand from the big ai companies drops out, we’ll still have a shortage - it’ll just being going into commodity data centers to run open weight models.


DinoV3 paper: https://arxiv.org/pdf/2508.10104#page=36

"we use a rough estimate of a total 9M GPU hours"

From CoreWeave, at current prices (~$2.46/hr spot to ~$6.16/hr on demand) would correspond to $22M–$55M.

The dataset is really where the cost is though - they used LVD-1689M - 1.6B images of curated web data from roughly 17B instagram images. This probably cost a huge amount of hours in human annotation, compute for algorithmic filtering, etc and not to mention probably a 20-50 person team working on this model.

You might want to change assumptions about how expensive these models are.


Thanks for the correction on the order of magnitude for the whole training process.

The 9M GPU hours includes the DINO v2 inference used in order to curate the data set.

The final training run used like 300000 dollars of compute.

Unfortunately we don't know how much RLVR + Agent training costs these companies. I'm just gonna say it's in the hundreds of millions, because they are supposedly making billions of profit on inference yet making billion dollar losses


I remember saving up for a year to buy the ATI Radeon 9600 XT (I think it was $200 MSRP) so I could play the game on high settings. Now we can play it inside a virtual machine on a crappy laptop. What a journey


In a few years todays high end AI models will run on your watch

Of course that assumes we maintain open access to compute that we've enjoyed for the last half century, and I doubt that very much.

Stallman warned about the dangers of software being closed [0] 30 years ago, and the majority of modern IT industry just laugh a that sort of stuff because you can't make a billion dollar startup with that attitude, but I think the restrictions on owning the hardware at all will probably come first.

[0] https://www.gnu.org/philosophy/right-to-read.en.html


> In a few years todays high end AI models will run on your watch

Although possible with cpu power, I dont think you will ever get enough ram in a watch to run a decent local LLM.

I also dont think the high ram requirements for running them will come down at all.


I remember when the game files got hacked before release, and you could run around in half completed maps and small area snippets. I spent hours running around in awe of the new physics engine


I was just going to say the same thing. I couldn't afford the rigs needed to run any of these games and never really played them. Now, it's running inside a browser on a laptop.


…while inside Jira working on a ticket.

https://github.com/wjkennedy/jira-quake3


Same here - splashed out crazy money upgrading my PC to play HL2.

After that moment I switched to consoles.


Always pains me to see all this innovation and cool software work being applies towards making a machine that extracts money from retail investors and makes a few people extremely wealthy. What a depressing and colossal waste of time.


Actually I think AI will largely automate software and math and really not much else in the short to medium term. (speaking as a computer/math person)


I don't see why that should be the case. The only reason software is getting focused on first is:

1. Software devs are obviously going to have a better idea how to apply AI to software development compared to other fields. So of course the coding tools are going to be the first things made.

2. Formal verification makes the problem easier by allowing for iterative feedback (compilers, proofs, etc.)

The second argument is, I think, somewhat valid, but ignores that a lot of other professions also have similar verification systems even if they're a bit less rigorous. The first argument just explains why things are the way they are now, it's not indicative of the future. I don't want to fall into the trap of thinking that other jobs than mine require less cognitive horsepower or whatever, but I don't see what's particularly special about other jobs if it can do hard STEM stuff.


I thought the same but i dont think so anymore. My wife is a senior manager at a big 4 consultancy gig and she says copilot became freaking good at understanding tax, multinational company structure etc etc. Even if you need a few partners and experts at the top to validate things you can cut huge amounts of workers.


Exactly. Regarding software, it is trained on a massive corpus of code and the feedback loop can be very fast (playing well into LLM's upsides) and results are ... mediocre.

Recently I had to go through some building regulations and Claude's advices were catastrophic.


Next time there is a fire at your house I will say "he's an adult who should have been careful playing with dangerous things like fire, we shouldnt waste society's money and resources on saving his house"


Sure! I can also be an adult, have fire detectors and insure my property against fires instead of hoping for the goodwill of the community or the government.


Fire insurance doesn't do anything for your house regarding it being on fire.

Fire departments are good for the community at large as well so the fire at your house doesn't become the fire at my house.


It's not goodwill though, I'm paying for public services like the fire department through my taxes. I like them to be owned by the government instead of a private entity because I don't want to pay the capitalist rent for borrowing their money; if we pool our resources we can cut out the middle-man and just fund it ourselves. Very typical human living arrangement I believe.


Just one thing: the government is a massive private entity and a middle man. The closest thing to the stated goal would be a contract between other insterested parties to pool resources for funding. The next one would be a local consumer co-op explicitly formed for that specific purpose.


Besides fire detectors you might also want some fire putter-outers as well. Oh, wait...


Can someone explain to me why local governments are so against datacenters? It seems like a golden opportunity to build electric infrastructure that's paid for by corporations and if AI is a bubble at least that infrastructure will remain and continue to provide cheap power.


Existing power infra is likely fine enough for current demands. And building new infra doesn't necessarily mean it will be any cheaper to operate or use. So with datacentres providing rather little local economic activity after being build and potential impact on electricity costs say during night overall they are not that beneficial.


It absolutely is not, power infrastructure throughout the US has been massively under-invested in for decades.

Nobody wants to admit this though because they hate AI and what it represents more than they want to rationally think about infrastructure.


Pretty sure its a populism/slopulism thing. I doubt anyone in govt ACTUALLY cares outside backlash thats prevalent in the media/socials so its easy to capitalize on the "we are doing something for the voters".


If your town users a peak of say 100MW and is powered by two 50MW feeds

Then a new DC creates another 300MW of demand and builds 300MW more feeds

Then the DC goes bust, you're left with 350MW of potential supply and 50MW of demand

Compare with a highway, you had a 2 lane road, and it was fine, with 1000 cars an hour, then someone expanded it to 8 lanes and filled them with another 3000 cars an hour.

Then they vanished, and you're left with an 8 lane road for 1000 cars an hour, paying 4 times the maintenance for extra unneeded capacity.


You're kidding yourself if you think people are laying down in front of bulldozers out of concern for trivial maintenance costs on under-utilized grid capacity.


The assertion was this was a "golden opportunity"

In reality there's nothing to gain in theory, and in practice it correlates with higher energy costs


> golden opportunity to build electric infrastructure that's paid for by corporations

The discussion here is precisely because it is not fully paid for by the corporation.


Read the article maybe? Data center corporations are not the ones getting billed for grid upgrades.


That part should be non negotiable - the tech companies should be funding the energy infrastructure. My bet is that they'd happily do it. They are starved for space and energy, not capital.


Fear of change due to AI, masquerading as concern for the one thing about technological progress that a local city council has the power to obstruct: building physical buildings (classic NIMBY-ism).

Basically every AI company using "catastrophizing" and "ragebait" as their marketing strategy is working so well that normies are afraid they're going to lose their cushy do-nothing desk jobs. Hence the braindead/conspiracy narratives that data centers are going to drink all the fresh water, give your kids cancer, kill the plants and quadruple your electric bill.

What's most hilarious about this; dramatically expanding the power grid is absolutely necessary to get off of fossil fuels, yet the same people who used to scream about climate change are now trying to block grid upgrades paid for by data centers. With zero awareness of the irony.

This will only get worse as time goes on because first world countries are aging at increasing rates. Old people hate change. They're deficit spending their children's future labor while obstructing the creation of anything productive that might dig us out of the hole they put us in.


What is the point of building energy outside of solar farms? I'm sincerely asking


An inexhaustible 24/7 production capable plant has many advantages over solar and maintaining large most types of battery banks.


Cost is like 90-99% of what matters. Last year, China installed 300GW of new renewables and 0GW of geothermal, despite geothermal being "an inexhaustible 24/7 production capable".

Geothermal will compete with solar if they can get the cost low enough. I hope they succeed!


Some people don't really want the planet covered in solar panels. Others are cool with it so long as they don't have to see it.


Night time? But batteries! Several cloudy days in a row? More batteries! Cost? -> a mix of sources becomes attractive


Batteries? You mean like digging a hole/making a wall and pumping water?

https://en.wikipedia.org/wiki/Taum_Sauk_Hydroelectric_Power_...

Or we could just spin some rocks around.

https://www.torus.co/torus-flywheel

Rather have solar/wind and this than shipping depleting goop from foreign nations.


https://imgur.com/a/dV8gk3R

can you find curves like this for any other power source?

also batteries are getting exponentially cheap too


These are typically representative of cost performance per watt of one part of a more complex deployed energy system. Things like the aluminum / steal for the container / framing, copper / aluminum for the transmission and wiring, land and labor for installation decline at much less aggressive rates or increase over time.

In almost all pareto optimal least cost energy system models that I've seen, high penetration of solar, wind, batteries plus some minority amount of (clean) baseload power is the most capital efficient energy system.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: