I have stopped reading Ed Zitron quite a while ago because of his long-winded and polemic style. But counting his failed predictions is maybe not that useful, in the sense that his core thesis is really just the AI Bubble. While it hasn't burst, all of his predictions remain wrong; when (if) it does, he'll have been right. Trying to guess the exact time or failure mode (whether it's open models or corporate sticker shock or a bond crisis) is a fool's errand when there are so many powerful actors all-in on keeping up appearances. Exposing the financial and corporate shenanigans of the AI world is nonetheless useful. I just wish someone would do it in a more level-headed way.
Note to self and others who have read Richard Preston's "Hot Zone": The 1989 Reston (US) incident was NOT caused by a single-base mutation of Ebola, but by a distinct monkey virus. Preston seems to have made that up to dramatize events.
I read that book when I was probably way too young for it, and that incident made me spend a lot of time staring at the ceiling in the middle of the night. I'll have to dig into it more!
Cheating and overfitting, as discussed in the article, are the most obvious problems with benchmarking LLMs. But there is also the aspect that, at least for closed models, the tokens still have to be sent to the provider's servers for inference. This makes the holdout set not as held out as it may appear. OpenAI and Antropic probably don't care about your private set of regex benchmarks, but for the headline "closed" benchmarks, I'd be surprised if they haven't collected a nice representative set of "holdout" problems to be examined at leisure.
... which is entirely unsurprising given that exhaled air is about 50.000 ppm CO2 and can vary by several 10.000s depending on depth and rate of breathing. I actually consider the recent wave of findings that CO2 levels as low as 500-1000 ppm measurably affect cognitive performance and well-being to be a great example of how you can prove literally anything with statistics and a sufficiently small sample size.
> I remember when Microsoft was the new darling not many years ago, because of VS Code and WSL
I was genuinely puzzled by that, actually. I thought it quite obvious from the start that Nadella is no longer interested in Windows and other Microsoft software as products and will be moving them to thin cloud wrappers, but for some reason people were really optimistic about the "New Microsoft".
"A guy on HN told me one time, 'Don't let yourself get attached to any cloud services you are not willing to walk out on in 30 seconds flat if you feel the heat around the corner.'" -- Robert de Niro
Theoretically yes, practically no. The ECJ can order the revision of national laws, but the country in question is responsible for implementation, and can send plaintiffs on a multi-decade merry chase. Several countries have also taken the view that they can refuse changes to their constitutions. This stands on shaky ground legally, but there is no real enforcement mechanism anyway.
It's not a crime if you do it with AI.
reply