Hacker Newsnew | past | comments | ask | show | jobs | submit | pizza's commentslogin

ironically since the swarm behavior can take place during rl training then the model could also be teaching itself to keep doing it more, as well as making the internet itself a place where this becomes more likely

They’ve provided insane value to the ai community. Imo is there really any obvious way to standardize model/preprocessing/inference/training/rl recipes across the space of frankenstein model family/architecture/modality combinations? It’s not even clear to me that the current frontier would be where it is today without HF doing what they did how they did it. Where else were people widely sharing datasets w that ease of reuse/redistribution, or sharing their retuned models etc? Would there even be as much interest/tinkerers today?


I think both can be true at the same time. They provided insane value, and while I'm not sure if they were first, they were most people's introduction during the initial LLM boom. Their code was also sloppy, and the org was mismanaged. Sometimes, success really is just 'right place, right time'


Compression is just counting. Probability, also, pretty much, just counting. For these reasons I think the role of information theory in describing the process of the development of reasoning and the gain of understanding has been overstated.


you can say the same thing of the watts in a person too


Had to butcher the title slightly to get it under the limit- original:

Compression And Decompression Under FHE Using Error-Correcting Codes and Copy-And-Recurse


For most tasks, at some future date, isn't there going to be some ambient baseline of capabilities you can get per $/tok, starting at ~0 for OSS models, such that eventually all tooling gets trivially transferable?


OP is correct; surprisal is outcome-dependent and entropy is distribution-dependent

- entropy is E_p[informativeness of measuring outcome x]

- take n outcomes, then a distribution over them lives on the simplex \delta ^ (n - 1). you can lift this to R^n via the log odds map p_k -> x_k = log p_k -- now x \in R^n can describe a histogram with n-1 degrees of freedom

- in log odds space, measurement is literally a linear functional from vector space of log probability onto the index of the outcome k.

- imo surprisal of some p(x) is best understood as "the length of a pointer", entropy "the rarity-weighted average length of a pointer", and collision entropy "how specific you would have to be to describe witnessing a specific outcome"

and in the same way, a single molecule of water, you might get by, calling dry


There’s a Dark Forest problem for evals. As soon as they’re made public they start running out of time to be useful. It’s also not clear how to predict how the model will perform on a task based on an eval. Or even whether, given two skills that the model can individually do well on in the evals, it still does well on their composition. It might at this point be better to be scientific in unscientific approaches, than to attribute more power to relatively weakly predictive evals than they actually have


I agree with your analysis but not the conclusion.

Evals are broken - OpenAI showed that SWE Bench Verified was in the training data - models were able to reconstruct the changes from memory (https://openai.com/index/why-we-no-longer-evaluate-swe-bench...)

However, this doesn't mean we should completely give up on benchmarking. In fact, as models get more intelligent, and we give them more autonomy, I believe that tracking agent alignment to your coding standards becomes even more important.

What I've been exploring is making a benchmark that is unique per-repo - answering the question of how does the coding agent perform in my repo doing my tasks with my context. No longer do we have to trust general benchmarks.

Of course there will still be difficulties and limitations, but it's a step towards giving devs more information about agent performance, and allowing them to use that information to tweak and optimize the agent further


Someone else already wrote it, but it's just too funny to not abuse:

Evals are bad because people learn and fit to them. So we do extremely small evals instead.


Is "Dark Forest problem" an actual name? I just heard of the hypothesis and it has nothing to do with how you used it in this context.


I meant in the sense of - you have benchmarkers and trainers. If you publicize your evaluation, trainers may likely have their models 'consume' it, even if only indirectly: another person creating their own benchmark from scratch may be influenced by yours, even if the new question sets are clean-room. That, and the rule of thumb that benchmark value dissipates like sqrt(age) [0]

So there is a definite advantage to never publicizing your internal benchmark. But then, no one else can replicate your findings. You should assume that the space of benchmarks that are actually decent at evaluating model performance is much larger and most of the good ones, the ones that were costliest to produce, are hidden, and might not even correspond very well with the public ones. And that the public expensive benchmarks are selective and have a bias towards marketing purposes.

[0] https://www.offconvex.org/2021/04/07/ripvanwinkle/


I believe the correct term is "Goodhart's Law": https://en.wikipedia.org/wiki/Goodhart%27s_law


In Singapore it seems 80% of people live in public housing https://en.wikipedia.org/wiki/Public_housing_in_Singapore though I can't speak as to what the effect is on its housing market


I mean. Sounds like the guy had existing long term goals, needed to overcome an activation threshold, and used AI as a catalyst to just get started. Seems like, behaviorally, AI was pivotal for him to learn things, even if the things he learned came from elsewhere / his own effort.


I suppose, yes, AI was like a kickstart. But the point is - he didn't just stick to AI, he realized that in terms of skill and fulfillment it's a no-go direction. Because you neither learn anything, nor create anything yourself.


I feel the same way. But this is a new economy now, software is cheap, and regarding the skill and fulfillment you derive writing it yourself, to quote Chris Farley: "that and a nickel will get you a nice hot cup of JACK SQUAT!!!"


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: