One of the reasons we’re focused on training efficiencies is because those are upstream of end user costs. We achieved these specific gains while cutting context windows down 4x. That will have major implications for inference speed and expense.
For us specifically, we want to target prosumers who aren’t gonna pay $1 for a few seconds of footage. It has to be 1/50th of the cost of the big guys. It helps that our ultimate focus/niche is animation, so we can get away with smaller models.
But to answer question more broadly, I think you have to ask the questions:
1) How will this lab front run the hyperscalers? Is there someone thing they’re doing (or if executed correctly) could break right such that they could have a hook for customers to use the product over hypscaler (eg crazy low cost of inference)?
2) If this lab lands the hook and gets small lead how could they maintain and grow it? This doesn’t have to be technical per say. It could be a clear land and expand sales motion at the enterprise level, for example.
Startup history is littered with examples of companies that seemed too close to a big guy to sprout up in the first place (eg stripes adjacency to PayPal).
From the outside looking in it might seem likes there no space for the lab to bloom; so you’ll have to bring the specifics of the lab and your on the ground expertise of the space / company’s situation to come to an answer for yourself.
A lot of current diffusion LLMs don't convert tokens to continous space before noising them. They add discrete "noise", which is often as simple as replacing some tokens with [MASK].
The real problem is that when using fewer sampling steps than output tokens, diffusion formulations fundamentally cannot represent distributions where output tokens are heavily codependent. Autoregressive formulations don't have this problem, they can represent any distribution (ignoring limitations of the underlying model).
Yes. Go to chat.com in an incognito window and ask it the same question that you ask 5.6-sol with thinking max, and compare. Depending on the question there is or is not a huge difference.
Also, remember the bulk of AI users are using free models, getting terrible answers, and wondering why people keep saying AI is going to take everybody's jobs away.
This algorithm: sample a bunch of latents, train the model using the one with the lowest error.
IWAE: sample a bunch of latents, weight the loss of training the model using each one by softmax(-error). For images and text where the errors have large variance, those weights become one-hot, yielding this algorithm.
Did the model really need to hack huggingface to get access to ExploitGym data? I'd imagine that once it had full internet access it could have just used the HF API or website (but the heavy prompting/nudging towards hacking made it do things the hard way).
Though, by the time we've replicated a complete industrial hinterland (up to and including a semiconductor supply chain) which the author describes, it seems like generating fuel and engines wouldn't be impossible.
Also to the authors last point (extremely long time scales causing degradation), it seems like we'd want high thrust capabilities regardless. i.e. maybe a small gravity well doesn't gain us anything, since we'd need big engines to get up to speed anyway.
I feel like the data should have been generated by a much less predictable policy.
It often feels like the model is ignoring my inputs and just doing what it would expect the bot to do (which is unsurprising if the model could predict what would happen next during training without paying attention to the inputs)
It's always "yudkowsky is a weird guy" or "they had an orgie once" but who gives a fuck?
I just want someone to lay out, in impersonal terms, the flaws of their reasoning and why I shouldn't agree with them.
My problem is with the whole genre of article rather than this one specifically.
reply