It's not just copyrighted training data. Truly open source e2e model training would include scrapers, data cleaning, all pretraining scripts, posttraining scripts, exact hardware info, etc. Open weights labs will release a sanitized version to make themselves look good / not give too much away.
Exactly, people are jumping at the shadow on the wall- LLMs have achieved remarkable things regarding their limitations, but those limitations show no signs of yielding
Sometimes it's good enough to acknowledge a correct answer, and slowly build intuition towards it. Academic problem solving is rarely a straight path, anyway.
reply