3 Comments

User's avatar
gwern's avatar

> We wonder whether this result holds up in 2026, but we’d bet that it does.

I would be a little more impressed if the result were on something a little larger than a 0.02b-parameter T5 without even any error bars for tiny effects from removing a few % of the total training data. I hesitate to generalize from that to any contemporary LLMs like a Fable at ~10,000b-parameters.

The frontier labs routinely investigate data mixes in extreme detail, and I am unaware of any LLM which doesn't seem to have practically memorized Arxiv, or any reports from, say, the Chinese labs that they have found their own data mix ablations to support the claim that "The Scientific Literature is Poisonous to LLMs".

Sean Murphy's avatar

It seems to me that the same flaws in the literature that are blocking LLMs from achieving an effective understanding would also be deleterious to human scientists ability to understand what's real.

1 more comment...

No posts

Ready for more?