🚨 LLMs can invent genuinely new things — here's the technical reason
Hard problems are novel compositions of known patterns and building blocks. The proof of an open theorem isn't in the training data, but every lemma and technique is.
Learn the operators, compose them in unseen orders, solve "untrained" problems.
Proposing is hard; checking is cheap. Run a massive test-time search, filter with a verifier, keep only the truth.
Concrete examples:
AlphaProof hit IMO silver by searching proof space with Lean as the verifier
FunSearch found new results on the cap set problem - genuinely new math, from an LLM proposing + a checker filtering
AlphaFold-style models generalize to proteins that labs haven't crystallized, because they learned the rules of protein folding
Biology and energy have exactly this shape: gigantic search spaces (proteins, materials, catalysts) + cheap verification (simulation, assays). Same with energy as well.
Nature's rules are in the distribution. And once you've learned the patterns of the universe, you can pretty much crack whatever can be cracked, including infinite energy or a cure for cancer 🚀
中文: LE0 LLM 可以发明真正新事物——以下是技术原因
难题是已知模式和构件的新颖构图。开放定理的证明不在训练数据中,但每个空数和技术都是。
学习操作者,以看不见的顺序撰写,解决“未经训练”的问题。
提出是困难的;检查成本低廉。运行一个大规模的测试时间搜索,使用验证器进行筛选,仅保留真相。
具体示例:
AlphaProof 通过以 Lean 为验证者搜索证明空间,成功获得 IMO 银牌
FunSearch 在设置上限时发现了新的结果——真正全新的数学,来自 LLM 的推荐 + 一个检查器筛选
AlphaFold 式模型对实验室尚未结晶的蛋白质进行通用化,因为它们学会了蛋白质折叠的规则
生物学和能量具有这种形状:巨大的搜索空间(蛋白质、材料、催化剂)以及廉价的验证(模拟、测定)。能量也是如此。
自然的规则在于分布。一旦你了解了宇宙的规律,几乎就能破解任何可能被破解的东西,包括无穷能量或治愈癌症的方法🚀