
All
AI 21
Geology 8
Benchmarking 5
China 4
Machine Learning 4
Complexity Theory 3
Reinforcement Learning 3
Cryptography 1
Data Visualization 1
Grokking 1
Reinforcement Learning

Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection
TL;DR Self-training systems typically degenerate because they lack an external criterion for judging data quality, which opens the door to reward hacking and semantic drift. The paper presents a proof-of-concept …
Generalising from Self-Produced Data: Model Training Beyond Human Constraints
TL;DR Current LLMs are bounded by human-derived training data and by a single level of abstraction that prevents definitive truth judgments about their own outputs. The paper proposes a framework in which AI agents …