EleutherAI Releases OLMo-3-7B Models for Reward-Hacking Research
EleutherAI releases two OLMo-3-7B models trained via GRPO to study reward-hacking dynamics and validate the hack-igni ...
News on open-weight models you can run locally
EleutherAI releases two OLMo-3-7B models trained via GRPO to study reward-hacking dynamics and validate the hack-igni ...