Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
MA-RLHF is a reinforcement learning framework that improves human feedback training by using macro actions with configurable termination conditions. It provides official code and hyperparameter configurations for training language models like Gemma on datasets such as TL;DR and HH RLHF.
Parse Score