HuatuoGPT 3 study expands medical reinforcement learning tests
An expanded HuatuoGPT 3 paper reports medical benchmark gains while warning that clinical validation remains necessary.

An expanded HuatuoGPT 3 paper submitted on October 5 adds analysis and larger model tests to earlier OnePO research. The method adapts language models to medicine without a separate supervised fine tuning stage.
OnePO uses teacher responses as temporary guidance during reinforcement learning. It strengthens learning from informative tokens, then removes teacher answers when the modelโs responses earn equal or higher rewards.
In controlled tests, the authors report better medical benchmark results than supervised training followed by reinforcement learning, using the same data and teacher source.
The expanded training produced 9B and 27B variants. The authors report 71.4 on HealthBench Professional for the larger model.
The project repository provides model weights, training code, data and a grader. This paper extends an existing project rather than establishing the familyโs first release.
Imperfect teachers and rewards remain limitations. The authors describe research prototypes requiring professional validation before clinical deployment. ByteForward has not reproduced the results.
Archival hospital photograph by Peachyeung316, taken in 2023. Resized and converted to WebP under Creative Commons Attribution ShareAlike 4.0. This illustrative image has no established connection to HuatuoGPT or its evaluation.



