Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.


More accurate to say RLHF aligns models to human preferences, most significantly to be helpful.


What makes you say that




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: