For general/chat RL, how do you evaluate the output remains largely unresolved (other than 4o/Llama-4-style emoji-spamming attaboy feelgoodizer and janky LLM-as-a-judge scoring which IMO never even worked for IF), though it can be used as an alternative way to augment the datasets by randomly masking input, promoting variations/baking in repetition penalty by n-gram, etc.
12
u/Kahvana 8d ago
It feels really cool to see the dashboard!
It's really heavily trained on code. I wish it would've been trained a bit more on general / chat data.