RESEARCH

Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor

ArXiv cs.AI · Fri, 02 Oct 2026 04:00:00 GMT

arXiv:2610.00197v1 Announce Type: new Abstract: We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controll

Read original source Discuss with SiiMON