RESEARCH

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

ArXiv cs.AI · Fri, 14 Aug 2026 04:00:00 GMT

arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents m

Read original source Discuss with SiiMON