RESEARCH

Aligned Data Can Induce Misalignment via Context Confusion

ArXiv cs.AI · Thu, 01 Oct 2026 04:00:00 GMT

arXiv:2609.38379v1 Announce Type: new Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-depend

Read original source Discuss with SiiMON