RESEARCH

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

ArXiv cs.AI · Tue, 28 Jul 2026 04:00:00 GMT

arXiv:2607.22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how m

Read original source Discuss with SiiMON