How should we evaluate an LLM response that is partly correct, contains unsupported phrases, or behaves differently across cultural contexts? Many existing evaluation methods rely on coarse labels or aggregate scores and therefore overlook important variations in factuality, evidential grounding, and cultural sensitivity. This talk presents three of our recent studies toward more fine-grained and
inclusive evaluation of large language models. First, I introduce an agentic framework for graded factuality verification that acquires external evidence and assigns scalar factuality scores. Second, I present a method for detecting hallucinated spans while aligning faithful output tokens with
supporting evidence in the input. Finally, I briefly introduce a multilingual benchmark for evaluating entity-centric cultural biases across Asian languages and cultures. Together, these studies highlight
the need to evaluate not only whether LLM outputs are correct, but also how correct they are, what evidence supports them, and how reliably they perform across cultural contexts.