Research
I research how to make language models more reliable when their answers affect real decisions.
A confident hallucination can matter far beyond a bad answer.
As AI moves into everyday products and high-stakes domains such as finance and payments, hallucination and overconfidence become practical reliability problems, not just benchmark failures.
I study how models behave under uncertainty, how post-training changes that behaviour, and whether undesirable changes can be predicted, detected and selectively controlled.
Research interests
Reliable language models
How to make model behaviour dependable when people rely on it.
Hallucination
Why models produce convincing answers that can mislead people.
Abstention
When a model should stop, defer, or simply say “I don’t know”.
Post-training behaviour
How training changes behaviour beyond what developers intended.
Safety
How useful capabilities can become harmful in the wrong context.
Knowledge and reasoning
Why models can learn something and still make consequential mistakes.
Selected first author publications
Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning
Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo, Erik Cambria, Xiuzhen Zhang
Accepted to EMNLP · 2026 · A* conference
A model can be confident and still be wrong because its memory reflects an older world.
Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Xiuzhen Zhang
Proceedings of the AAAI Conference on Artificial Intelligence · 2026 · A* conference
When a question feels familiar, a model jumps to the first reading that fits. I bring the other readings back before it commits.
Bootstrap Wayfinding Questions to Elicit Emotion Shift Reasoning with Large Language Models
Vy Nguyen, Xiuzhen Zhang, Feng Xia
IEEE Transactions on Affective Computing · 2026 · Q1 journal
When emotion changes in a conversation, the question is what caused the change, not what each sentence felt. I start from that change and reason towards its cause.
Emotion Flip Reasoning via Stacked Instruction Finetuning of LLMs
Vy Nguyen, Xiuzhen Zhang
Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024) · 2024
Naming the trigger in one prediction hides the judgments that identify it. I train those judgments in steps.
A more complete list is on Google Scholar.