← Publications

Publications / January 2026

AAAI-26 Undergraduate Consortium

A curious robot examining a question about language.

Can You Trust What I Think? Analyzing and Improving Verbalized Uncertainty and Factuality in Reasoning-Based Large Language Models

Tianruo Rose Xu.

Abstract

Reasoning-based large language models often produce natural-language thinking traces with their answers, but it remains unclear whether the verbalized uncertainties expressed in thinking traces faithfully reflect model’s knowledge. We study this question on long-form, knowledge-intensive biography generation. Our pipeline decomposes thinking traces and responses into atomic facts, filters out planning content, labels factual reasoning by certainty, and aligns response facts to their supporting reasoning, enabling plan-based filtering, self-verification, and a classifier that predicts factuality from facts and associated reasoning. Preliminary results suggest that high-certainty reasoning is more likely to be included and correct and that structured use of these signals can improve factuality, though broader validation across models and dataset will be needed.

Published in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2026), 40(48), 41534-41536. Presented at the AAAI-26 Undergraduate Consortium in Singapore, January 2026.

Poster Β· Oral presentation

Cite this work

Tianruo Rose Xu. Can You Trust What I Think? Analyzing and Improving Verbalized Uncertainty and Factuality in Reasoning-Based Large Language Models. AAAI-26 Undergraduate Consortium, Singapore, January 2026.