Kshitij Mishra
I am a Postdoctoral Associate at MBZUAI, where I study how to build AI agents that can adapt and improve without compromising consistency, reliability, auditability, or alignment with requirements. My research spans reasoning and post-training, self-improving agents, continual and multi-agent learning, context attribution, and adversarial evaluation, with a particular focus on learning from feedback, transferring knowledge, and using tools safely over long interactions. I welcome collaborations on intelligence, adaptability, introspection, and trust.
During my Ph.D. at IIT Patna, I developed reward models to steer conversational agents toward politeness and empathy in user-oriented settings. We also deployed 14 multilingual systems through Sevak, an Indian-language chatbot platform serving the railway, judicial, and healthcare sectors.
Outside research, I enjoy traveling and driving around India, spending time with family and friends, and reading about history and philosophy.
जिज्ञासायाः अधिगमः, अधिगमात् अनुभवः, अनुभवात् नवाचारः ॥
From curiosity comes learning; from learning, experience; from experience, innovation. Innovation opens new questions.
Latest
All publications ↗PROBE: Learning to Audit Policy Compliance in Tool-Using LLM Agents — accepted to NeurIPS 2026.
TRACER: Trace-Regularized Training Against Indirect Prompt Injection — accepted to Findings of EMNLP 2026.
CoRe: Collaborative Reasoning via Cross Teaching — accepted to ICML 2026.
SD-E²: Semantic Exploration for Reasoning Under Token Budgets — published in Findings of EACL 2026.
Research Interests
Consistent: evidence, beliefs, and decisions.
CAST connects self-improvement and adaptation with consistent reasoning, strategic interaction, and AI systems that remain reliable, auditable, and trustworthy.
Current questions
Learning from textual feedback. Reasoning through investigation. Faithful attribution. Strategic interaction. Safe tool use, prompt-injection defense, and adversarial auditing.
Looking ahead
Vision–language agents, memory grounded in evidence, cross-modal belief revision, and reliable long-horizon interaction.