Evaluating Long-Form Text
How reliable are the metrics and model judges we use to score generated text? I study self-inconsistency in LLM-as-a-judge frameworks and what it means for benchmarks built on top of them.
PhD Candidate · Computer Science
University of Illinois Urbana-Champaign
I am a PhD candidate in Computer Science at the University of Illinois Urbana-Champaign, advised by Prof. Julia Hockenmaier in the Hockenmaier Lab. My research sits at the intersection of natural language processing and software — I study how language models understand and describe code, and how we should (and should not) measure the quality of the long-form text they generate.
Most recently I have been working on the reliability of LLM-as-a-judge evaluation, showing how much of a model's verdict is signal and how much is noise it invents on rerun. Earlier work spans semantic code search, code summarization, and event extraction — some of it done during research internships at IBM Research, where it also became three granted US patents.
Before Illinois, I earned my B.Tech. in Computer Science and Engineering from the Indian Institute of Technology Kharagpur.
How reliable are the metrics and model judges we use to score generated text? I study self-inconsistency in LLM-as-a-judge frameworks and what it means for benchmarks built on top of them.
Semantic code search, code summarization, and matching natural language to programs — including what large language models actually capture when they explain a function.
Event extraction, schema induction, and knowledge graphs — recovering the structure implicit in unstructured text so it can be searched, compared, and reasoned over.
Full and continuously updated lists are on Google Scholar, DBLP, and the ACL Anthology.
Findings of the Association for Computational Linguistics: EMNLP 2025
Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)
58th Annual Meeting of the Association for Computational Linguistics (ACL)
IEEE Transactions on Knowledge and Data Engineering, 32(11)
51st ACM Technical Symposium on Computer Science Education (SIGCSE)
15th International Wireless Communications & Mobile Computing Conference (IWCMC)
2018 Conference of the North American Chapter of the ACL: Demonstrations (NAACL-HLT)
Granted US patents, all assigned to International Business Machines Corporation.
University of Illinois Urbana-Champaign · Urbana, IL
Advised by Prof. Julia Hockenmaier
Schlumberger · Menlo Park, CA
IBM Research · Yorktown Heights, NY
Indian Institute of Technology Kharagpur · Kharagpur, India
University of Illinois Urbana-Champaign · Urbana, IL
University of Illinois Urbana-Champaign · Urbana, IL
Happy to talk about LLM evaluation, code intelligence, or anything adjacent.
rhaldar2@illinois.edu