Welcome to my homepage! I am a Ph.D. candidate in Engineering Science at Harvard School of Engineering and Applied Sciences, fortunate to be advised by Prof. Michael Lingzhi Li. Prior to that, I obtained my Bachelor’s degree in Statistics from the University of Science and Technology of China (USTC), where I had the privilege of being advised by Prof. Yanran Wang and the opportunity to work with Prof. Atlas Wang. I was honored to receive the 41st Guo Moruo Scholarship, the highest honor for USTC undergraduates.

My research asks how AI systems can perform reliably in real-world healthcare workflows. I focus on multimodal and longitudinal clinical data, LLM and agent evaluation, and trustworthy reasoning. My work includes clinical data infrastructure and interpretable medical imaging, alongside collaborations on document parsing, chart understanding, visual representations, and multi-agent reasoning.

I welcome collaborations on healthcare AI, multimodal benchmarks, and LLM/agent evaluation. Feel free to reach out at xiaolongluo@g.harvard.edu.

I am seeking research internships for summer 2027, particularly in healthcare AI, LLM/agent evaluation, and multimodal learning. Please feel free to reach out!

🔬 Research Interests

  • Reliable LLMs and Agents for Healthcare: I study how to evaluate reasoning over longitudinal clinical data, with an emphasis on fidelity, calibration, robustness, and failure analysis. I am interested in evaluation that helps us understand when an AI system can be trusted within a clinical workflow.

  • Multimodal Learning and Evaluation: I work on interpretable models for medical imaging and clinical data, and contribute to benchmarks for difficult document parsing and faithful chart understanding. These projects connect data quality, reproducible evaluation, and model reliability.

🌱 Long-term Vision

My long-term goal is to build AI systems that reliably support people in real healthcare workflows. I want to connect rigorous evaluation with useful tools, so that improvements in model capability translate into better clinical work and more accessible care.

🔥 News

  • [2026.09] New submissions on multi-agent debate, visual representations, and multimodal benchmarks; see the publication list below.
  • [2026.09] Teaching Fellow for PHIL 166AI: Artificial Agency and Society at Harvard.
  • [2026.06] Served as a challenge organizer for the DataMFM workshop at CVPR 2026, covering document parsing and chart understanding.
  • [2026.05] CRISP is now available in the ML4H 2025 proceedings (PMLR 297).
  • [2025.11] Our paper “The CRITICAL Records Integrated Standardization Pipeline (CRISP)” was accepted by ML4H 2025 as Spotlight.
  • [2025.11] Great experience presenting “Towards Interpretable, Sequential Multiple Instance Learning: An Application to Clinical Imaging” at AMIA 2025 in Atlanta! The conference was fantastic!
  • [2025.10] Successfully passed my qualification exam and officially became a Ph.D. candidate! 🎉
  • [2025.09] New paper on arXiv: “The CRITICAL Records Integrated Standardization Pipeline (CRISP)” - Check it out & open to collaboration!
  • [2025.06] Our paper accepted by AMIA 2025! See you in Atlanta this November!
  • [2025.05] Successfully completed all PhD coursework and earned CS Master’s degree en route to PhD!

📝 Publications

Authors marked with * contributed equally.

  • Beyond Solo and Consistency: Vindicating Multi-Agent Debate through Conditional Progressive Pruning
    Ruosong Ye, Caiqi Zhang, Jiahao Li, Haijun Wu, Xiaolong Luo, et al., Dimitris N. Metaxas.
    Submitted to ICLR 2027, September 2026. Under review.

  • SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
    Shuang Liang, Lejun Liao, Shiyuan Zhang, Max Cheng Zhang, Xiaolong Luo, et al., Yuan Yuan.
    Submitted to ICLR 2027, September 2026. Under review.

  • Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
    Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo, Zhenting Qi, Konwoo Kim, Longtian Ye, Xiaolong Luo, et al., Zexue He.
    Submitted to ICLR 2027, September 2026. Under review.
    [Preprint]

  • ChartNet-Bench: A Benchmark for Faithful Multimodal Chart Understanding
    Pengyuan Li, Isaac Sanchez, Dhiraj Joshi, Xiaolong Luo, et al., Rogerio Feris.
    Manuscript under review, 2026.

  • The CRITICAL Records Integrated Standardization Pipeline (CRISP): End-to-End Processing of Large-scale Multi-institutional OMOP CDM Data
    Xiaolong Luo, Michael Lingzhi Li.
    Machine Learning for Health (ML4H) 2025, PMLR 297:1592–1608, 2026.
    [Paper] [Code]

  • Towards Interpretable, Sequential Multiple Instance Learning: An Application to Clinical Imaging
    Xiaolong Luo, Hsin-Hsiao Scott Wang, Michael Lingzhi Li
    American Medical Informatics Association (AMIA), Nov. 2025

  • AI Transformers for Radiation Dose Reduction in Serial Whole-Body PET Scans
    YR Wang, L Qu, ND Sheybani, X Luo, J Wang, KE Hawk, AJ Theruvath, et al.
    Radiology: Artificial Intelligence, Apr. 2023; [Link]

  • Learning Pruning-Friendly Networks via Frank-Wolfe: One-Shot, Any-Sparsity, and No Retraining
    Miao Lu*, Xiaolong Luo*, Tianlong Chen, Wuyang Chen, Dong Liu, Zhangyang Wang
    International Conference on Learning Representations (ICLR), Spotlight Presentation, Mar. 2022
    [Paper] [Code]

Working Projects

  • Healthcare AI Safety and Reliability Research
    Investigating how to evaluate LLMs and agents on longitudinal clinical data from the CRITICAL dataset, with a focus on reasoning fidelity, robustness, and failure analysis.

  • Multimodal Benchmark Evaluation
    Contributing to Dr. DocBench and ChartNet-Bench. My document-parsing evaluation work includes submission validation, containerized scoring, metric aggregation, and reproducibility checks.

🎤 Invited Talks

  • “Towards Interpretable, Sequential Multiple Instance Learning: An Application to Clinical Imaging”
    INFORMS Annual Meeting, Seattle (2024)

📝 Professional Service

Challenge Organizer

Reviewer

  • CVPR 2026 - Conference on Computer Vision and Pattern Recognition
  • ICLR 2026 - International Conference on Learning Representations
  • NeurIPS 2025 Workshop Imageomics - Imageomics Workshop
  • NeurIPS 2025 Workshop GenAI4Health - Generative AI for Health Workshop
  • ML4H 2025 - Machine Learning for Health Symposium
  • ICLR 2025 - International Conference on Learning Representations
  • ACL ARR 2025 - ACL Rolling Review
  • Pattern Recognition - Journal reviewer

📚 Education

  • Harvard University, Cambridge, Massachusetts (2022-Present)
    Ph.D. Candidate in Engineering Science

  • Harvard University, Cambridge, Massachusetts (2022-2025)
    S.M. in Computer Science

  • University of Science and Technology of China, Anhui, China (2018-2022)
    Bachelor of Science in Statistics

🏆 Honors and Awards

Academic Awards

  • The 41st Guo Moruo Scholarship (Top 1%, highest honor at USTC) (2021)
  • National Scholarship (Top 1%, from Ministry of Education of China) (2019)
  • Outstanding Student Scholarship, Golden Award (Top 5%) (2020)
  • Chinese Mathematics Competitions, Anhui, The Second Prize (2019)

Leadership & Entrepreneurship

💻 Technical Skills

  • Programming: Python, C/C++, R, JavaScript, TypeScript
  • Machine Learning: PyTorch, TensorFlow, PyG, scikit-learn, OpenCV
  • LLM Tools: vLLM, LangChain, OpenAI SDK, Anthropic SDK
  • Evaluation and Development: EvalAI, Docker, Linux, Git, LaTeX
  • Web: React, HTML

Check out my GitHub profile for code and projects.

👨‍🏫 Teaching

Head Teaching Fellow

AM101: Statistical Inference for Scientists and Engineers (Spring 2024)

  • Instructor: Prof. Rob Howe
  • Class size: 55 students

Teaching Fellow

PHIL 166AI: Artificial Agency and Society (Fall 2026)

  • Instructor: Dr. Matthew Kopec

ENG-SCI 139/239: Innovation in Science and Engineering (Fall 2025)

  • Instructor: Prof. David Ricketts
  • Class size: 113 students
  • Jointly offered with Graduate School of Design as SCI 6272

COMPSCI 1090B: Data Science 2: Advanced Topics in Data Science (Spring 2025)

  • Instructor: Prof. Pavlos Protopapas (SEAS) & Natesh Pillai (Statistics)
  • Class size: 277 students

NEURO 240: Biological and Artificial Intelligence (Spring 2025)

  • Instructor: Prof. Gabriel Kreiman
  • Class size: 142 students
  • Course Website

CS 182: Artificial Intelligence (Fall 2023)

  • Instructors: Prof. Stephanie Gil; Prof. Milind Tambe (Harvard SEAS)
  • Class size: 138 students

Stat 139: Introduction to Linear Models (Fall 2023)

  • Instructor: Prof. James Xenakis (Harvard GSAS)
  • Class size: 83 students

Probability Theory and Mathematical Statistics (Fall 2021)

  • Instructor: Prof. Canwen Hong (Applied Math, USTC)
  • Class size: 97 students

Differential Equation I (Fall 2020)

  • Instructor: Prof. Wuqing Ning (Applied Math, USTC)
  • Class size: 156 students

📧 Contact


Last updated: September 2026