← People

Siwei Han

Source-listed: First-year Ph.D. student · University of North Carolina at Chapel Hill

  • Multimodal AI
  • Language models
  • Information retrieval

Selected work

5
  1. MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsContributor · 2024
  2. MDocAgent: A Multi-Modal Multi-Agent Framework for Document UnderstandingContributor · 2025
  3. Generating Chain-of-Thoughts with a Direct Pairwise-Comparison Approach to Searching for the Most Promising Intermediate ThoughtContributor · 2024
  4. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?Contributor · 2025
  5. MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video GenerationContributor · 2025