Siwei Han
Source-listed: First-year Ph.D. student · University of North Carolina at Chapel Hill
- Multimodal AI
- Language models
- Information retrieval
Selected work
5- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsContributor · 2024
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document UnderstandingContributor · 2025
- Generating Chain-of-Thoughts with a Direct Pairwise-Comparison Approach to Searching for the Most Promising Intermediate ThoughtContributor · 2024
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?Contributor · 2025
- MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video GenerationContributor · 2025