← People

Paul Voigtlaender

Source-listed: Research Scientist · Google

  • Multimodal AI
  • Language models

Selected work

5
  1. Image Generators are Generalist Vision LearnersContributor · 2026
  2. PaliGemma: A versatile 3B VLM for transferAuthor · 2024
  3. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextAuthor · 2024
  4. Point-VOS: Pointing Up Video Object SegmentationContributor · 2024
  5. Connecting Vision and Language with Video Localized NarrativesContributor · 2023