Paul Voigtlaender
Source-listed: Research Scientist · Google
- Multimodal AI
- Language models
Selected work
5- Image Generators are Generalist Vision LearnersContributor · 2026
- PaliGemma: A versatile 3B VLM for transferAuthor · 2024
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextAuthor · 2024
- Point-VOS: Pointing Up Video Object SegmentationContributor · 2024
- Connecting Vision and Language with Video Localized NarrativesContributor · 2023