Find the people behind AI papers and projects.
/
Search public work, explore its connections, and check the sources. How coverage works
4,006 saved profiles
Matching saved profiles. Coverage and source-listed affiliations may be incomplete or historical.
Search results
1–80 of 159
- Ahmad Al-DahleAffiliation: AirbnbWork: Coarse-to-fine Optimization for Speech Enhancement
- Alexei BaevskiAffiliation: Not recordedWork: data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- Andrés AlvaradoAffiliation: Not recordedWork: Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions
- Andrew Yan-Tak NgAffiliation: LandingAIWork: Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
- Annie DongAffiliation: Not recordedWork: SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities
- Arthur HinsvarkAffiliation: Meta Superintelligence LabsWork: Scaling Speech Tokenizers with Diffusion Autoencoders
- Brian GamidoAffiliation: MetaWork: The Llama 3 Herd of Models
- Can BaliogluAffiliation: MetaWork: Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
- Changhan WangAffiliation: MetaWork: Seamless: Multilingual Expressive and Streaming Speech Translation
- Charlotte CaucheteuxAffiliation: Google DeepMindWork: Evidence of a predictive coding hierarchy in the human brain listening to speech
- Cynthia GaoAffiliation: MetaWork: Seamless: Multilingual Expressive and Streaming Speech Translation
- Danny WyattAffiliation: Microsoft AIWork: Inferring Colocation and Conversation Networks from Privacy-Sensitive Audio with Implications for Computational Social Science
- Greg BrockmanAffiliation: OpenAIWork: Robust Speech Recognition via Large-Scale Weak Supervision
- Irina-Elena VelicheAffiliation: MetaWork: Improving Fairness and Robustness in End-to-End Speech Recognition through unsupervised clustering
- Jennifer DingAffiliation: DecagonWork: StarCoder: may the source be with you!
- Jin XuAffiliation: AlibabaWork: Qwen2-Audio
- Jinzheng HeAffiliation: Alibaba GroupWork: Qwen2-Audio
- Keshav DhandhaniaAffiliation: SierraWork: τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
- Mike HeatonAffiliation: Not recordedWork: Introducing next-generation audio models in the API
- Tao XuAffiliation: Not recordedWork: Robust Speech Recognition via Large-Scale Weak Supervision
- Ye JiaAffiliation: Not recordedWork: CVSS Corpus and Massively Multilingual Speech-to-Speech Translation
- Yichong LengAffiliation: Not recordedWork: NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
- Yuan ManAffiliation: Not recordedWork: AI Audio Datasets (AI-ADS)
- Yunfei ChuAffiliation: Alibaba GroupWork: Qwen-Audio
- Zhifang GuoAffiliation: Not recordedWork: Qwen2-Audio Technical Report
- Jian LiAffiliation: Not recordedWork: USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
- Ning DongAffiliation: Not recordedWork: Unified Speech-Text Pre-training for Speech Translation and Recognition
- Xuewei WuAffiliation: Not recordedWork: The SYSU System for the Interspeech 2015 Automatic Speaker Verification Spoofing and Countermeasures Challenge
- Lara TumehAffiliation: Not recordedWork: Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
- Pengwei LiAffiliation: Not recordedWork: Seamless: Multilingual Expressive and Streaming Speech Translation
- Jeff PitmanAffiliation: Not recordedWork: Simultaneous Speech Translation in Google Translate
- Nanxin ChenAffiliation: Not recordedWork: Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Qian YangAffiliation: Not recordedWork: Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Corinne WongAffiliation: Not recordedWork: Seamless: Multilingual Expressive and Streaming Speech Translation
- Ginger PerngAffiliation: Not recordedWork: Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
- Andy CrawfordAffiliation: Not recordedWork: Temporal Synchronization of Multiple Audio Signals
- Ben LimonchikAffiliation: Not recordedWork: Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Chunyang WuAffiliation: Not recordedWork: AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
- Duc LeAffiliation: MetaWork: MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
- Jason SandersAffiliation: Not recordedWork: Background audio identification for query disambiguation
- Jiaming KongAffiliation: Not recordedWork: Simultaneous Speech-to-Text Translation Web Application for Estonian
- Simon TokumineAffiliation: Not recordedWork: NotebookLM update: Video Overview support in 80 languages and deeper Audio Overviews
- Thilo KoehlerAffiliation: Not recordedWork: Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
- Haojie WeiAffiliation: Not recordedWork: Qwen2-Audio Technical Report
- Sten SootlaAffiliation: MetaWork: AECMOS: A speech quality assessment metric for echo impairment
- Jitesh Vinod PunjabiAffiliation: Not recordedWork: Relaxed Context-Aware Machine Learning Middleware (RCAMM) for Android: A Step towards Sustainability
- Ozlem KalinliAffiliation: MetaWork: AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
- Han LuAffiliation: Not recordedWork: Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- Qiaochu (Frank) ZhangAffiliation: Not recordedWork: Scaling ASR Improves Zero and Few Shot Learning
- James QinAffiliation: Not recordedWork: AudioPaLM: A Large Language Model That Can Speak and Listen
- Michael L. SeltzerAffiliation: MetaWork: AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
- Ruize GaoAffiliation: Not recordedWork: Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- Xiaohuan ZhouAffiliation: ByteDanceWork: Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Shibo WangAffiliation: Not recordedWork: Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Ajay KannanAffiliation: Google DeepMindWork: Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Jan-Thorsten PeterAffiliation: GoogleWork: Sisyphus, a Workflow Manager Designed for Machine Translation and Automatic Speech Recognition
- Jason RiesaAffiliation: Google ResearchWork: Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Lukáš ŽilkaAffiliation: Not recordedWork: AudioPaLM: A Large Language Model That Can Speak and Listen
- Tomer BitonAffiliation: Not recordedWork: AI Invoice OCR and Structured Data Extraction Pipeline
- Anmol GulatiAffiliation: Not recordedWork: Conformer: Convolution-augmented Transformer for Speech Recognition
- Igor TufanovAffiliation: Not recordedWork: Joint speech and text machine translation for up to 100 languages
- Jay MahadeokarAffiliation: MetaWork: Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
- Martin BäumlAffiliation: Not recordedWork: Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Ke LiAffiliation: Not recordedWork: AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
- Nayan SinghalAffiliation: Not recordedWork: Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities
- Karén SimonyanAffiliation: Not recordedWork: Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Vilobh MeshramAffiliation: Not recordedWork: Gemini 3.1 Flash TTS: the next generation of expressive AI speech
- Dalia El BadawyAffiliation: Not recordedWork: AudioPaLM: A Large Language Model That Can Speak and Listen
- Kushal LakhotiaAffiliation: Not recordedWork: Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
- Saisuresh KrishnakumaranAffiliation: GoogleWork: Entity-aware Joint Translation of Query and Semantic Parse
- Ankur BapnaAffiliation: MetaWork: Advanced audio dialog and generation with Gemini 2.5
- Izhak ShafranAffiliation: GoogleWork: SLM: Bridge the thin gap between speech and text foundation models
- Karan GoelAffiliation: CartesiaWork: It's Raw! Audio Generation with State-Space Models
- Mikołaj BińkowskiAffiliation: Not recordedWork: High Fidelity Speech Synthesis with Adversarial Networks
- Aarush KattaAffiliation: LAIONWork: LAION-5B: An open large-scale dataset for training next generation image-text models
- Luyu WangAffiliation: Google DeepMindWork: Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Seb NouryAffiliation: Google DeepMindWork: Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Shirin BadiezadeganAffiliation: Google DeepMindWork: Reconstructing incomplete and unreliable speech spectrogram for robust automatic speech recognition
- Steven Chu-Hong HoiAffiliation: Not recordedWork: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Marvin RitterAffiliation: GoogleWork: Audio Set: An ontology and human-labeled dataset for audio events