Audio General Intelligence

Open models that listen, understand, and reason about all sound

GAMMA Lab at the University of Maryland builds AI systems that reason deeply about speech, music, and the acoustic world — developing open benchmarks, models, and methods that advance the science of auditory intelligence.

Audio General Intelligence — the capacity of AI agents to deeply understand and reason about all types of auditory input, including speech, environmental sounds, and music — is crucial for enabling AI to interact seamlessly and naturally with our world.

Despite this importance, audio intelligence has traditionally lagged behind advancements in vision and language processing. This gap arises from significant challenges: limited datasets, the complexity of audio signals, and a shortage of advanced neural architectures and effective training methodologies tailored specifically for audio. Recent breakthroughs in Large Language Models have begun to transform the landscape — offering promising pathways to enhance foundational audio tasks like ASR, cross-modal retrieval, and audio captioning, while giving emergence to new tasks like complex Audio Question Answering.

At GAMMA Lab, our mission is to accelerate progress toward Audio General Intelligence through open and accessible innovation. Our flagship Audio Flamingo series — spanning GAMA, AF2, AF3, and now Audio Flamingo Next — features specialized architectures, optimized audio encoders, and meticulously curated alignment datasets that excel in complex reasoning, long-form audio understanding, and hallucination robustness across speech, sound, and music.

We are actively expanding beyond audio into full omni-modal intelligence. MMOU and EgoAVU push reasoning over long, complex video. Through open-source models, rigorous benchmarks like MMAU and MMOU, and synthetic data frameworks such as Synthio, GAMMA Lab fosters transparency and collaboration — ensuring audio and multimodal intelligence remains inclusive, impactful, and accessible worldwide.

NVIDIANVIDIA
MetaMeta
GoogleGoogle
AdobeAdobe
DolbyDolby
Scale AIScale AI
SesameSesame
CentificCentific

Latest Papers

View all →
Preprint
SteerDuplex: Steerable Duplex Speech Dialogue Models
SpeechDialogue
Tyagi, Selvakumar et al.
EMNLP 2026
TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models
Audio LLMsTemporal Grounding
Kulkarni, Jayakumar et al.
EMNLP 2026
VIBE: Video Instruction-aligned Background Music Generation
GenerationMultimodal
Bhosale, Lokegaonkar et al.
Preprint
DuplexWorld: Can Voice Agents Help You Get Through the Day?
BenchmarksVoice Agents
Bhosale, Rajgarhia et al.
Preprint
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
BenchmarksUnified Audio
Bhosale, Rajgarhia et al.
Preprint
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
Audio-VisualAudio LLMs
Ghosh, Goel et al.
ACL 2026
FIGMA: Towards Fine-Grained Music Retrieval
Music RetrievalAudio-Text
Anand, Seth et al.
Sep 2026
SteerDuplex released — steerable full-duplex speech dialogue models that shift tone, persona, speaking rate, and voice style on instruction, with the SteerBench benchmark and two-stage reinforcement learning. Project page ↗
Sep 2026
🎉 Two papers accepted at EMNLP 2026.
Aug 2026
DuplexWorld released — a benchmark for speech-to-speech voice agents across six real-world domains, evaluating agentic, conversational, and speech-naturalness quality over 156 scenarios and 350+ hours of dialogue. Project page ↗
Jul 2026
TORUS released — the first self-coherence benchmark for unified audio models, testing whether a model's generation and understanding heads agree about the same audio. Project page ↗
Jul 2026
Audio-Visual Flamingo released — fully open audio-visual LLM for long, complex video with timestamp-grounded chain-of-thought reasoning. Project page ↗
Jun 2026
FIGMA accepted to ACL 2026 — fine-grained music retrieval with multi-view contrastive learning. Project page ↗
Apr 2026
Audio Flamingo Next released — next-generation open audio-language model supporting 30-minute audio, temporal reasoning, and 1M+ hours of training data. Project page ↗
Apr 2026
Apr 2026
Do Audio-Visual Large Language Models Really See and Hear? accepted to CVPR 2026 — mechanistic interpretability study revealing a strong vision bias in audio-visual LLMs.
Mar 2026
MMOU released — omni-modal benchmark for reasoning over long, complex real-world videos across visual, audio, and text signals.
Mar 2026
AHA: Audio Hallucination Attacks released — probing the reliability of large audio language models with 6,500 adversarial QA pairs.
Dec 2025
MultiVox accepted as Oral at EMNLP 2025 — multimodal voice assistant benchmark over audio-visual content.
Nov 2025
Audio Flamingo 3 selected as Spotlight at NeurIPS 2025. New state-of-the-art on MMAU and Air-Bench.
Oct 2025
EH-MAM accepted as Oral at EMNLP 2025 — easy-to-hard masked acoustic modeling for self-supervised pre-training.
Sep 2025
Three papers at NAACL 2025: ProSE (Oral), PAT (Oral), and Do Audio-Language Models Understand Linguistic Variations?
Jul 2025
Audio Flamingo 2 accepted to ICML 2025. Long-form audio understanding up to 5 minutes across speech, music, and environmental sound.
May 2025
MMAU accepted as Spotlight at ICLR 2025. Music Flamingo accepted at ICLR 2026.
Nov 2024
GAMA accepted as Oral at EMNLP 2024 — general-purpose audio understanding via instruction-tuned LLMs.
Oct 2025
DCASE 2025 Audio Question Answering Challenge — Workshop, Oct 30–31, Barcelona, Spain.
Jun 2025
JSALT 2025 Summer Workshop — June 9 – August 1, Brno, Czechia.
Apr 2025
SALMA Workshop @ ICASSP 2025 — Sound and Language Multimodal Analysis. April 6–11, Hyderabad, India.