Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
-
Updated
Jul 24, 2026 - C
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities.
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
Code for DeSTA2.5-Audio, general-purpose LALM
Outlining and demonstrating how language models are able to understand image, video, and text content.
[Interspeech 2026] Official Implementation of "ALARM: Audio–Language Alignment for Reasoning Models"
Repository of paper "Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against LALMs" (ICML'26)
[EMNLP 2025] Official code repository of paper titled "TrojanWave: Exploiting Prompt Learning for Stealthy Backdoor Attacks on Large Audio-Language Models" accepted in EMNLP 2025 conference.
Source code of our paper "When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models", ICASSP 2026
Core shared libraries for multimodal Kani extensions.
Offline audio QA, transcription, translation, emotion, temporal grounding and timed summaries on Apple Silicon — native GigaChat Audio MLX.
Bridge audio encoders to LLMs for audio captioning and sound-event understanding — pure-PyTorch, offline-by-default
Pointer-based structural recall for full-song generation. Choruses that return instead of being re-sung — asymmetric section reuse over hierarchical audio tokens.
Neural audio codec and tokenizer for audio language models — SEANet encoder with residual vector quantization in PyTorch
音频语言模型框架:把音频编码器 / 离散音频 token 接入 LLM,实现音频理解、描述、问答与检索(纯 numpy 核心,可选 torch)
Research preview: timing-aware turn-taking for spoken language models with external Time Adapters and action-token control.
Pytorch implementation of a Moshi-inspired audio LM based on multiple backbones, inlcuding Qwen and a transformer decoder.
低比特率神经音频编解码器 + 残差矢量量化:把连续音频转成离散 token,服务音频语言模型
Add a description, image, and links to the audio-language-model topic page so that developers can more easily learn about it.
To associate your repository with the audio-language-model topic, visit your repo's landing page and select "manage topics."