This forum presents research on low-bitrate neural speech coding. We introduce FMelCodec, an ultra-low-rate neural codec built on an encode-refine-reconstruct framework for high-quality reconstruction at 250 bps. We also discuss VoCodec for adaptive streaming quantization and AKDQ for efficient residual vector quantization.
This forum evaluates paralinguistic capabilities for natural human-computer interaction. We present MINT-Bench for multilingual instruction-following assessment of TTS models, and NVV-SuperBench with 45 categories of Chinese-English non-verbal vocalizations. We analyze core challenges in paralinguistic control.
This forum presents non-invasive intelligent detection of cognitive impairment from patients' natural conversational speech. We explore pathological speech feature extraction and cross-modal fusion modeling between speech and text, establishing an end-to-end diagnostic framework for elderly cognitive health.
This forum explores speech enhancement inspired by human auditory perception mechanisms. We cover auditory masking-based endpoint detection, a dual-stream noise-speech perception framework, and a harmonic-compensated auditory network combining cochlear frequency selectivity with pitch perception.
This forum introduces the "Wanjuan · Silk Road" multimodal low-resource language corpus for the Belt and Road Initiative. We discuss innovative construction of high-quality audio-visual and text corpora through expert-intelligent processing collaboration, and share enterprise globalization application experience.
This forum investigates music's physical properties and deep neural dynamics. Using intracranial EEG, we examine music's effects on dopamine reward circuits, and build a surrogate brain architecture enabling real-time deep reward state inference from scalp EEG for music-based emotional disorder therapy.
This forum addresses dysarthric and cognitively-impaired speech recognition challenges. We discuss transferring pretrained speech foundation models to disordered speech through model fusion, discrete token representations, semi-supervised learning, and adaptive modeling techniques.
This forum presents Xiaohongshu dots team's speech synthesis advances, including the open-source fully continuous autoregressive TTS model dots.tts, unified speech representation HoliTok, and fine-grained speech editing model dots.tts.edit supporting text-emotion-prosody-pause editing.
This forum showcases breakthroughs in audio reasoning and emotional speech generation, including Audio-DeepThinker with progressive reasoning, PhyAVBench for physical commonsense evaluation, EmoSteer-TTS for fine-grained emotion control, and AudioGenie multi-agent collaboration framework.
This forum focuses on low-resource Tibetan speech synthesis for the Lhasa dialect. We explore linguistic knowledge mining and end-to-end acoustic modeling across speech structure analysis, text feature analysis, prosody prediction, and acoustic modeling dimensions.
NCMMSC 2026 Young Researchers Forum invites proposals from young scholars worldwide. We welcome presentations on speech and language intelligence, multimodal interaction, affective computing, audio and music processing, and other frontier research directions.