Young Researchers Forum

NCMMSC 2026 Young Researchers Forum
Forum 1

Research on Low-Bitrate Neural Speech Codec Methods

This forum presents research on low-bitrate neural speech coding. We introduce FMelCodec, an ultra-low-rate neural codec built on an encode-refine-reconstruct framework for high-quality reconstruction at 250 bps. We also discuss VoCodec for adaptive streaming quantization and AKDQ for efficient residual vector quantization.

Yang Ai
Yang Ai
National Engineering Research Center of Speech and Language Information Processing, USTC · Associate Researcher
Forum 2

Evaluating Paralinguistic Capabilities of Speech for Natural Human-Computer Interaction: From Instruction Following to Non-verbal Vocalization

This forum evaluates paralinguistic capabilities for natural human-computer interaction. We present MINT-Bench for multilingual instruction-following assessment of TTS models, and NVV-SuperBench with 45 categories of Chinese-English non-verbal vocalizations. We analyze core challenges in paralinguistic control.

Liumeng Xue
Liumeng Xue
School of Intelligence Science and Technology, Nanjing University · Tenure-track Assistant Professor
Forum 3

Intelligent Diagnosis of Cognitive Health in the Elderly Based on Doctor-Patient Dialogue

This forum presents non-invasive intelligent detection of cognitive impairment from patients' natural conversational speech. We explore pathological speech feature extraction and cross-modal fusion modeling between speech and text, establishing an end-to-end diagnostic framework for elderly cognitive health.

Yilin Pan
Yilin Pan
School of Artificial Intelligence, Dalian Maritime University · Lecturer, Master Supervisor
Forum 4

Speech Enhancement Methods Inspired by Auditory Perception Mechanisms

This forum explores speech enhancement inspired by human auditory perception mechanisms. We cover auditory masking-based endpoint detection, a dual-stream noise-speech perception framework, and a harmonic-compensated auditory network combining cochlear frequency selectivity with pitch perception.

Nan Li
Nan Li
School of Computer and Information Engineering, Tianjin Normal University · Lecturer
Forum 5

Innovation and Practice of the "Wanjuan · Silk Road" Multimodal Low-resource Language Corpus

This forum introduces the "Wanjuan · Silk Road" multimodal low-resource language corpus for the Belt and Road Initiative. We discuss innovative construction of high-quality audio-visual and text corpora through expert-intelligent processing collaboration, and share enterprise globalization application experience.

Xiaoman Wang
Xiaoman Wang
Shanghai Artificial Intelligence Laboratory · Senior Engineer
Forum 6

Mechanistic Study on the Quantitative Characterization of Neural Entrainment Effects of Musical Physical Properties on the Dopamine System

This forum investigates music's physical properties and deep neural dynamics. Using intracranial EEG, we examine music's effects on dopamine reward circuits, and build a surrogate brain architecture enabling real-time deep reward state inference from scalp EEG for music-based emotional disorder therapy.

Jing Xia
Jing Xia
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences · Assistant Researcher
Forum 7

Dysarthric and Cognitive-impaired Speech Recognition with Speech Foundation Models

This forum addresses dysarthric and cognitively-impaired speech recognition challenges. We discuss transferring pretrained speech foundation models to disordered speech through model fusion, discrete token representations, semi-supervised learning, and adaptive modeling techniques.

Xurong Xie
Xurong Xie
Institute of Software, Chinese Academy of Sciences · Associate Researcher
Forum 8

Towards More Stable Fully Continuous Autoregressive Speech Synthesis

This forum presents Xiaohongshu dots team's speech synthesis advances, including the open-source fully continuous autoregressive TTS model dots.tts, unified speech representation HoliTok, and fine-grained speech editing model dots.tts.edit supporting text-emotion-prosody-pause editing.

Da Zheng
Da Zheng
Xiaohongshu (RED) dots Speech Team
Hankun Wang
Hankun Wang
Shanghai Jiao Tong University X-LANCE · Ph.D. Student
Bohan Li
Bohan Li
Shanghai Jiao Tong University · Ph.D. Student
Forum 9

Advances in Audio Reasoning and Anthropomorphic Emotional Speech Generation

This forum showcases breakthroughs in audio reasoning and emotional speech generation, including Audio-DeepThinker with progressive reasoning, PhyAVBench for physical commonsense evaluation, EmoSteer-TTS for fine-grained emotion control, and AudioGenie multi-agent collaboration framework.

Li Liu
Li Liu
Hong Kong University of Science and Technology (Guangzhou) · Associate Researcher, Ph.D. Supervisor
Forum 10

Research on Low-resource Tibetan Speech Synthesis Optimization by Integrating Linguistic Prior Knowledge

This forum focuses on low-resource Tibetan speech synthesis for the Lhasa dialect. We explore linguistic knowledge mining and end-to-end acoustic modeling across speech structure analysis, text feature analysis, prosody prediction, and acoustic modeling dimensions.

Labadunzhu
Labadunzhu
School of Information Science and Technology, Tibet University · Associate Professor, Ph.D. Supervisor
Call

NCMMSC 2026 Young Researchers Forum Call for Proposals

NCMMSC 2026 Young Researchers Forum invites proposals from young scholars worldwide. We welcome presentations on speech and language intelligence, multimodal interaction, affective computing, audio and music processing, and other frontier research directions.

Zhijian Ou
Zhijian Ou
Tsinghua University
Junfeng Li
Junfeng Li
Institute of Acoustics, CAS
Nan Yan
Nan Yan
Shenzhen Institute of Advanced Technology, CAS