语音交互+情感大模型:从判别式情感识别迈向生成式情感理解
Leveraging the rich vocabulary and multimodal perception capabilities of multimodal large models to transition from discriminative affective recognition to fine-grained, interpretable generative affective understanding, covering EMER, OV-MER, AffectGPT, EmoPrefer, and AffectGPT-RL.
空间多模态听觉感知:面向真实复杂场景的说话人定位与提取
Systematically exploring how to organically integrate visual spatial information with multi-microphone array audio signals, covering 3D sound source localization, multimodal target speaker extraction, and 3D spatial audio-visual reasoning.
操控生成:面向音频生成与编辑的条件扩散、桥、与引导方法
Introducing audio generation and editing methods based on diffusion models, covering fundamental theory, cross-modal consistency training and inference, bridge models, and efficient content processing such as audio super-resolution and content translation.
面向深度伪造语音的话者隐私保护:语音匿名化与水印
Focusing on voice anonymization and watermarking technologies, systematically introducing their principles, representative methods, development history, and current challenges to reduce deepfake generation risks.
鲁棒对话智能:复杂场景下人机语音交互感知与表达
Building a complete dialogue intelligence system of "robust perception - precise understanding - empathetic expression", covering multimodal perception and human-like conversational speech synthesis technologies.
语音增强前沿进展:从预测式估计迈向生成式建模
Systematically reviewing single-channel speech enhancement with a unified analysis of predictive and generative paradigms, covering problem formulation, speech representation, paradigm and architecture evolution, and emerging directions such as universal speech enhancement driven by speech foundation models and large language models.
教程报告提案征集
We cordially invite experts and scholars from home and abroad to submit tutorial proposals. The conference welcomes high-quality tutorials covering relevant research areas, with particular encouragement for proposals on emerging research directions that are closely related to the conference theme and possess cutting-edge, interdisciplinary, and applied value.