语音交互+情感大模型:从判别式情感识别迈向生成式情感理解
Leveraging the rich vocabulary and multimodal perception capabilities of multimodal large models to transition from discriminative affective recognition to fine-grained, interpretable generative affective understanding, covering EMER, OV-MER, AffectGPT, EmoPrefer, and AffectGPT-RL.
空间多模态听觉感知:面向真实复杂场景的说话人定位与提取
Systematically exploring how to organically integrate visual spatial information with multi-microphone array audio signals, covering 3D sound source localization, multimodal target speaker extraction, and 3D spatial audio-visual reasoning.
操控生成:面向音频生成与编辑的条件扩散、桥、与引导方法
Introducing audio generation and editing methods based on diffusion models, covering fundamental theory, cross-modal consistency training and inference, bridge models, and efficient content processing such as audio super-resolution and content translation.
面向深度伪造语音的话者隐私保护:语音匿名化与水印
Focusing on voice anonymization and watermarking technologies, systematically introducing their principles, representative methods, development history, and current challenges to reduce deepfake generation risks.
鲁棒对话智能:复杂场景下人机语音交互感知与表达
Building a complete dialogue intelligence system of "robust perception - precise understanding - empathetic expression", covering multimodal perception and human-like conversational speech synthesis technologies.
教程报告提案征集
We cordially invite experts and scholars from home and abroad to submit tutorial proposals. The conference welcomes high-quality tutorials covering relevant research areas, with particular encouragement for proposals on emerging research directions that are closely related to the conference theme and possess cutting-edge, interdisciplinary, and applied value.