About Me
I am a Senior Research Scientist at Tencent, working on Text-to-Speech, Audio Generation, Speech LLMs, and MLLMs. I received my Ph.D. from Institute of Automation, Chinese Academy of Sciences in 2026, advised by Professor Jianhua Tao and Associate Professor Jiangyan Yi. Before that, I received my B.Eng. degree from Tsinghua University in 2020. I have also worked at Tencent AI Lab and StepFun. I have published some papers at the top international AI or Speech conferences such as ICML2026, ACL2025, ICASSP 2023/2024/2025, Interspeech 2025/2026 with Google Scholar · 966 citations.
🔥 News
- 2026.06Released End-to-End Training for Discrete Token LLM based TTS System.
- 2026.06Edit Content, Preserve Acoustics was accepted to Interspeech 2026 (Oral).
- 2026.05Evaluating and Rewarding LALMs for Expressive Role-Play TTS was accepted to ICML 2026.
- 2026.01OV-InstructTTS was accepted to ICASSP 2026 (Oral).
- 2025.05Hearing from Silence was accepted to Interspeech 2025.
- 2024.12STA-V2A was accepted to ICASSP 2025.
- 2023.12TiCodec was accepted to ICASSP 2024.
📝 Selected Publications
End-to-End Training for Discrete Token LLM based TTS System
Changfeng Gao, Yong Ren, Jun Yuan, Ye Bai, Zhao You, Shidong Shang
- Jointly optimizes the speech tokenizer, LLM, flow-matching decoder, and reward model to reduce cascade mismatch in discrete-token TTS.
Yong Ren, Jiangyan Yi, Jianhua Tao, Zhengqi Wen, Tao Wang
- Combines semantic-space editing, flow-matching reconstruction, and self-consistency rewards for seamless text-based speech editing.
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
Yong Ren*, Jingbei Li*, Haiyang Sun, Yujie Chen, Cheng Yi, Yechang Huang, Hao Gu, Ye Bai, Xuerui Yang
- Introduces mean continuation log-probability as an interpretable metric and reinforcement-learning reward for expressive role-play TTS.
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
Yong Ren, Jiangyan Yi, Jianhua Tao, Haiyang Sun, Zhengqi Wen, Hao Gu, Le Xu, Ye Bai
- Uses reasoning over flexible natural-language instructions to infer emotional, acoustic, and paralinguistic attributes for expressive TTS.
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren, Chenxing Li, Le Xu, Hao Gu, Duzhen Zhang, Yujie Chen, Manjie Xu, Ruibo Fu, Shan Yang, Dong Yu
- Introduces silent-video audio-description reasoning and improves downstream video-to-audio generation through chain-of-thought supervision.
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
Yong Ren, Chenxing Li, Manjie Xu, Wei Liang, Yu Gu, Rilin Chen, Dong Yu
- Aligns local temporal and global semantic video features with text guidance for higher-quality, better-synchronized video-to-audio generation.
Fewer-Token Neural Speech Codec with Time-Invariant Codes
Yong Ren, Tao Wang, Jiangyan Yi, Le Xu, Jianhua Tao, Chuyuan Zhang, Junzuo Zhou
- Separates time-invariant information from frame-level codes to improve zero-shot TTS while using fewer speech tokens.
🎖 Honors and Recognition
📖 Education and Experience
- 2026.07–present · Senior Research Scientist, Tencent YuanBao.
- 2026.01–2026.07 · Intern, Tencent YuanBao.
- 2025.05-2026.01 · Intern, StepFun.
- 2024.04-2025.03 · Intern, Tencent AI Lab.
- 2020.09-2026.06 · Ph.D. in Institute of Automation, Chinese Academy of Sciences.
- 2016.09-2020.06 · B.Eng. in Department of Automation. Tsinghua University.