About Me

I am a Senior Research Scientist at Tencent, working on Text-to-Speech, Audio Generation, Speech LLMs, and MLLMs. I received my Ph.D. from Institute of Automation, Chinese Academy of Sciences in 2026, advised by Professor Jianhua Tao and Associate Professor Jiangyan Yi. Before that, I received my B.Eng. degree from Tsinghua University in 2020. I have also worked at Tencent AI Lab and StepFun. I have published some papers at the top international AI or Speech conferences such as ICML2026, ACL2025, ICASSP 2023/2024/2025, Interspeech 2025/2026 with Google Scholar · 966 citations.

🔥 News

📝 Selected Publications

arXiv 2026 Overview of the end-to-end discrete-token TTS training framework

End-to-End Training for Discrete Token LLM based TTS System

Changfeng Gao, Yong Ren, Jun Yuan, Ye Bai, Zhao You, Shidong Shang

Paper

  • Jointly optimizes the speech tokenizer, LLM, flow-matching decoder, and reward model to reduce cascade mismatch in discrete-token TTS.
Interspeech 2026 Oral Overview of the imperceptible text-based speech editing framework

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards

Yong Ren, Jiangyan Yi, Jianhua Tao, Zhengqi Wen, Tao Wang

Paper

  • Combines semantic-space editing, flow-matching reconstruction, and self-consistency rewards for seamless text-based speech editing.
ICML 2026 Overview of the MCLP evaluation and reward framework for role-play TTS

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

Yong Ren*, Jingbei Li*, Haiyang Sun, Yujie Chen, Cheng Yi, Yechang Huang, Hao Gu, Ye Bai, Xuerui Yang

Paper Code HF

  • Introduces mean continuation log-probability as an interpretable metric and reinforcement-learning reward for expressive role-play TTS.
ICASSP 2026 Oral Architecture and example of the OV-InstructTTS-TEP framework

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

Yong Ren, Jiangyan Yi, Jianhua Tao, Haiyang Sun, Zhengqi Wen, Hao Gu, Le Xu, Ye Bai

Paper Code HF

  • Uses reasoning over flexible natural-language instructions to infer emotional, acoustic, and paralinguistic attributes for expressive TTS.
Interspeech 2025 Overview of the SVAD training and reasoning framework

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Yong Ren, Chenxing Li, Le Xu, Hao Gu, Duzhen Zhang, Yujie Chen, Manjie Xu, Ruibo Fu, Shan Yang, Dong Yu

Paper Code

  • Introduces silent-video audio-description reasoning and improves downstream video-to-audio generation through chain-of-thought supervision.
ICASSP 2025 Overview of the STA-V2A semantic and temporal alignment framework

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Yong Ren, Chenxing Li, Manjie Xu, Wei Liang, Yu Gu, Rilin Chen, Dong Yu

Paper Code HF

  • Aligns local temporal and global semantic video features with text guidance for higher-quality, better-synchronized video-to-audio generation.
ICASSP 2024 Overview of the TiCodec architecture and time-invariant representation module

Fewer-Token Neural Speech Codec with Time-Invariant Codes

Yong Ren, Tao Wang, Jiangyan Yi, Le Xu, Jianhua Tao, Chuyuan Zhang, Junzuo Zhou

Paper Code HF

  • Separates time-invariant information from frame-level codes to improve zero-shot TTS while using fewer speech tokens.

🎖 Honors and Recognition

📖 Education and Experience

  • 2026.07–present · Senior Research Scientist, Tencent YuanBao.
  • 2026.01–2026.07 · Intern, Tencent YuanBao.
  • 2025.05-2026.01 · Intern, StepFun.
  • 2024.04-2025.03 · Intern, Tencent AI Lab.
  • 2020.09-2026.06 · Ph.D. in Institute of Automation, Chinese Academy of Sciences.
  • 2016.09-2020.06 · B.Eng. in Department of Automation. Tsinghua University.

🧑‍⚖️ Academic Services

Journal Reviewer
Transactions on Audio, Speech and Language Processing (TASLP)
Conference Reviewer
ICML 2026 NeurIPS 2026 AAAI 2027 ACM MM 2026 ICASSP ICME 2026 SLT 2026

💬 Invited Talks

2025.12
I gave an invited talk on InstructTTS, hosted by the Metaverse Technical Committee of the Chinese Association for Artificial Intelligence.
2025.04
I gave an invited talk on Audio Watermarking at the Xmart Student Forum.