https://arxiv.org/abs/2408.16532
jishengpeng
novateur
AI & ML interests
speech language model, discrete codec, text to speech
Recent Activity
authored a paper about 4 hours ago
Omni Interaction Agent Technical Report authored a paper about 4 hours ago
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for
Zero-Shot Speech Synthesis authored a paper about 4 hours ago
WavChat: A Survey of Spoken Dialogue Models