{"id":2327,"date":"2025-06-25T16:35:33","date_gmt":"2025-06-25T07:35:33","guid":{"rendered":"https:\/\/aida.korea.ac.kr\/?page_id=2327"},"modified":"2025-06-25T16:46:15","modified_gmt":"2025-06-25T07:46:15","slug":"631-2","status":"publish","type":"page","link":"https:\/\/aida.korea.ac.kr\/?page_id=2327","title":{"rendered":""},"content":{"rendered":"\n<h1 class=\"wp-block-heading\">Deep Learning \u2013 Speech Synthesis<\/h1>\n\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong><strong>RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching<\/strong><\/strong><\/h2>\n\n\n\n<p><strong>Objective<\/strong><\/p>\n\n\n\n<p>Text-to-speech (TTS), also known as speech synthesis, aims to synthesize high-fidelity speech, given an input text. While previous Ordinary Differential Equations (ODE)-based TTS models, such as diffusion and flow matching, have demonstrated strong performance, they still suffer from slow inference due to the need for many generation steps. In this work, we propose RapFlow-TTS, which improves both inference efficiency and synthesis quality by leveraging consistency flow matching and enhanced training strategies.\n <\/p>\n\n\n\n<p><strong>Data<\/strong><\/p>\n\n\n\n<p>We use the LJSpeech [1] and VCTK [2] dataset which are single- and multi-speaker English corpus dataset. <\/p>\n\n\n\n<p class=\"has-small-font-size\">[1] K. Ito and L. Johnson, \u201cThe LJ speech dataset,\u201d https:\/\/keithito.com\/LJ-Speech-Dataset\/, 2017 <br>\n[2] J. Yamagishi, C. Veaux, and K. MacDonald, \u201cCSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),\u201d 2019\n\n\n<\/p>\n\n\n\n<p><strong>Related Work<\/strong><\/p>\n\n\n\n<p>Grad-TTS[3] is a diffusion-based approach that requires a large number of inference steps due to its complex ODE trajectories.<br>\nMatcha-TTS[4] mitigates this by leveraging flow matching to linearize the ODE paths, enabling high-quality speech synthesis with fewer inference steps. However, it still demands a considerable number of steps for generation.\n<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/RapFlow-tts1.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [3] V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, \u201cGrad-TTS: A diffusion probabilistic model for text-to-speech,\u201d in Proc. ICML, 2021, pp. 8599\u20138608<br>\n[4] S. Mehta, R. Tu, J. Beskow, E. Sz \u00b4 ekely, and G. E. Henter,  \u201cMatcha-TTS: A fast tts architecture with conditional flow matching,\u201d in Proc. ICASSP, 2024, pp. 11 341\u201311 345\n    <\/figcaption>\n<\/figure>\n\n\n<p><strong>Proposed Method<\/strong><\/p>\n\n\n\n<p>RapFlow-TTS leverages both flow matching and the concept of consistency to learn consistency along linearized trajectories. Unlike previous methods, this enables high-quality speech synthesis with significantly fewer inference steps. In addition, RapFlow-TTS incorporates several enhanced strategies\u2014such as time-delta scheduling, Huber loss, and adversarial learning\u2014to further improve synthesis quality.<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/RapFlow-tts2.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [5] H. J. Park, et al. &#8220;RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching \u201c, Interspeech 2025\n    <\/figcaption>\n<\/figure>\n\n\n<p>RapFlow-TTS achieves outstanding performance while requiring 5 to 10 times fewer inference steps compared to previous ODE-based models. Remarkably, it offers inference speed comparable to FastSpeech2, one of the fastest models available.<br>\n<br>\nIn summary, RapFlow-TTS synthesizes high-quality speech with fewer steps, enabled by consistency-based flow matching and improved training strategies, effectively overcoming the inference speed limitations of prior ODE-based TTS models.<br>\n<br>\nMoreover, visualization of 2-step inference demonstrates that our method can generate detailed frequency representations even with extremely few steps.<br>\n<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/RapFlow-tts3.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [5] H. J. Park, et al. &#8220;RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching \u201c, Interspeech 2025\n    <\/figcaption>\n<\/figure>\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong><strong>DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability<\/strong><\/strong><\/h2>\n\n\n\n<p><strong>Objective<\/strong><\/p>\n\n\n\n<p>Reference-based TTS is the task of synthesizing more expressive and natural speech by representing style characteristics from a reference speech sample. To enhance style representation and generalization, we propose DEX-TTS which is a diffusion-based TTS model. DEX-TTS differentiates styles into time-invariant and time-variant categories. By designing encoders and adapters tailored to each style characteristic, rich style representation and high generalization can be achieved.\n <\/p>\n\n\n\n<p><strong>Data<\/strong><\/p>\n\n\n\n<p>We use the multi-speaker English dataset, VCTK [1], and emotional dataset, ESD [2].<\/p>\n\n\n\n<p class=\"has-small-font-size\">[1] Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald. Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92). In University of Edinburgh. The Centre for Speech Technology Research (CSTR), 2019.<br>\n[2] Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li. Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 920\u2013924. IEEE, 2021.\n\n<\/p>\n\n\n\n<p><strong>Related Work<\/strong><\/p>\n\n\n\n<p>StyleTTS [3] adopted AdaIN for flexible style adaptation.<br>\nGenerSpeech [4] adopted a multi-level adapter for rich style adaptation.\n<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/Dex-tts1.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [3] Yinghao Aaron Li, Cong Han, and Nima Mesgarani. Styletts: A style-based generative model for natural and diverse text-to-speech synthesis. arXiv preprint arXiv:2205.15439, 2022. <br>\n        [4] Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao. Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech synthesis. arXiv preprint arXiv:2205.07211, 2022.\n    <\/figcaption>\n<\/figure>\n\n\n<p><strong>Proposed Method<\/strong><\/p>\n\n\n\n<p>DEX-TTS is a reference-based TTS with enhanced style representations. DEX-TTS consists of encoders and adapters to extract and represent reference styles categorized into time-invariant (T-IV) and time-variant (T-V) styles. By leveraging style modeling on time variability, DEX-TTS can represent enriched styles representations. In addition, DEX-TTS can enhance the performance by building the framework on the diffusion system.<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/Dex-tts2.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [5] Park, Hyun Joon, et al. \u201cDEX-TTS: Diffusion based Expressive Text-to-Speech with style modeling on time variability\u201d, Knowledge-Based Systems 2025\n    <\/figcaption>\n<\/figure>\n\n\n<p>DEX-TTS outperforms previous methods in both objective and subjective evaluations on the VCTK and ESD datasets. It consistently achieves high similarity scores, demonstrating the effectiveness of its style modeling in extracting and reflecting rich styles from reference speech. Additionally, DEX-TTS shows strong generalization in zero-shot scenarios while maintaining high speech quality. Unlike previous approaches, it achieves excellent performance using only mel-spectrograms without relying on pre-trained models.<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/Dex-tts3.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [5] Park, Hyun Joon, et al. \u201cDEX-TTS: Diffusion based Expressive Text-to-Speech with style modeling on time variability\u201d, Knowledge-Based Systems 2025\n    <\/figcaption>\n<\/figure>\n\n\n<p>We visualize the extracted T-IV and T-V styles using T-SNE and analyze their clustering behavior. The results show that our style representations effectively capture speaker and emotional information without explicit labels.<\/p>\n\n\n<figure class=\"wp-block-image aligncenter size-full\">\n    <img decoding=\"async\" src=\"https:\/\/aida.korea.ac.kr\/wp-content\/uploads\/2025\/06\/Dex-tts4.png\" alt=\"\" class=\"wp-image-1727\"\/>\n    <figcaption class=\"wp-element-caption\">\n        [5] Park, Hyun Joon, et al. \u201cDEX-TTS: Diffusion based Expressive Text-to-Speech with style modeling on time variability\u201d, Knowledge-Based Systems 2025\n    <\/figcaption>\n<\/figure>\n\n","protected":false},"excerpt":{"rendered":"<p>Deep Learning \u2013 Speech Synthesis RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching Objective Text-to-speech (TTS), also known as speech synthesis, aims to synthesize high-fidelity speech, given an input text. While previous Ordinary Differential Equations (ODE)-based TTS models, such as diffusion and flow matching, have demonstrated strong performance, they still suffer from slow &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/aida.korea.ac.kr\/?page_id=2327\" class=\"more-link\">Read more<span class=\"screen-reader-text\"> &#8220;&#8221;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-2327","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages\/2327","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2327"}],"version-history":[{"count":5,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages\/2327\/revisions"}],"predecessor-version":[{"id":2363,"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=\/wp\/v2\/pages\/2327\/revisions\/2363"}],"wp:attachment":[{"href":"https:\/\/aida.korea.ac.kr\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2327"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}