高品質音声合成のための韻律制御

Research Project

Project/Area Number	12132202
Research Category	Grant-in-Aid for Scientific Research on Priority Areas
Allocation Type	Single-year Grants
Review Section	Science and Engineering
Research Institution	The University of Tokyo
Principal Investigator	広瀬啓吉東京大学, 大学院・新領域創成科学研究科, 教授 (50111472)
Co-Investigator(Kenkyū-buntansha)	小林隆夫東京工業大学, 大学院・総合理工学研究科, 教授 (70153616) 石塚満東京大学, 大学院・情報理工学系研究科, 教授 (50114369) 西田豊明東京大学, 大学院・情報理工学系研究科, 教授 (70135531) 徳田恵一名古屋工業大学, 工学部, 助教授 (20217483) WARD Nigel 東京大学, 大学院・情報理工学系研究科, 助教授 (00242008)
Project Period (FY)	2000 – 2003
Project Status	Completed (Fiscal Year 2003)
Budget Amount *help	¥56,800,000 (Direct Cost: ¥56,800,000) Fiscal Year 2003: ¥18,700,000 (Direct Cost: ¥18,700,000) Fiscal Year 2002: ¥19,400,000 (Direct Cost: ¥19,400,000) Fiscal Year 2001: ¥18,700,000 (Direct Cost: ¥18,700,000)
Keywords	音声合成 / 生成過程モデル / 回帰木 / 音声対話システム / 感情音声 / 発話速度 / HMM音声合成 / 平均声モデル / 固有音声 / 統計的基本周波数パターン生成 / 談話情報 / 話者適応 / 多空間確率分布HMM / パラ・非言語情報 / 統計的F0パターン生成 / 対話調音声 / Fillerの韻律 / 基本周波数抽出
Research Abstract	種々の調子の音声を従来になく人間らしい抑揚で合成する技術を確立した上でユーザフレンドリな応答音声生成システムを構築することを目的として研究を進め、以下の成果を達成した。 1.テキストを入力として、文節境界、(FOパターン生成過程モデルの)フレーズ・アクセント指令、音素長を回帰木により推定する統合的な手法を開発した。学習用コーパスのモデルの指令は、自動的に抽出しているが、フレーズ指令に制約をかけることにより抽出精度を向上させ、より良好な指令の推定を達成した。HMM音声合成に組み込み、朗読音声の他、感情音声について実験を行い、怒りの表出が適正に行われていることを確認した。 2.昨年度開発した、エージェント音声対話システムで応答文の概念から音声合成を一貫して行う手法において、その韻律の面からの品質向上を行い、その効果を聴取実験により確認した。 3.プレゼンテーションを想定し、書き言葉で表記された文内容から、ジェスチャータグ付き話し言葉の文を自動的に生成する手法を開発した。生成された文から音声合成を行うための韻律制御の手法を検討した。 4.感情表現機能付きマルチモーダルプレゼンテーション記述言語MPMLの開発を進め、出力される感情音声の観点から評価を行った。また、試用実験により、感情音声がユーザに与える効果を調べた。 5.適応技術により任意の話者・調子の音声を生成するためのHMM音声合成用平均声モデルの品質改良を行った。学習データ量が限られている場合への対処手法として、文脈クラスタと話者適応学習を導入することを行った。また、モーフィングにより、種々の話者・調子の音声の合成が可能なことを示した。 6.スペクトル・F0・継続長を統一的に扱うHMM音声合成において、音素HMMの適応により感情音声を合成することを行った。聴取実験により、意図した感情が有効に伝達されることを確認した。

Report

(4 results)

Research Products
(59 results)

All Other

All Publications (59 results)

[Publications] Atsuhiro Sakurai: "Data-driven generation of F0 contours using a superpositional model"Speech Communication. 40・4. 535-549 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keikichi Hirose: "Corpus-based synthesis of F0 contours for emotional speech using the generation process model"Proceedings 15th International Congress of Phonetic Sciences. 3. 2945-2948 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keikichi Hirose: "Use of linguistic information for automatic extraction of F0 contour generation process model parameters"Proceedings 8th European Conference on Speech Communication and Technology. 1. 141-144 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keikichi Hirose: "Corpus-based synthesis of fundamental frequency contours of Japanese using automatically-generated prosodic corpus and generation process model"Proccedings 8th European Conference on Speech Communication and Technology. 1. 333-336 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keikichi Hirose: "Speech generation from concept for realizing conversation with an agent in a virtual room"Proceedings 8th European Conference on Speech Communication and Technology. 3. 1693-1696 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keikichi Hirose: "Speech prosody in spoken language processing(invited)"Proccedings International Conference on Computer and Information Technology. 1. 20-27 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keikichi Hirose: "Emotional speech synthesis with corpus-based generation of F_0 contours using generation process model"Proceedings of International Conference on Speech Prosody. 417-420 (2004)
- Related Report
  2003 Annual Research Report
[Publications] Shuichi Narusawa: "Evaluation of an improved method for automatic extraction of model parameters from fundamental frequency contours of speech"Proceedings of International Conference on Speech Prosody. 443-446 (2004)
- Related Report
  2003 Annual Research Report
[Publications] Qing Li: "Highlighting multimodal syhchronization for embodied conversational agent"Proceedings of the 2nd International Conference on Information Technology for Application(ICITA 2004). 17-20 (2004)
- Related Report
  2003 Annual Research Report
[Publications] Helmut Prendinger: "Designing and evaluating animated agents as social actors"IEICE Transactions on Information and Systems. E86-D・8. 1378-1385 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Zhenglu Yang: "A two-model framework for multimodal presentation with life-like characters in flash medium"Proc.of 7th IASTED Int'l Conf.On Software Engineering and Applications(SEA 2003). 769-774 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Junichi Yamagishi: "A training method of average voice model for HMM-based speech synthesis"IEICE Trans.Fundamentals of Electronics, Communications and Computer Sciences. E86-A・8. 1956-1963 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Junich Yamagishi: "Modeling of various speaking styles and emotions for HMM-based speech synthesis"Proceedings 8th European Conference on Speech Communication and Technology. 3. 2461-2464 (2003)
- Related Report
  2003 Annual Research Report
[Publications] 都築亮介: "HMM音声合成における感情表現のモデル化"電子情報通信学会技術研究報告. 103・206. 25-30 (2003)
- Related Report
  2003 Annual Research Report
[Publications] Keiichi Tokuda: "Text-to-Speech Synthesis : New Paradigms and Advances(Edited by S.Narayanan, A.Alwan)"Prentice Hall. 23 (2004)
- Related Report
  2003 Annual Research Report
[Publications] Shinya Kiriyama: "Development and evaluation of a spoken dialogue system for academic document retrieval with a focus on reply generation"Systems and Computers in Japan. 33・4. 25-39 (2002)
- Related Report
  2002 Annual Research Report
[Publications] 成澤修一: "音声の基本周波数パターン生成過程モデルのパラメータ自動抽出法"情報処理学会論文誌. 43・7. 2155-2168 (2002)
- Related Report
  2002 Annual Research Report
[Publications] Nobuaki Minematsu: "Automatic estimation of accentual attribute values of words for accent sandhi rules of Japanese text-to-speech conversion"IEICE Trans. Information and Systems. E86-D・1. 550-557 (2003)
- Related Report
  2002 Annual Research Report
[Publications] Atsuhiro Sakurai: "Data-driven generation of F0 contours using a superpositional model"Speech Communication. (発表予定). (2003)
- Related Report
  2002 Annual Research Report
[Publications] Keikichi Hirose: "Improved corpus-based synthesis of fundamental frequency contours using generation process model"Proc. International Conference on Spoken Language Processing. 2085-2088 (2002)
- Related Report
  2002 Annual Research Report
[Publications] 多胡順司: "エージェント対話システムにおける音声応答生成手法"日本音響学会平成15年度春季研究発表会講演論文集. 1(発表予定). (2003)
- Related Report
  2002 Annual Research Report
[Publications] Keikichi Hirose: "Corpus-based synthesis of F0 contours for emotional speech using the generation process model"Proceedings 15th International Congress of Phonetic Sciences. (発表予定). (2003)
- Related Report
  2002 Annual Research Report
[Publications] 西田悠介: "料理教示発話の構造解析"言語処理学会第9回年次大会論文集. (発表予定). (2003)
- Related Report
  2002 Annual Research Report
[Publications] Nigel Ward: "Automatic user-adaptive speaking rate selection for information delivery"Proc. International Conference on Spoken Language Processing. 1. 549-552 (2002)
- Related Report
  2002 Annual Research Report
[Publications] Masafumi Okamoto: "Quantitative estimation of the meanings of the phonetic components of back-channels"Proc. 35th Spoken Language Understanding and Discourse Workshop. 47-52 (2002)
- Related Report
  2002 Annual Research Report
[Publications] 田村正統: "HMMに基づく音声合成におけるピッチ・スペクトルの話者適応"電子情報通信学会論文誌. J85-D-II・4. 545-553 (2002)
- Related Report
  2002 Annual Research Report
[Publications] Junichi Yamagishi: "A context clustering technique for average voice models"IEICE Trans. on Information and Systems. E86-D・3. 534-542 (2003)
- Related Report
  2002 Annual Research Report
[Publications] Keiichi Tokuda: "An HMM-based speech synthesis system applied to English"Proc. IEEE Speech Synthesis Workshop. (CD-ROM). (2002)
- Related Report
  2002 Annual Research Report
[Publications] Kengo Shichiri: "Eigenvoices for HMM-based speech synthesis"Proc. International Conference on Spoken Language Processing. 2. 1269-1272 (2002)
- Related Report
  2002 Annual Research Report
[Publications] 広瀬啓吉: "Temporal rate change of dialogue speech in prosodic units as compared to read speech"Speech Communication. 36・1-2. 97-111 (2002)
- Related Report
  2001 Annual Research Report
[Publications] 桐山伸也: "Development and evaluation of a spoken dialogue system for academic document retrieval with a focus on reply generation"Systems and Computers in Japan. (掲載予定). (2002)
- Related Report
  2001 Annual Research Report
[Publications] 広瀬啓吉: "Corpus-based synthesis of fundamental frequency contours based on a generation process model"Proc. European Conference on Speech Communication and Technology. 3. 2255-2258 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 桐山伸也: "Control of prosodic focuses for reply speech generation in a spoken dialogue system of information retrieval on academic documents"Proc. Speech Prosody 2002. (発表予定). (2002)
- Related Report
  2001 Annual Research Report
[Publications] 広瀬啓吉: "Data-driven synthesis of fundamental frequency contours for TTS systems based on a generation process model"Proc. Speech Prosody 2002. (発表予定). (2002)
- Related Report
  2001 Annual Research Report
[Publications] 成澤修一: "A method for automatic extraction of model parameters from fundamental frequency contours of speech"Proc. IEEE International Conference on Acoustics, Speech, & Signal Processing. (発表予定). (2002)
- Related Report
  2001 Annual Research Report
[Publications] 西田豊明: "知の創造と学習のための会話型コンテンツ"『情報技術と経済文化』,NTT出版(今井賢一編). (印刷中). (2002)
- Related Report
  2001 Annual Research Report
[Publications] 西田豊明: "Social intelligence design for knowledge creating communities"Proc. International Conference on Intelligent Agent Technology. 23-26 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 塚原渉: "Responding to subtle, fleeting changes in the user's internal state"CHI Letters. 3・1. 77-84 (2001)
- Related Report
  2001 Annual Research Report
[Publications] WARD Nigel: "Conversational grunts and real-time interaction(招待)"Proc. International Conference on Speech Processing. 1. 53-58 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 田村正統: "Text-to-speech synthesis with arbitrary speaker's voice from average voice"Proc. European Conference on Speech Communication and Technology. 3. 345-348 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 田村正統: "HMM音声合成におけるMLLRを用いたピッチ・スペクトルの話者適応"電子情報通信学会技術研究報告(音声研究会). 101・86. 15-20 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 全炳河: "有声/無声境界の動的特徴量を考慮したピッチのモデル化"電子情報通信学会技術研究報告(音声研究会). 101・325. 53-58 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 石松喜信: "HMM音声合成におけるガンマ分布状態継続長モデルの検討"電子情報通信学会技術研究報告(音声研究会). 101・352. 57-62 (2001)
- Related Report
  2001 Annual Research Report
[Publications] 広瀬啓吉: "Analytical and perceptual study on the role of acoustic features in realizing emotional speech"Proc.International Conf.on Spoken Language Processing. 2. 369-372 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 桜井淳宏: "Data-driven intonation modeling using a neural network and a command response model"Proc.International Conf.on Spoken Language Processing. 3. 223-226 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 桜井淳宏: "Modeling and generation of accentual phrase F0 contours based on discrete HMMs synchronized at mora-unit transitions"Proc.International Conf.on Spoken Language Processing. 3. 259-262 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 桐山伸也: "応答生成に着目した学術文献音声対話システムの構築とその評価"電子情報通信学会論文誌. J83-D-II・11. 2318-2329 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 桐山伸也: "Development and evaluation of a spoken dialogue system for academic document retrieval with a focus on reply generation"Proc.8-th Australian International Conference on Speech Science and Technology. 32-37 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 広瀬啓吉: "Temporal rate change of dialogue speech in prosodic units as compared to read speech"Speech Communication. (発表予定). (2001)
- Related Report
  2000 Annual Research Report
[Publications] 桜井淳宏: "Generation of FO contours using model-constrained data-driven method"Proc.IEEE International Conf.on Acoustics, Speech,& Signal Processing. (発表予定). (2001)
- Related Report
  2000 Annual Research Report
[Publications] 西田豊明: "Towards dynamic knowledge interaction (Keynote Paper)"Proc.4th International Conf.on Knowledge-based Intelligent Engineering Systems & Allied Technologies. 1-12 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 久保田秀和: "EgoChat Agent : A talking virtualized member for supporting community knowledge creation"Proc.AAAI Fall Symposium "Society Intelligent Agents-The Human in the Loop". (発表予定). (2001)
- Related Report
  2000 Annual Research Report
[Publications] WARD,Nigel: "Prosodic features which cue back-channel feedback in English and Japanese"Journal of Pragmatics. 32. 1177-1207 (2000)
- Related Report
  2000 Annual Research Report
[Publications] WARD,Nigel: "The challenge of non-lexical speech sounds"Proc.International Conf.on Spoken Language Processing. 2. 571-574 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 益子貴史: "多空間確率分布HMMによるピッチパターン生成"電子情報通信学会論文誌. J83-D-II・7. 1600-1609 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 田中智宏: "瞬時周波数振幅スペクトルに基づくピッチ抽出法の検討"電子情報通信学会技術報告(音声研究会). (発表予定). (2001)
- Related Report
  2000 Annual Research Report
[Publications] 徳田恵一: "多空間上の確率分布に基づいたHMM"電子情報通信学会論文誌. J83-D-II・7. 1579-1589 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 吉村貴克: "HMMに基づく音声合成におけるスペクトル・ピッチ・継続長の同時モデル化"電子情報通信学会論文誌. J83-D-II・7. 2099-2107 (2000)
- Related Report
  2000 Annual Research Report
[Publications] 徳田恵一: "Speech parameter generation algorithms for HMM-based speech synthesis"Proc.IEEE International Conf.on Acoustics, Speech, & Signal Processing. 3. 1315-1318 (2000)
- Related Report
  2000 Annual Research Report

高品質音声合成のための韻律制御

Principal Investigator

広瀬 啓吉 東京大学, 大学院・新領域創成科学研究科, 教授 (50111472)

¥56,800,000 (Direct Cost: ¥56,800,000)

Report

Research Products

[Publications] Atsuhiro Sakurai: "Data-driven generation of F0 contours using a superpositional model"Speech Communication. 40・4. 535-549 (2003)

Related Report

[Publications] Keikichi Hirose: "Corpus-based synthesis of F0 contours for emotional speech using the generation process model"Proceedings 15th International Congress of Phonetic Sciences. 3. 2945-2948 (2003)

Related Report

[Publications] Keikichi Hirose: "Use of linguistic information for automatic extraction of F0 contour generation process model parameters"Proceedings 8th European Conference on Speech Communication and Technology. 1. 141-144 (2003)

Related Report

[Publications] Keikichi Hirose: "Corpus-based synthesis of fundamental frequency contours of Japanese using automatically-generated prosodic corpus and generation process model"Proccedings 8th European Conference on Speech Communication and Technology. 1. 333-336 (2003)

Related Report

[Publications] Keikichi Hirose: "Speech generation from concept for realizing conversation with an agent in a virtual room"Proceedings 8th European Conference on Speech Communication and Technology. 3. 1693-1696 (2003)

Related Report

[Publications] Keikichi Hirose: "Speech prosody in spoken language processing(invited)"Proccedings International Conference on Computer and Information Technology. 1. 20-27 (2003)

Related Report

[Publications] Keikichi Hirose: "Emotional speech synthesis with corpus-based generation of F_0 contours using generation process model"Proceedings of International Conference on Speech Prosody. 417-420 (2004)

Related Report

[Publications] Shuichi Narusawa: "Evaluation of an improved method for automatic extraction of model parameters from fundamental frequency contours of speech"Proceedings of International Conference on Speech Prosody. 443-446 (2004)

Related Report

[Publications] Qing Li: "Highlighting multimodal syhchronization for embodied conversational agent"Proceedings of the 2nd International Conference on Information Technology for Application(ICITA 2004). 17-20 (2004)

Related Report

[Publications] Helmut Prendinger: "Designing and evaluating animated agents as social actors"IEICE Transactions on Information and Systems. E86-D・8. 1378-1385 (2003)

Related Report

[Publications] Zhenglu Yang: "A two-model framework for multimodal presentation with life-like characters in flash medium"Proc.of 7th IASTED Int'l Conf.On Software Engineering and Applications(SEA 2003). 769-774 (2003)

Related Report

[Publications] Junichi Yamagishi: "A training method of average voice model for HMM-based speech synthesis"IEICE Trans.Fundamentals of Electronics, Communications and Computer Sciences. E86-A・8. 1956-1963 (2003)

Related Report

[Publications] Junich Yamagishi: "Modeling of various speaking styles and emotions for HMM-based speech synthesis"Proceedings 8th European Conference on Speech Communication and Technology. 3. 2461-2464 (2003)

Related Report

[Publications] 都築亮介: "HMM音声合成における感情表現のモデル化"電子情報通信学会技術研究報告. 103・206. 25-30 (2003)

Related Report

[Publications] Keiichi Tokuda: "Text-to-Speech Synthesis : New Paradigms and Advances(Edited by S.Narayanan, A.Alwan)"Prentice Hall. 23 (2004)

Related Report

[Publications] Shinya Kiriyama: "Development and evaluation of a spoken dialogue system for academic document retrieval with a focus on reply generation"Systems and Computers in Japan. 33・4. 25-39 (2002)

Related Report

[Publications] 成澤修一: "音声の基本周波数パターン生成過程モデルのパラメータ自動抽出法"情報処理学会論文誌. 43・7. 2155-2168 (2002)

Related Report

[Publications] Nobuaki Minematsu: "Automatic estimation of accentual attribute values of words for accent sandhi rules of Japanese text-to-speech conversion"IEICE Trans. Information and Systems. E86-D・1. 550-557 (2003)

Related Report

[Publications] Atsuhiro Sakurai: "Data-driven generation of F0 contours using a superpositional model"Speech Communication. (発表予定). (2003)

Related Report

[Publications] Keikichi Hirose: "Improved corpus-based synthesis of fundamental frequency contours using generation process model"Proc. International Conference on Spoken Language Processing. 2085-2088 (2002)

Related Report

[Publications] 多胡 順司: "エージェント対話システムにおける音声応答生成手法"日本音響学会平成15年度春季研究発表会講演論文集. 1(発表予定). (2003)

Related Report

[Publications] Keikichi Hirose: "Corpus-based synthesis of F0 contours for emotional speech using the generation process model"Proceedings 15th International Congress of Phonetic Sciences. (発表予定). (2003)

Related Report

[Publications] 西田悠介: "料理教示発話の構造解析"言語処理学会第9回年次大会論文集. (発表予定). (2003)

Related Report

[Publications] Nigel Ward: "Automatic user-adaptive speaking rate selection for information delivery"Proc. International Conference on Spoken Language Processing. 1. 549-552 (2002)

Related Report

[Publications] Masafumi Okamoto: "Quantitative estimation of the meanings of the phonetic components of back-channels"Proc. 35th Spoken Language Understanding and Discourse Workshop. 47-52 (2002)

Related Report

[Publications] 田村正統: "HMMに基づく音声合成におけるピッチ・スペクトルの話者適応"電子情報通信学会論文誌. J85-D-II・4. 545-553 (2002)

Related Report

[Publications] Junichi Yamagishi: "A context clustering technique for average voice models"IEICE Trans. on Information and Systems. E86-D・3. 534-542 (2003)

Related Report

[Publications] Keiichi Tokuda: "An HMM-based speech synthesis system applied to English"Proc. IEEE Speech Synthesis Workshop. (CD-ROM). (2002)

Related Report

[Publications] Kengo Shichiri: "Eigenvoices for HMM-based speech synthesis"Proc. International Conference on Spoken Language Processing. 2. 1269-1272 (2002)

Related Report

[Publications] 広瀬啓吉: "Temporal rate change of dialogue speech in prosodic units as compared to read speech"Speech Communication. 36・1-2. 97-111 (2002)

Related Report

[Publications] 桐山伸也: "Development and evaluation of a spoken dialogue system for academic document retrieval with a focus on reply generation"Systems and Computers in Japan. (掲載予定). (2002)

Related Report

[Publications] 広瀬啓吉: "Corpus-based synthesis of fundamental frequency contours based on a generation process model"Proc. European Conference on Speech Communication and Technology. 3. 2255-2258 (2001)

Related Report

[Publications] 桐山伸也: "Control of prosodic focuses for reply speech generation in a spoken dialogue system of information retrieval on academic documents"Proc. Speech Prosody 2002. (発表予定). (2002)

Related Report

[Publications] 広瀬啓吉: "Data-driven synthesis of fundamental frequency contours for TTS systems based on a generation process model"Proc. Speech Prosody 2002. (発表予定). (2002)

Related Report

[Publications] 成澤修一: "A method for automatic extraction of model parameters from fundamental frequency contours of speech"Proc. IEEE International Conference on Acoustics, Speech, & Signal Processing. (発表予定). (2002)

Related Report

[Publications] 西田豊明: "知の創造と学習のための会話型コンテンツ"『情報技術と経済文化』,NTT出版(今井賢一編). (印刷中). (2002)

Related Report

[Publications] 西田豊明: "Social intelligence design for knowledge creating communities"Proc. International Conference on Intelligent Agent Technology. 23-26 (2001)

Related Report

広瀬啓吉東京大学, 大学院・新領域創成科学研究科, 教授 (50111472)

[Publications] 多胡順司: "エージェント対話システムにおける音声応答生成手法"日本音響学会平成15年度春季研究発表会講演論文集. 1(発表予定). (2003)