Articulatory synthesis

File:Modeling-Consonant-Vowel-Coarticulation-for-Articulatory-Speech-Synthesis-pone.0060603.s008.ogv

3D vocal tract model for Articulatory synthesis Based on Consonant-Vowel Coarticulation modeling, German sentence “Leo und Doreen mögen Bananen” was reproduced from a naturally spoken sentence in terms of the fundamental frequency and the phone durations.^[1]

Articulatory synthesis refers to computational techniques for synthesizing speech based on models of the human vocal tract and the articulation processes occurring there. The shape of the vocal tract can be controlled in a number of ways which usually involves modifying the position of the speech articulators, such as the tongue, jaw, and lips. Speech is created by digitally simulating the flow of air through the representation of the vocal tract.

Mechanical talking heads

There is a long history of attempts to build mechanical "talking heads." ^[2] Mechanical devices.Gerbert (d. 1003), Albertus Magnus (1198–1280) and Roger Bacon (1214–1294) are all said to have built speaking heads (Wheatstone 1837). However, historically confirmed speech synthesis begins with Wolfgang von Kempelen (1734–1804), who published an account of his research in 1791 (see also Dudley & Tarnoczy 1950).

Electrical vocal tract analogs

The first electrical vocal tract analogs were static, like those of Dunn (1950), Ken Stevens and colleagues (1953), Gunnar Fant (1960). Rosen (1958) built a dynamic vocal tract (DAVO), which Dennis (1963) later attempted to control by computer. Dennis et al. (1964), Hiki et al. (1968) and Baxter and Strong (1969) have also described hardware vocal-tract analogs. Kelly and Lochbaum (1962) made the first computer simulation; later digital computer simulations have been made, e.g. by Nakata and Mitsuoka (1965), Matsui (1968) and Paul Mermelstein (1971). Honda et al. (1968) have made an analog computer simulation.

Haskins and Maeda models

The first software articulatory synthesizer regularly used for laboratory experiments was developed at Haskins Laboratories in the mid-1970s by Philip Rubin, Tom Baer, and Paul Mermelstein. This synthesizer, known as ASY,^[3] was a computational model of speech production based on vocal tract models developed at Bell Laboratories in the 1960s and 1970s by Paul Mermelstein, Cecil Coker, and colleagues. Another popular model that has been frequently used is that of Shinji Maeda, which uses a factor-based approach to control tongue shape.

Modern models

Recent progress in speech production imaging, articulatory control modeling, and tongue biomechanics modeling has led to changes in the way articulatory synthesis is performed [1]. Examples include the Haskins CASY model (Configurable Articulatory Synthesis),^[4] designed by Philip Rubin, Mark Tiede [2], and Louis Goldstein [3], which matches midsagittal vocal tracts to actual magnetic resonance imaging (MRI) data, and uses MRI data to construct a 3D model of the vocal tract. A full 3D articulatory synthesis model has been described by Olov Engwall. A geometrically based 3D articulatory speech synthesizer has been developed by Peter Birkholz (VocalTractLab^[5]). The ArtiSynth project,^[6] headed by Sidney Fels [4] at the University of British Columbia, is a 3D biomechanical modeling toolkit for the human vocal tract and upper airway. Biomechanical modeling of articulators such as the tongue has been pioneered by a number of scientists, including Reiner Wilhelms-Tricarico [5], Yohan Payan [6] and Jean-Michel Gerard [7], Jianwu Dang and Kiyoshi Honda [8].

Commercial models

One of the few commercial articulatory speech synthesis systems is the NeXT-based system originally developed and marketed by Trillium Sound Research, a spin-off company of the University of Calgary, where much of the original research was conducted. Following the demise of the various incarnations of NeXT (started by Steve Jobs in the late 1980s and merged with Apple Computer in 1997), the Trillium software was published under a GNU General Public Licence, with work continuing as gnuspeech. The system, first marketed in 1994, provides full articulatory-based text-to-speech conversion using a waveguide or transmission-line analog of the human oral and nasal tracts controlled by Rene Carré's "distinctive region model".^[7]

Footnotes

↑ Lua error in package.lua at line 80: module 'strict' not found.
↑ Talking Heads
↑ ASY
↑ CASY
↑ VocalTractLab
↑ Artisynth
↑ Real-time articulatory speech-synthesis-by-rules

Bibliography

Baxter, Brent, and William J. Strong. (1969). WINDBAG—a vocal-tract analog speech synthesizer. Journal of the Acoustical Society of America, 45, 309(A).
Birkholz P, Jackel D, Kröger BJ (2007) Simulation of losses due to turbulence in the time-varying vocal system. IEEE Transactions on Audio, Speech, and Language Processing 15: 1218-1225
Birkholz P, Jackel D, Kröger BJ (2006) Construction and control of a three-dimensional vocal tract model. Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2006) (Toulouse, France) pp. 873–876
Coker. C. H. (1968). Speech synthesis with a parametric articulatory model. Proc. Speech. Symp., Kyoto, Japan, paper A-4.
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Dennis, Jack B. (1963). Computer control of an analog vocal tract. Journal of the Acoustical Society of America, 35, 1115(A).
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Engwall, O. (2003). Combining MRI, EMA & EPG measurements in a three-dimensional tongue model. Speech Communication, 41, 303-329.
Fant, C. Gunnar M. (1960). Acoustic theory of speech production. The Hague, Mouton.
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Henke, W. L. (1966). Dynamic Articulatory Model of Speech Production Using Computer Simulation. Unpublished doctoral dissertation, MIT, Cambridge, MA.
Honda, Takashi, Seiichi Inoue, and Yasuo Ogawa. (1968). A hybrid control system of a human vocal tract simulator. Reports of the 6th International Congress on Acoustics, ed. by Y. Kohasi, pp. 175–8. Tokyo, International Council of Scientific Unions.
Kelly, John L., and Carol Lochbaum. (1962). Speech synthesis. Proceedings of the Speech Communications Seminar, paper F7. Stockholm, Speech Transmission Laboratory, Royal Institute of Technology.
Kempelen, Wolfgang R. Von. (1791). Mechanismus der menschlichen Sprache nebst der Beschreibung seiner sprechenden Maschine. Wien, J. B. Degen.
Maeda, S. (1988). Improved articulatory model. Journal of the Acoustical Society of America, 84, Sup. 1, S146.
Maeda, S. (1990). Compensatory articulation during speech: evidence from the analysis and synthesis of vocal-tract shapes using an articulatory model. In W. J. Hardcastle and A. Marchal (Eds.), Speech Production and Speech Modelling, Kluwer Academic, Dordrecht, 131-149.
Matsui, Eiichi. (1968). Computer-simulated vocal organs. Reports of the 6th International Congress on Acoustics, ed. by Y. Kohasi, pp. 151–4. Tokyo, International Council of Scientific Unions.
Mermelstein, Paul. (1969). Computer simulation of articulatory activity in speech production. Proceedings of the International Joint Conference on Artificial Intelligence, Washington, D.C., 1969, ed. by D. E. Walker and L. M. Norton. New York, Gordon & Breach.
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Lua error in package.lua at line 80: module 'strict' not found.
Rubin, P., Saltzman, E., Goldstein, L., McGowan, R., Tiede, M., & Browman, C. (1996). CASY and extensions to the task-dynamic model. Proceedings of the 1st ESCA Tutorial and Research Workshop on Speech Producing Modeling - 4th Speech Production Seminar, 125-128.
Lua error in package.lua at line 80: module 'strict' not found.

External links

Lua error in package.lua at line 80: module 'strict' not found.
Praat
Lua error in package.lua at line 80: module 'strict' not found.
Introduction to Articulatory Speech Synthesis
Simulated singing with the singing robot Pavarobotti or a description from the BBC on how the robot synthesized the singing.

[1] Lua error in package.lua at line 80: module 'strict' not found.

[2] Talking Heads

[3] ASY

[4] CASY

[5] VocalTractLab

[6] Artisynth

[7] Real-time articulatory speech-synthesis-by-rules

[1]

[2]

[3]

[4]

[5]

[6]

[7]

v t e Speech synthesis
Free software	eSpeak Gnopernicus Gnuspeech Orca Festival Speech Synthesis System FreeTTS Sinsy Automatik Text Reader
Proprietary software	DECtalk Software Automatic Mouth Talk It! Microsoft Agent Microsoft Speech API Microsoft text-to-speech voices Readspeaker Voice browser CoolSpeech VoiceWeb BrowseAloud LaLaVoice Vocaloid Cantor Symphonic Choirs IVONA CereProc Utau Voiceroid NIAONiao Virtual Singer Vocalina Realivox CeVIO Creative Studio Chipspeech Alter/Ego PPG Phonem
Machine	Echo 2 Pattern playback Phasor RIAS Texas Instruments LPC Speech Chips TuVox
Applications	AOLbyPhone DialogOS Dr. Sbaitso MBROLA Microsoft Narrator Microsoft Speech Server PlainTalk Voice font
Protocols	Speech Synthesis Markup Language SABLE VoiceXML
Developers/ Researchers	Alan W. Black Catherine Browman Franklin Seaney Cooper Gunnar Fant Haskins Laboratories Wolfgang von Kempelen Ignatius Mattingly Philip Rubin Yamaha
Process	Articulatory synthesis Concatenative synthesis Currah Inverse filter PSOLA Phase vocoder Self-voicing

Articulatory synthesis

Contents

Mechanical talking heads

Electrical vocal tract analogs

Haskins and Maeda models

Modern models

Commercial models

See also

Footnotes

Bibliography

External links

Navigation menu

Personal tools

Namespaces

Variants

Views

More

Search

Navigation

Tools