2009年12月19日 星期六

簡介
  在人類語言連鎖活動的過程中,最具體而可以直接觀察的部分是語音。語音可以聽得見,而且也可以記錄下來。從最具體的層次看,語音是聲波,聲波是空氣分子(particle)振動所導致的氣壓變化,是物理的現象,聲學語音學是研究語音的物理特性以及這些特性在語言系統中所具有的功能。簡單來說,聲學語音學就是一門研究「語音的物理基礎」的學問。
那麼我們為何要研究聲學語音學?因為那聽起來似乎是物理學家的工作,但是由人所發出的語音,其聲學上的特性也是語言學研究的一部份。現在,就讓我們簡單的說明幾點研究聲學語音學的理由吧!第一,我們對聲音的感覺取決於語音的聲學性質,因此,明瞭語音聲學性質的原理基本上是很重要的事。第二,語音不容易以發音動作來描述,因此以聲學特性為基礎比較容易解釋,比如母音可以以所謂的「特強頻率帶」(formant)的不同來顯示其差別。另外,聲學特性也方便我們去解釋一些容易混淆的語音。第三,說話的聲音是很短暫的,語音會隨著時間而消失。雖然語音可以模仿,甚至可以用符號記錄下來,但那與原本語音並不是同一回事,因此,獲得永久性的語音記錄是研究語音的一種助力。第四,由於近代聲學儀器的發明,我們不只可以使語音有重現上的可能,我們更可以使語音變成視覺上的記錄(如聲波圖、聲譜圖spectrograph等)。這麼一來,語音不只可以聽見,還可以“看見”。第五,在研究分析上,我們可以更進一步的克服語音瞬間消失的困難,使我們可以更詳盡的分析語音的聲學特質,以便和相關學科做結合,如:聽力學、特殊教育的聽障教學、語言治療等。
共振峰

共振峰是用來描述聲學共振現象的一種概念,[1]在語音科學及語音學中,描述的是人類聲道中的共振情形。常用的量測方法是由頻譜分析或聲譜圖(spectrogram,見右圖)中,尋找頻譜中的峰值。但假如說話者,用比較高的基頻發出母音,例如小孩或女性的聲音,則頻譜上看起來比較像是寬帶狀,比較無法看出明顯的峰值。在聲學中,共振峰是用來描述聲源內部的共振,特別是對樂器而言,指的是共嗚箱內的共振。

人類說話或唱歌產生的聲音包含許多不同的頻率,共振峰是這些頻率中較有意義的部份。定義上,人類若想分辨幾個不同的母音,我們所需要的資訊是完全可以被量化的。共振峰是使聽者能夠區分母音的關鍵泛音。大部份的這些共振峰是由管內或腔體的共振產生,

頻率最低的共振峰頻率稱為f1,第二低的是f2,而第三低的是f3。絕大多部分的情形是,前兩個共振峰,f1 和 f2就足以劃分不同母音。這兩個共振峰可以描述母音的開/閉、前/後兩個維度(過去傳統上把這和舌頭的位置聯結在一起,不過這不是完全精確)。因此開母音(例如[a])有比較高的第一共振峰頻率f1,而閉母音(例如 [i] 或 [u])的則比較低;前母音(例如[i])的第二共振峰頻率f2較高,後母音(例如[u])的則比較低。[2][3]母音幾乎都有四個以上的共振峰,有時還會超過六個。然而,前兩個共振峰還是最關鍵的。通常我們會用第一共振峰對第二共振峰的 關係圖描述不同母音的性質。[4] 但這不足以描述某些母音的性質,例如圓唇與否。[5]

2009年12月17日 星期四

Spectrogram

A spectrogram is an image that shows how the spectral density of a signal varies with time. Also known as spectral waterfalls, sonograms, voiceprints, or voicegrams, spectrograms are used to identify phonetic sounds. The instrument that generates a spectrogram is called a spectrograph or sonograph.

The most common format is a graph with two geometric dimensions: the horizontal axis represents time, the vertical axis is frequency; a third dimension indicating the amplitude of a particular frequency at a particular time is represented by the intensity or colour of each point in the image.

2009年12月16日 星期三

overview of Acoustic phonetics

Acoustic phonetics is a subfield of phonetics which deals with acoustic aspects of speech sounds. Acoustic phonetics investigates properties like the mean squared amplitude of a waveform, its duration, its fundamental frequency, or other properties of its frequency spectrum.

The study of acoustic phonetics was greatly enhanced in the late 19th century by the invention of the Edison phonograph. The phonograph allowed the speech signal to be recorded and then later processed and analyzed. By replaying the same speech signal from the phonograph several times, filtering it each time with a different band-pass filter, a spectrogram of the speech utterance could be built up. A series of papers by Ludimar Hermann published in Pflüger's Archiv in the last two decades of the 19th century investigated the spectral properties of vowels and consonants using the Edison phonograph, and it was in these papers that the term formant was first introduced. Hermann also played back vowel recordings made with the Edison phonograph at different speeds to distinguish between Willis' and Wheatstone's theories of vowel production.

Further advances in acoustic phonetics were made possible by the development of the telephone industry. (Incidentally, Alexander Graham Bell's father, Alexander Melville Bell, was a phonetician.) During World War II, work at the Bell Telephone Laboratories (which invented the spectrograph) greatly facilitated the systematic study of the spectral properties of periodic and aperiodic speech sounds, vocal tract resonances and vowel formants, voice quality, prosody, etc.

2009年11月9日 星期一

2009年10月20日 星期二

Speech synthesis

Speech synthesis is the artificial production of human speech. A computer system used for this purpose is called a speech synthesizer, and can be implemented in software or hardware. A text-to-speech (TTS) system converts normal language text into speech;

Synthesized speech can be created by concatenating pieces of recorded speech that are stored in a database. Systems differ in the size of the stored speech units; a system that stores phones or diphones provides the largest output range, but may lack clarity. For specific usage domains, the storage of entire words or sentences allows for high-quality output. Alternatively, a synthesizer can incorporate a model of the vocal tract and other human voice characteristics to create a completely "synthetic" voice output.

The quality of a speech synthesizer is judged by its similarity to the human voice and by its ability to be understood. An intelligible text-to-speech program allows people with visual impairments or reading disabilities to listen to written works on a home computer.

from http://en.wikipedia.org/wiki/Speech_synthesis

2009年10月6日 星期二

Alveolar trill

The alveolar trill is a type of consonantal sound, The symbol in the International Phonetic Alphabet that represents dental, alveolar, and postalveolar trills is [r].

In the majority of Indo-European languages, this sound is at least occasionally allophonic with an alveolar tap [ɾ], particularly in unstressed positions. Exceptions to this include Spanish, and Portuguese which treat them as separate phonemes.

Features of the alveolar trill:

Its manner of articulation is trill, which means it is produced by vibrations of the tongue against the place of articulation.

在一些語言, 例如捷克語, 齒齦顫音可以作為成節輔音 , 例如捷克語 krk (頸).

多數漢語中的方言都沒有此音。但湖北的中北部的一部分中原官話區和西南官話區里,即當陽江陵鍾祥京山一帶,直至神農架北部地區,存在有顫音/r/。而且此種顫音/r/,多半是由詞尾「子」演變而來。另外,某些普通話官話方言的民歌如鳳陽花鼓裡,歌詞「得兒叮噹飄一飄」裡,如果讀時舌肌放鬆,就會讀出清齒齦顫音。

from http://en.wikipedia.org/wiki/Alveolar_trill