The short answer

Prosody carries meaning the words do not. It marks questions, signals which word is the important one, indicates whether the speaker has finished, and conveys attitude. For voice agents it cuts both ways: prosody in the caller’s speech is a signal for turn detection, and prosody in the agent’s speech is most of what makes it sound human or robotic.

In detail

The measurable components are fundamental frequency, which the ear hears as pitch; duration, including segment lengthening and pause placement; intensity; and voice quality features such as breathiness or creak. Together these produce the intonation contour, the stress pattern, and the phrasing of an utterance.

Prosody disambiguates sentences that are lexically identical. Rising terminal pitch turns a statement into a question. Contrastive stress relocates the point entirely: "I said Tuesday" denies who said it, while "I said Tuesday" denies which day. A transcript preserves none of this, which is one reason a text-only view of a call can mislead a reviewer about what actually happened.

On the synthesis side, prosody is the difference between a voice that is intelligible and one that is comfortable to listen to. Neural models infer it from text automatically and do so well for ordinary prose, but they have no way to know that a particular number is a confirmation code to be read slowly, or that a name should carry emphasis. That is what markup and pronunciation control exist to supply.

Back to the glossary

All 53 glossary terms