The short answer
Prosody carries meaning the words do not. It marks questions, signals which word is the important one, indicates whether the speaker has finished, and conveys attitude. For voice agents it cuts both ways: prosody in the caller’s speech is a signal for turn detection, and prosody in the agent’s speech is most of what makes it sound human or robotic.
In detail
The measurable components are fundamental frequency, which the ear hears as pitch; duration, including segment lengthening and pause placement; intensity; and voice quality features such as breathiness or creak. Together these produce the intonation contour, the stress pattern, and the phrasing of an utterance.
Prosody disambiguates sentences that are lexically identical. Rising terminal pitch turns a statement into a question. Contrastive stress relocates the point entirely: "I said Tuesday" denies who said it, while "I said Tuesday" denies which day. A transcript preserves none of this, which is one reason a text-only view of a call can mislead a reviewer about what actually happened.
On the synthesis side, prosody is the difference between a voice that is intelligible and one that is comfortable to listen to. Neural models infer it from text automatically and do so well for ordinary prose, but they have no way to know that a particular number is a confirmation code to be read slowly, or that a name should carry emphasis. That is what markup and pronunciation control exist to supply.
Related terms
Text-to-speech
Text-to-speech is the synthesis of natural-sounding spoken audio from written text.
SSML
SSML is an XML markup language for annotating text with instructions that control how a speech synthesiser renders it.
Turn detection
Turn detection is the decision that the floor has passed from one participant in a conversation to the other.
Voice cloning
Voice cloning is the creation of a synthetic voice that reproduces the vocal characteristics of a specific real person.