← All guidesVoice

Expressive AI voiceover: laughs, whispers and pauses

Flat AI narration loses viewers. Audio tags like [laughs], [whispers] and [pause] make the voice act. Which tags exist, where to put them, and which voice models perform them.

By , maker of Seedflows · Updated · 6 min read

Most AI narration has the same problem: every sentence sounds equally calm. Real narrators laugh, hesitate, whisper and leave silence before the big moment. Modern voice models can do all of that, if you tell them where.

What audio tags are

Audio tags are short instructions in square brackets, written directly into the script. The voice performs them instead of reading them:

[whispers] Don't move. [nervous] Something is down here with us, [pause] and I don't think it's friendly.

Hear the difference

These samples come straight out of Seedflows, unedited:

Joy

[excited] We did it, we actually found it! [laughs] After all those nights, the river was right there under the mountain the whole time!

Tension

[whispers] Don't move. [nervous] Something is down here with us, [pause] and I don't think it's friendly.

Four moods in one take

[curious] Hold on, what's carved into this door? [gasps] It's our names, both of them, written hundreds of years ago! [sarcastic] Oh, perfect, that's not creepy at all. [mischievously] Well, there's only one way to find out what's behind it.

The tags you can use

KindTags
Sounds[laugh], [chuckle], [sigh], [gasp], [cough], [clears throat], [sniff], [groan], [crying]
Delivery[whispers], [shouts], [excited], [sad], [angry], [sarcastic], [curious]
Narration[serious], [softly], [warmly], [concerned], [urgent], [thoughtful], [hopeful], [ominous]
Timing[pause], [long pause]

Short free-form tags such as [nervous] or [mischievously] work too, as long as they are one or two English words, even if the script is in another language.

Where to put them

  • Before the words they colour.[sad] He never came back. A tag at the end of a sentence comes too late and is often ignored.
  • Where the mood changes, not on every sentence. For narration, one tag every two to four sentences is plenty.
  • [pause] before a reveal, not as decoration. Silence works because it is rare.

Let Enhance do the first pass. In the script editor, Enhance reads the whole script and adds fitting tags without changing a single word. Then keep what fits and delete the rest.

Which voices perform the tags

Not every voice model understands tags. Seedflows keeps the tags in your script and adapts them to the voice you pick, so you never have to rewrite anything:

  • ElevenLabs v3 and v4 perform all of them, including free-form tags.
  • ElevenLabs Multilingual v2 and Flash only turn [pause] and [long pause] into real pauses; other tags are removed before sending.
  • Chatterbox Turbo performs the sounds such as laughs, sighs and coughs.
  • Other engines read the plain text; tags are removed automatically.

To use v4 with an ElevenLabs voice, open the voice under My Voices, choose the model and save.

Several characters, several voices

Dialogue gets much better when every character has their own voice. Write the script with speakers, give each one a voice, and Seedflows voices the whole conversation, tags included. Combined with consistent faces, your characters become people viewers recognise.