Performance directives let you tell the voice engine how a specific line should be delivered — whispered, shouted, hesitant, pleading — rather than accepting one flat default delivery for every line a character speaks. Used well, this is one of the highest-leverage tools available for making a scene feel genuinely performed rather than simply read.
How the syntax works
A performance directive is written as a short phrase in asterisks directly before the line it applies to — for example, *whispering* before a line of dialogue. The system recognizes a set of known styles (whispering, shouting, crying, nervous, laughing, angry, sad, excited, frightened, hesitant, detached, pleading, sarcastic, exhausted, trembling, desperate, bitter, tender, shocked, and calm) and adjusts the voice's delivery accordingly. If you use a style that isn't specifically recognized, it still passes through with a sensible default rather than breaking the line — so you're never blocked from trying something.
Where directives matter most
- Emotional turning points. The line where a scene's tone shifts — from calm to angry, from steady to frightened — is exactly where a directive earns its place.
- Subtext-heavy lines. A line that means something different depending on delivery (sarcastic versus sincere) needs the directive to land correctly, since the words alone don't disambiguate it.
- Quiet or intimate moments. Whispered or hesitant delivery rarely happens by default, and often needs to be called out explicitly to read correctly.
Where to leave lines alone
Not every line needs a directive. Neutral, functional dialogue — a character simply stating something plainly — usually reads fine at a character's baseline delivery. Overusing directives on every single line tends to flatten their impact rather than sharpen it; save them for the moments that actually call for a shift.
Testing and iterating
Since directives affect voice delivery meaningfully, it's worth generating a scene once, listening for any line where the delivery feels off, and adjusting or adding a directive specifically to that line rather than guessing across the whole scene upfront. This iterative approach tends to get better results than trying to perfectly direct every line before hearing any of it.