Best practices for creating AI audio
This is a guide for anyone using the AI audio tools to turn your script into an audio file. These recommendations emphasize the importance of the script.
For more detailed instructions, read Creating AI audio from a script.
When to use AI audio
Bloomberg Connects recommends recording audio with people whenever possible. If that is beyond your resources, consider whether the AI audio tool is a good fit for your content. Here are a few examples in which it might be the right choice:
- You’re making audio content that gets updated frequently.
- Many pieces of content need a single, consistent voice over a long period of time.
- Your deadline doesn’t give you enough time to record and edit audio sessions.
- You want to hear a rough cut of a script before committing to a human recording.
- You need audio in a second language, but don’t have fluent speakers on staff.
AI audio might not be the best choice in the following scenarios:
- You and your team are already well equipped to produce audio.
- Your audio file needs to have multiple speakers.
- The speaker needs to be a specific person, like an artist or historian.
- You need fine control over tone of voice, intonation, or specific pronunciation.
Writing scripts
Creating a high-quality script will ensure the best results with this tool. First, we’ll cover general audio recommendations. After that, you’ll find recommendations that are specific to the AI audio tool.
General recommendations
Write for the ear, not the page
- Be succinct: Use relatively short sentences. Turn parenthetical phrases into separate sentences (because a listener can’t see them). What looks simplistic on the page will sound clear and coherent out loud.
- Use informal language, including contractions (it's, you'll, they're, isn't). You can start a sentence with “And” or “But.”
- Read every script aloud to ensure that it fits the natural rhythm of a spoken voice.
Engage the listener
- Write as if you're speaking directly to one person: Use “you” and “we” to make the experience feel personal and inclusive.
- Use sensory detail to bring the material to life. Invite the listener to notice what they can feel, hear, or see.
- Ask reflective or rhetorical questions. For example, “What might it have felt like to stand here, centuries ago?” or “Can you imagine the noise?”
Structure each stop as a self-contained story
- Give the most important or immediate information first. Start with the things the listener can notice right now. Add more detail as the stop continues.
- Hook the listener. Think about why they might care about the material.
- Cover one or two points per stop. A point might highlight a particular feature, technique, piece of historical context, etc.
- Order your points the way you want listeners to look at the work, so each one flows into the next.
AI audio recommendations
Clean up the script before generating your audio
- Remove any characters that are not meant to be part of the audio, including notes, bullet points, etc.
Use punctuation to shape pacing
Listen carefully to the full audio file, especially to the pace of the reading and the pronunciation of names. For most languages, you can use standard punctuation to add pauses that make the recording sound more natural.
- Commas (,) add the shortest pauses.
- Hyphens and dashes (—) add brief pauses.
- Periods (.) add full pauses, and a falling tone to indicate the end of a sentence.
- Ellipses (...) add the longest pauses.
Spell out numbers and symbols
- Use numerals for 1–999, then words for larger numbers (“one hundred and twenty-three thousand, four hundred and fifty-six,” “million”).
- Clarify currency, times, dates, symbols, and numerals (e.g., “twenty dollars”, “9 a.m.”, “nineteen ninety-seven”, “Henry the Eighth,” “20 percent”).
- Instead of using “.com”, spell out URLs with “dot com”.
Phonetic spelling
- If the AI mispronounces a name, try spelling the name phonetically. For reference, see Carnegie Mellon’s phonetic spelling instructions.
- Watch out for homographs (words spelled alike but pronounced differently, such as “read” or “tear”).
Audio length and other specifications
- We generally recommend keeping audio files under 2 minutes, or 250–300 words. Shorter audio is easier for visitors to follow, but the AI audio tool can generate files up to 5–6 minutes if your guide requires it.
- File naming: use consistent, descriptive names so files are easy to reassemble and track.
Other considerations
- Choose your voices based on what fits your institution and audio goals. We recommend limiting the number of voices across your guide to three. Visitors will enjoy the consistency and familiarity that comes with a predictable voice, or set of voices.
- Listen through to the entire clip before publishing it. Follow the guidance above to adjust your script, or find more detailed instructions in Creating AI audio from a script.
- Do not use the tool deceptively, such as a claim that it represents a recording of a particular person.
- Audio is best used as a tool to tell compelling stories about your institution. General overviews, visitor information, or other non-interpretive content may be better suited in text.
Disclosing AI to visitors
The integrity of your collection is our top priority, and we strongly recommend that you disclose any use of AI tools. When deciding where that disclosure appears in your guide, only you can know what’s best for your institution and your visitors. Connects, accordingly, does not automatically disclose the use of these tools on your behalf.
Our recommendation is that you include a disclosure in the Description field of the Audio form. For example, “[Writer name]’s script is read by an automated voice,” or “In this recording, our script is read by an AI voice.”
Other places to disclose the use of AI audio include:
- In the Description field for exhibition or tour that features AI audio files.
- In the script for the audio file itself, at the beginning or the end.
- In the Credit field for the item that features an AI audio file.
We do not recommend putting your disclosure in your audio’s Transcript field (this would imply that the transcript is AI generated), or in your guide’s “About” description (which would make it difficult for visitors to know which content was AI-assisted).