Shape the sound Fictional AI presenter; AI-generated narration. Shape the sound Welcome to module 5. I'm Alex, a fictional AI presenter for Vidmoat Learning. In this lesson, we will keep speech clear, use music purposefully and distinguish verified audio from AI estimates. Let's work through a practical example. Make the voice comfortable first A viewer may tolerate a slightly plain picture while learning something useful. A voice that is painful, muffled or buried is much harder to follow. Start with the original recording and headphones at a comfortable level. Listen for room noise, sudden level changes and distortion. Make the voice comfortable first Lowering the volume of a distorted recording does not undo the distortion. Noise reduction can soften steady background noise, but too much may make a voice watery or metallic. Compression reduces the difference between louder and quieter moments; it is not a cure for every recording problem. Make one modest adjustment and compare it with the untreated version. Make the voice comfortable first Select the intended clip before changing A/V or audio properties. If the voice and music were recorded together, ordinary clip volume changes both. Separating them perfectly is not something to assume. When a line is beyond repair, re-recording it in a quieter place may take less time and produce a better result than adding more processing. Try it in Vidmoat Here is a way to try this. Mute added music temporarily and listen to the main voice. Select the speaking clip and open its A/V or audio properties. Adjust obvious level differences so sentences are comfortable to hear. If needed, apply a small amount of noise reduction and compare before and after. Listen on headphones and a phone speaker. Re-record an unusable line if possible. Make the decision yourself Now pause the lesson and make your own version. Record the same sentence near a fan and in a quiet corner. Try a gentle cleanup on the noisy take, then compare it with the quieter recording at similar listening levels. When you are ready, compare your result with this standard. The voice is intelligible and natural. Your decision is based on listening, not on how many processing controls are enabled. Choose music that supports the edit Music changes how a scene feels before a viewer understands the words. A bright, light rhythm can suit a friendly tech review. A dense vocal song can compete with the review because the listener is trying to follow two sets of words at once. Choose by the job the music must do, not just by whether you like the track. Choose music that supports the edit Start with the voice at a comfortable level, then bring the music up from silence until it supports the pace. If the first sentence becomes harder to understand, the music is too prominent for that moment. Use fades to avoid abrupt starts and stops. Ducking lowers music while speech is present; check that it does not audibly pump between every word. Choose music that supports the edit The included practice-bed.wav is an original thirty-second rhythm for this course. It is useful for practising levels and fades, not a promise that one track will suit every genre. You can also use music you own or are licensed to use. A recognisable commercial song is not automatically permitted because it is available online. Try it in Vidmoat Here is a way to try this. Import practice-bed.wav through the media library and place it on an audio track. Trim it to the edit length and start it quietly under the voice. Apply short fades at the opening and ending. Use Duck under speech if available, or lower the bed manually during narration. Listen to the quietest sentence and the ending. Adjust levels and fades before exporting. Make the decision yourself Now pause the lesson and make your own version. Make one mix with music too loud, then a second with the voice clearly leading. Compare at the same device volume. Write what became easier to hear. When you are ready, compare your result with this standard. Every word is understandable, the music enters smoothly, and the ending sounds intentional. Ask the agent to listen and verify what it reports Speech transcription and audio understanding answer different questions. A transcript asks what was said. Audio inspection can help describe sounds, locate energy changes or estimate a beat. Lyric listening asks about sung or rapped words. A failed speech transcript does not prove that a song has no vocals. Ask the agent to listen and verify what it reports Give the agent a short range and a specific question. In this edition, source audio inspection normally begins with thirty seconds and supports explicitly selected excerpts up to two minutes. Results describe the inspected excerpt, not necessarily the final mix with every effect and track. Ask whether reported times are source, clip-local or timeline times before cutting on them. Ask the agent to listen and verify what it reports Lyric phrases may include uncertainty and their timing may drift. Listen to the actual recording and keep unclear words unresolved until you can verify them. Sound descriptions cannot reliably establish the song title or artist. The agent can also fail or be unavailable; that should lead to a manual check, not a guessed soundtrack description. Try it in Vidmoat Here is a way to try this. Use sound-check.wav from the resource kit as a simple control. Ask the agent to inspect the first ten seconds and describe the audible events. Listen yourself: the file has two short tones beginning at two and six seconds. For your own vocal music, request lyric listening and uncertainty explicitly. Confirm words and edit points manually before turning the result into captions or cuts. Make the decision yourself Now pause the lesson and make your own version. Inspect the two-tone control and compare the report with what you hear. Then record or use an original sung phrase and test lyric listening. If unavailable, annotate it yourself and continue. When you are ready, compare your result with this standard. You distinguish measured timing, model description and your own listening. You do not treat estimated lyric phrases as word-perfect karaoke captions. Keep one useful decision Before moving on, write down one edit you made, why you chose it, and what changed for the viewer. The course workbook and practice files are available with this lesson. Continue when you can explain your decision in your own words.