Transcripts and speakers
Word-level transcription and speaker detection are what let you edit, caption and search footage by its words.
The transcript is the backbone of editing in Eddie. Because it has word-level timing, you can edit, caption and search your footage by what was actually said.
Word-level timing
Every word is time-stamped, so cuts land precisely and captions highlight on the beat.
Speaker changes
Eddie marks where the person talking changes — useful for interviews and panels, and it's what makes prompts like "keep only the guest's answers" possible.
Why it matters for prompting
You can point Eddie at moments by content — "the part about pricing", "when she says 'the key thing is'" — because it's reading the transcript, not guessing from the waveform.
Languages
Eddie transcribes 100+ languages — including Chinese (Mandarin and Cantonese), Japanese, Korean and Thai, plus a large second tier and regional variants like Canadian French and Latin-American Spanish. Editing, search and captions all work in-script: captions render and wrap correctly in each writing system. See Languages for the full picture and the current right-to-left caveat.
Spell names right with attached context
Attach a Backgrounder (or lean on your brand kit) before importing and Eddie hands your names, product terms and proper nouns to the transcription engine as a keyterm dictionary — biasing it toward the right spelling as it listens. It's the most reliable way to get people's names and brand terms correct throughout titles, captions and summaries.
Getting the best transcript
- Import source footage with clean audio where possible.
- Multi-track audio (lav + camera mic) is detected and mixed on import, which improves clarity.
- Names, products and unusual spellings: attach a Backgrounder dictionary before importing so they're recognized from the first pass.
Related: Languages, Captions and subtitles.