Transcripts and speakers
Word-level transcription and speaker detection are what let you edit, caption and search footage by its words.
The transcript is the backbone of editing in Eddie. Because it has word-level timing, you can edit, caption and search your footage by what was actually said.
Word-level timing
Every word is time-stamped, so cuts land precisely and captions highlight on the beat.
Speaker changes
Eddie marks where the person talking changes — useful for interviews and panels, and it's what makes prompts like "keep only the guest's answers" possible.
Why it matters for prompting
You can point Eddie at moments by content — "the part about pricing", "when she says 'the key thing is'" — because it's reading the transcript, not guessing from the waveform.
Languages
Eddie transcribes 100+ languages — including Chinese (Mandarin and Cantonese), Japanese, Korean and Thai, plus a large second tier and regional variants like Canadian French and Latin-American Spanish. Editing, search and captions all work in-script: captions render and wrap correctly in each writing system. See Languages for the full picture and the current right-to-left caveat.
Spell names right with attached context
Attach a Backgrounder (or lean on your brand kit) before importing and Eddie hands your names, product terms and proper nouns to the transcription engine as a keyterm dictionary — biasing it toward the right spelling as it listens. It's the most reliable way to get people's names and brand terms correct throughout titles, captions and summaries.
Fixing a name or word after the fact
There's currently no tool to edit a single word's spelling directly inside an existing transcript or caption track — Eddie can only remove transcript content outright, or re-transcribe a source from scratch, neither of which lets it swap one spelling for another after the fact (a full re-transcription can't distinguish two spellings that sound identical anyway). If you spot a misspelled name or term after importing, the reliable fix is to add it to a Backgrounder dictionary and re-import or re-transcribe the source before you rely on captions or titles, so the correct spelling is used from the first pass. Catching it before import is much more effective than trying to correct it afterward.
Getting the best transcript
- Import source footage with clean audio where possible.
- Multi-track audio (lav + camera mic) is detected and mixed on import, which improves clarity.
- Names, products and unusual spellings: attach a Backgrounder dictionary before importing so they're recognized from the first pass.
Related: Languages, Captions and subtitles.