I know this is an old post, but I wonder if anyone has any thoughts about this scenario:
I already have verbatim manuscripts of videos. I want to automate the creation of subtitle files. Will any existing speech to text programs spit out a transcript with time codes? Then I could write a program that would match the imprecise text in the transcript to the precise text in the manuscript. That way my subtitles would be perfect even though the speech to text is imperfect. Better yet, is there already a program that does this kind of thing?