-
Spoken Word Search in Video Editing Applications
My name is Kip Watters – I am the Director or Product Management for Media at Nexidia. Nexidia is the company whose technology powers ScriptSync in Avid’s Media Composer. ScriptSync is one way to utilize Nexidia’s “phonetic search” capability – jumping to specific lines in scenes based on text selected from the shooting script.
I am very interested in gauging the interest in and ultimately bringing these types of features to the Apple community as well.
First, a couple notes on how Nexidia’s technology is different from Speech-to-text:
1) Rather than producing a text transcript, Nexidia’s indexing process (which produces a small index file for each media file) is several hundred times faster than speech-to-text. As an example, an hour long video can be indexed in several seconds with Nexidia. The index file is not human readable, but can be searched for any number of terms/phrases once it is created.
2) Phonetic search, unlike speech to text, does not require every word to be in the “dictionary” of possible terms for it to be found. This is especially important when searching regional or topical content with unique proper name and colloquialisms. Nexidia is also more tolerant of background noise and does not require a perfect recording to be useful.
I am also interested in hearing your thoughts on other ways that search of spoken word content in audio and video would benefit the professional video editing community.
For example, would having access to a search box through Avid client applications that located all occurrences of a word or phrase (anything you typed in) within a local and or network-based content archive – similar to a Google search for untranscribed audio and video – be valuable? Search results would jump you directly to each place in the audio where the search term was spoken.
This search would not rely on manual tagging, descriptions, captions, etc. of content, and is independent of any existing metadata (title, tags, descriptions, character names, etc.) that may already be searchable. In fact, dialogue/spoken-word and metadata-based search could be combined to ensure that all information, including what the actors say, is searched together.
Would an audio search be more useful in a particular client application that you use, and if so, which applications? Does a server-based search solution, able to access content on networked drives and from multiple integrated Avid tools on the network be more useful?
Do you have examples of jobs you’re working on in which fast, on-demand access to target content via keyword/phrase search would make it easier or possible to do things that you cannot do today?
