- Number
- 11182431
- Published
- 2021-11-23
- Filed
- 2014-12-11
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Wang; Jing X., Arana; Mark, Drake; Edward, Chen; Alexander C.
- CPC
- G06F16/433; G06F16/48; G06F16/7844; G06F16/90332; G06F16/683; G06F16/7867; G10L25/54; G06F16/44; H04N21/4826; H04N21/4828; G06F16/4387
- Verdict
- Set aside voice content search, business
- Source
- Google Patents · FreePatentsOnline
Abstract
Systems and methods for voice searching media content based on metadata or subtitles are provided. Metadata associated with media content can be pre-processed at a media server. Upon receiving a vocal command representative of a search for an aspect of the media content, the media server performs a search for one or more portions of the media content relevant to the aspect of the media content being searched for. The media performs the search by matching the aspect of the media content being searched for with the pre-processed metadata.
Background
TECHNICAL FIELD (1) The present disclosure relates generally to media content playback and interaction. DESCRIPTION OF THE RELATED ART (2) Traditional methods of interacting with media content via a digital video disk (DVD) or video cassette recorder (VCR) generally rely on actuating playback buttons or controls. For example, a user may fast forward or rewind through portions of the media content, e.g., scenes of a movie, to achieve playback of a particular portion of the media content that the user wishes to view or experience. Media interaction on devices such as smart phones, laptop personal computers (PCs), and the like mimic such controls during playback of media content being streamed or downloaded to the device. BRIEF SUMMARY OF THE DISCLOSURE (3) In accordance with one embodiment, a computer-implemented method comprises receiving vocal input from a user via a user device. The computer-implemented method further comprises searching for at least one portion of media content based on the vocal user input. Further still, the computer-implemented method comprises providing access to the at least one portion of the media content via the user device. (4) In accordance with another embodiment, an apparatus comprises a content database containing one or more media content files. The apparatus further comprises a speech recognition unit configured to recognize voice commands representative of a search for at least one portion of the one or more media content files. Additionally
Claims
1. A computer-implemented method, comprising: receiving, by a media content server, media content that includes a first scene associated with first metadata and a second scene associated with second metadata, wherein the first metadata comprises information provided by an originator of the media content, wherein the second metadata comprises automatically generated information provided by a production or editing process of the media content; initiating a metadata generation process to generate, based on feedback received from at least a first-user device, third metadata associated with a third scene included in the media content; providing access to the media content by the media content server to a second-user device; initiating a speech-to-text recognition process to translate a vocal user input to one or more search criteria, wherein the vocal user input is received from the second-user device; determining, using a hierarchical search, a plurality of scenes relevant to the one or more search criteria based on at least the first metadata, the second metadata, and the third metadata, the plurality of scenes being non-contiguous in the media content, wherein the plurality of scenes is customized based on second-user classifications of second-user favorite scenes; generating derivative media content by stitching together the plurality of scenes relevant to the one or more search criteria into a single media content file based on second-user selection of at least one of a plurality of stitching options; and providing access to the derivative media content by the media content server to the second-user device, wherein the second-user device is configured to display the derivative media content on a display of the second-user device. ||
4. An apparatus, comprising: a content database containing media content; a memory including computer program code, the memory and the computer program code configured to, with a processor, cause the apparatus to perform at least the following: receive, by a media content server, media content that includes a first scene associated with first metadata and a second scene associated with second metadata, wherein the first metadata comprises information provided by an originator of the media content, wherein the second metadata comprises automatically generated information provided by a production or editing process of the media content; initiate a metadata generation process to generate, based on feedback received from at least a first-user device, third metadata associated with a third scene included in the media content; provide access to the media content by the media content server to a second-user device; receive vocal user input from the second-user device; initiate a speech-to-text recognition process to translate the vocal user input to one or more search criteria; determine, using a hierarchical search, a plurality of scenes relevant to the one or more search criteria based on at least the first metadata, the second metadata, and the third metadata, the plurality of scenes being non-contiguous in the media content, wherein the plurality of scenes is customized based on second-user classifications of second-user favorite scenes; generate derivative media content by stitching together the plurality of scenes relevant to the one or more search criteria into a single media content file based on second-user selection of at least one of a plurality of stitching options; and provide access to the derivative media content by the media content server to the second-user device, wherein the second-user device is configured to display the derivative media content on a display of the second-user device. ||
5. A media content server, comprising: a processor; and a memory including computer program code executable by the processor to perform at least the following: receive media content that includes a first scene associated with first metadata and a second scene associated with second metadata, wherein the first metadata comprises information provided by an originator of the media content, wherein the second metadata comprises automatically generated information provided by a production or editing process of the media content; initiate a metadata generation process to generate, based on feedback received from at least a first-user device, third metadata associated with a third scene included in the media content; provide access to the media content by the media content server to a second-user device; receive vocal user input from the second-user device; initiate a speech-to-text recognition process to translate the vocal user input to one or more search criteria; determine, using a hierarchical search, a plurality of scenes relevant to the one or more search criteria based on at least the first metadata, the second metadata, and the third metadata, the plurality of scenes being non-contiguous in the media content, wherein the plurality of scenes is customized based on second-user classifications of second-user favorite scenes; generate derivative media content by stitching together the plurality of scenes relevant to the one or more search criteria into a single media content file based on second-user selection of at least one of a plurality of stitching options; and provide access to the derivative media content by the media content server to the second-user device, wherein the second-user device is configured to display the derivative media content on a display of the second-user device.