Outer Rim Archives
Archives · 2025 · 20250280239

Application (pre-grant publication)

AUTOMATIC DETECTION OF ALIGNMENT BETWEEN TWO AUDIO SIGNALS

Number
20250280239
Published
2025-09-04
Filed
2024-07-18
Assignee
Disney Enterprises, Inc.
Inventors
Abecassis; Eitan et al.
CPC
G10L25/57; H04R5/04; G10L21/055; G11B27/10; G11B27/36; H04N21/4394; H04N21/8106; H04S7/301
Verdict
Set aside audio-signal alignment, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

In some embodiments, a method analyzes a first sample of a first audio signal to determine a first representation in a space. A plurality of second samples for a second audio signal is analyzed to determine a plurality of second representations in the space. The method compares the first representation and the plurality of second representations in the space to select a second representation. An offset is determined between the first sample and a second sample that is associated with the second representation. The offset is output.

Background

BACKGROUND

Audio quality control is largely a manual process for a company. Audio errors can occur in a variety of places in a content pipeline starting from content mastering and extending all the way to content distribution. The misalignment between audio and video is one of the most distracting quality defects in media consumption today. Even a slight offset in the audio and video may be noticeable to a viewer. In the film industry, audio and video synchronization issues may be key drivers to viewer disengagement. While media synchronization issues happen for a myriad of reasons, one area that is more susceptible to these errors is dubbing. In some examples, the visual component of the dubbed media is unchanged, but the audio track is replaced with a translated rendition of the original language. Producing and inserting the new audio track may result in a de-synchronization with the original audio track. As dubbed content consistently grows in popularity, the problem of unidentified synchronization errors becomes increasingly prevalent.

The quality control of dubbed media may mostly be a manual process. The quality checks may include users who are specialized in listening for differences in the audio or quality control operators visually comparing raw waveforms for any errors. These subjective evaluations may be costly, inefficient, and also not able to detect small isolated errors.

Claims

1. A method comprising: analyzing a first sample of a first audio signal to determine a first representation in a space; analyzing a plurality of second samples for a second audio signal to determine a plurality of second representations in the space; comparing distances in the space between the first representation and the plurality of second representations in the space to select a second representation; determining an offset between the first sample in the first audio signal and a second sample in the second audio signal that is associated with the second representation; and outputting the offset. || 18. A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for: analyzing a first sample of a first audio signal to determine a first representation in a space; analyzing a plurality of second samples for a second audio signal to determine a plurality of second representations in the space; comparing the first representation and the plurality of second representations in the space to select a second representation; determining an offset between the first sample and a second sample that is associated with the second representation; and outputting the offset. || 20. An apparatus comprising: one or more computer processors; and a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for: analyzing a first sample of a first audio signal to determine a first representation in a space; analyzing a plurality of second samples for a second audio signal to determine a plurality of second representations in the space; comparing the first representation and the plurality of second representations in the space to select a second representation; determining an offset between the first sample and a second sample that is associated with the second representation; and outputting the offset.