Outer Rim Archives
Archives · 2020 · 20200394999

Application (pre-grant publication)

SYSTEM AND METHOD FOR MUSIC AND EFFECTS SOUND MIX CREATION IN AUDIO SOUNDTRACK VERSIONING

Number
20200394999
Published
2020-12-17
Filed
2019-06-11
Assignee
Lucasfilm Entertainment Company Ltd. LLC
Inventors
Levine; Scott, Morris; Stephen
CPC
G11B27/031; G10H1/0091; G06F40/58; G06F16/683; G10L25/51; G10L15/063; G06F40/279; G06N20/00; G10L15/005; G10H1/0025; G06F40/56
Verdict
Set aside audio soundtrack versioning tool, generic production software
Source
Google Patents · FreePatentsOnline

Abstract

Implementations of the disclosure describe systems and methods that leverage machine learning to automate the process of creating music and effects mixes from original sound mixes including domestic dialogue. In some implementations, a method includes: receiving a sound mix including human dialogue; extracting metadata from the sound mix, where the extracted metadata categorizes the sound mix; extracting content feature data from the sound mix, the extracted content feature data including an identification of the human dialogue and instances or times the human dialogue occurs within the sound mix; automatically calculating, with a trained model, content feature data of a music and effects (M&E) sound mix using at least the extracted metadata and the extracted content feature data of the sound mix; and deriving the M&E sound mix using at least the calculated content feature data.

Background

BRIEF SUMMARY OF THE DISCLOSURE

Implementations of the disclosure describe systems and methods that leverage machine learning to automate the process of creating music and effects (M&E) sound mixes using an original sound mix having domestic dialogue.

In one embodiment, a method includes: receiving a sound mix comprising human dialogue; extracting metadata from the sound mix, wherein the extracted metadata categorizes the sound mix; extracting content feature data from the sound mix, the extracted content feature data comprising an identification of the human dialogue and instances or times the human dialogue occurs within the sound mix; automatically calculating, with a trained model, content feature data of a music and effects (M&E) sound mix using at least the extracted metadata and the extracted content feature data of the sound mix; and deriving the M&E sound mix using at least the calculated content feature data. The content feature data extracted from the sound mix may further include one or more of: human dialogue-related data other than the identification of the human dialogue and times the human dialogue occurs within the sound mix, music-related data, and other sound data besides human dialogue-related data and music content-related data. The extracted metadata may identify one or more of the following categories of the sound mix: a domestic language, a production studio, a genre, a filmmaker, a type of media content, a re-recording mixer, a first frame

Claims

1. A method, comprising: receiving a sound mix comprising human dialogue; extracting metadata from the sound mix, wherein the extracted metadata categorizes the sound mix; extracting content feature data from the sound mix, the extracted content feature data comprising an identification of the human dialogue and instances the human dialogue occurs within the sound mix; automatically calculating, with a trained model, content feature data of a music and effects (M&E) sound mix using at least the extracted metadata and the extracted content feature data of the sound mix; and deriving the M&E sound mix using at least the calculated content feature data. 16. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, performs operations of: receiving a sound mix comprising human dialogue; extracting metadata from the sound mix, wherein the extracted metadata categorizes the sound mix; extracting content feature data from the sound mix, the extracted content feature data comprising an identification of the human dialogue and instances the human dialogue occurs within the sound mix; automatically calculating, with a trained model, content feature data of a music and effects (M&E) sound mix using at least the extracted metadata and the extracted content feature data of the sound mix; and deriving the M&E sound mix using at least the calculated content feature data.