- Number
- 11581970
- Published
- 2023-02-14
- Filed
- 2021-04-21
- Assignee
- Lucasfilm Entertainment Company Ltd. LLC
- Inventors
- Morris; Stephen et al.
- CPC
- G06F3/165; G11B27/031; H04S3/002; H04H60/04; G10L25/57; G06V20/46
- Verdict
- Set aside audio-production pipeline tool, business
- Source
- Google Patents · FreePatentsOnline
Abstract
Some implementations of the disclosure relate to using a model trained on mixing console data of sound mixes to automate the process of sound mix creation. In one implementation, a non-transitory computer-readable medium has executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix obtaining mixing metadata; automatically calculating with a trained model, using at least the mixing metadata and the first audio features, mixing console features; and deriving a second version of the sound mix using at least the mixing console features calculated by the trained model.
Background
BRIEF SUMMARY OF THE DISCLOSURE (1) Implementations of the disclosure describe systems and methods that leverage machine learning to automate the process of creating various versions of sound mixes. (2) In one embodiment, a non-transitory computer-readable medium has executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix obtaining mixing metadata; automatically calculating with a trained model, using at least the mixing metadata and the first audio features, mixing console features; and deriving a second version of the sound mix using at least the mixing console features calculated by the trained model. (3) In some implementations, deriving the second version of the sound mix, comprises: inputting the mixing console features derived by the trained model into a mixing console for playback; and recording an output of the playback. (4) In some implementations, deriving the second version of the sound mix, comprises: displaying to a user, in a human readable format, one or more of the mixing console features derived by the trained model. In some implementations, deriving the second version of the sound mix, further comprises: receiving data corresponding to one or more modifications input by the user modifying one or more of the displayed mixing console features derived by the trained model; an
Claims
1. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix obtaining mixing metadata; automatically calculating with a trained model, using at least the mixing metadata and the first audio features, mixing console features, the mixing console features comprising console automation data including time-domain control values for one or more audio processing components for an audio channel; and deriving, using at least the mixing console features calculated by the trained model, a second version of the sound mix. ||
3. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix obtaining mixing metadata; automatically calculating with a trained model, using at least the mixing metadata and the first audio features, mixing console features; and deriving, using at least the mixing console features calculated by the trained model, a second version of the sound mix, wherein deriving the second version of the sound mix, comprises: displaying to a user, in a human readable format, one or more of the mixing console features derived by the trained model. ||
5. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix obtaining mixing metadata; extracting video features from video corresponding to the first version of the sound mix; automatically calculating with a trained model, using at least the mixing metadata, the first audio features, and the video features, mixing console features; and deriving, using at least the mixing console features calculated by the trained model, a second version of the sound mix. ||
11. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix; obtaining mixing metadata; automatically calculating with a trained model, using at least the mixing metadata and the first audio features, mixing console features; automatically calculating with the trained model, using at least the mixing metadata and the first audio features, second audio features for deriving a second version of the sound mix; displaying to a user a first option to derive the second version of the sound mix using the mixing console features, and a second option to derive the second version of the sound mix using the second audio features; receiving input from the user selecting the first option; and deriving, using at least the mixing console features calculated by the trained model, the second version of the sound mix. ||
12. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, causes the processor to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix extracting video features from video corresponding to the first version of the sound mix; obtaining mixing metadata; and automatically calculating with a trained model, using at least the mixing metadata, the first audio features, and the video features: second audio features corresponding to a second version of the sound mix; or pulse-code modulation (PCM) audio or coded audio corresponding to a second version of the sound mix. ||
15. A sound mixing system, comprising: one or more processors; and one or more non-transitory computer-readable mediums having executable instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: obtaining a first version of a sound mix; extracting first audio features from the first version of the sound mix obtaining mixing metadata; automatically calculating with a trained model, using at least the mixing metadata and the first audio features, mixing console features, the mixing console features comprising console automation data including time-domain control values for one or more audio processing components for an audio channel; and deriving, using at least the mixing console features calculated by the trained model, a second version of the sound mix.