Outer Rim Archives
Archives · 2026 · 20260212199

Application (pre-grant publication)

ADAPTIVE MOTION CONTROL VIA MULTI-OBJECTIVE REINFORCEMENT LEARNING

Number
20260212199
Published
2026-07-23
Filed
2026-01-22
Assignee
DISNEY ENTERPRISES, INC.
Inventors
BÄCHER; Moritz Niklaus, KNOOP; Lars Espen, GRANDIA; Ruben Jelle, SERIFI; Agon, ALEGRE; Lucas Nunes, CAO; David Benjamin
CPC
G06N3/006; G06N3/092; G06N3/008; G06N3/044; G06N3/045; G06N3/047; G06N3/0475; G06N3/08; G06N3/084; G06N3/088; G06N7/01; G06N20/00
Verdict
High Notable software
First reported
2026-W30 (2026-07-24)
Source
Google Patents · FreePatentsOnline

The keeper's note

One embodiment of the present invention sets forth a technique for controlling motion in an articulated object.

Abstract

One embodiment of the present invention sets forth a technique for controlling motion in an articulated object. The technique includes generating, via execution of a machine learning model, one or more actions based on (i) a first state of the articulated object at a first time and (ii) a first plurality of weights associated with a plurality of rewards. The technique also includes generating, via execution of the machine learning model, one or more additional actions based on (i) a second state of the articulated object at a second time and (ii) a second plurality of weights associated with the plurality of rewards. The technique further includes causing a task associated with the articulated object to be performed based on the one or more actions and the one or more additional actions.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to motion tracking and reinforcement learning and, more specifically, to adaptive motion control via multi-objective reinforcement learning. Description of the Related Art

Physics-based character control is a technique for generating motion in physical and/or virtual characters in a physically realistic and robust manner. To achieve this type of motion, a controller computes actions (e.g., joint torques, target positions, actuator commands, etc.) that cause a character to move in a desired manner while respecting physics constraints such as (but not limited to) gravity, momentum, friction, and/or contact forces. The actions are used to update joints of the character and produce physically plausible motion in a robot, game, animation, simulation, and/or another application involving the character.

Existing approaches for performing physics-based character control include the use of reinforcement learning (RL) to train a control policy to output actions that maximize a reward function. The reward function can include one or more objectives related to the accuracy with which a reference motion is tracked. When multiple objectives are included in the reward function, a set of weights is used to control the relative priorities and/or effects of the objectives on the outputted actions.

However, these approaches have traditionally used reward functions with

Claims

1. A computer-implemented method for controlling motion in an articulated object, the method comprising: generating, via execution of a machine learning model, one or more actions based on (i) a first state of the articulated object at a first time and (ii) a first plurality of weights associated with a plurality of rewards; generating, via execution of the machine learning model, one or more additional actions based on (i) a second state of the articulated object at a second time and (ii) a second plurality of weights associated with the plurality of rewards; and generating the motion based on the one or more actions and the one or more additional actions. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: generating, via execution of a machine learning model, one or more actions based on (i) a first state of an articulated object at a first time and (ii) a first plurality of weights associated with a plurality of rewards; generating, via execution of the machine learning model, one or more additional actions based on (i) a second state of the articulated object at a second time and (ii) a second plurality of weights associated with the plurality of rewards; and generating a motion based on the one or more actions and the one or more additional actions. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: generating, via execution of a machine learning model, one or more actions at a first time based on (i) a first state of an articulated object at a first time step and (ii) a first plurality of weights associated with a plurality of rewards; generating, via execution of the machine learning model, one or more additional actions at a second time based on (i) a second state of the articulated object at a second time step and (ii) a second plurality of weights associated with the plurality of rewards; and computing a plurality of reward values for the plurality of rewards based on at least one of the first state, the one or more actions, or the second state; computing one or more losses based on the plurality of reward values and the first plurality of weights; and training the machine learning model based on the one or more losses.