Outer Rim Archives
Archives · 2023 · 20230237261

Application (pre-grant publication)

Extended Vocabulary Including Similarity-Weighted Vector Representations

Number
20230237261
Published
2023-07-27
Filed
2022-01-21
Assignee
Disney Enterprises, Inc.
Inventors
Farre Guiu; Miquel Angel et al.
CPC
G06F16/48; G06F40/279; G06F16/3344; G06F40/169; G06F40/242; G06F40/284; G06F40/30; G06N3/08; G06N20/00
Verdict
Set aside NLP vocabulary research, generic/business
Source
Google Patents · FreePatentsOnline

Abstract

According to one implementation, a system includes a computing platform having processing hardware, and a system memory storing a software code. The processing hardware is configured to execute the software code to receive a vocabulary, identify words from the vocabulary for use in extending the vocabulary, pair each of those words with every other of those words to provide word pairs, and output the word pairs to a vocabulary administrator. The software code also receives word pair characterizations identifying each of the word pairs as one of similar, dissimilar, or neither similar nor dissimilar, configures, based on the word pair characterizations, a multi-dimensional vector space including multiple embedding vectors each corresponding respectively to one of the identified words, and cross-references each of those words with its corresponding embedding vector to produce an extended vocabulary corresponding to the received vocabulary.

Background

BACKGROUND

The annotation, or “tagging,” of digital media content may be performed manually by human taggers, or in an automated or substantially automated process, based on a predetermined taxonomy including a vocabulary of words that may be applied as annotation tags. Each word included in the taxonomy is purposefully selected for use as a tag and has a carefully defined scope and intended application that is well understood by the librarians or administrators of the taxonomy. Nevertheless, given the subjective nature of manual tagging, and the ambiguity associated with some of the features of content to which tags are to be applied, both human and automated taggers may reinterpret tags and apply them in a manner that is inconsistent with their intended use.

Due to its popularity, ever more digital media content is being produced and made available to consumers. As a result, the efficiency and accuracy with which digital media content can be tagged and managed has become increasingly important to the producers, owners, and distributors of that content. For example, tagging of video is an important part of the production, distribution, and recommendation processes for television (TV) content and movies. Consequently, there is a need in the art for systems and methods enabling the consistently accurate application of the tags included in annotation taxonomies by automated and human taggers alike.

Claims

1. A system comprising: a computing platform including processing hardware and a system memory storing a software code; the processing hardware configured to execute the software code to: receive a vocabulary including a first plurality of words; identify, from among the first plurality of words, a second plurality of words for extending the received vocabulary; pair each word of the second plurality of words with every other word included among the second plurality of words to provide a plurality of word pairs; output the plurality of word pairs to a vocabulary administrator; receive a plurality of word pair characterizations identifying each of the plurality of word pairs as one of similar, dissimilar, or neither similar nor dissimilar; configure, based on the plurality of word pair characterizations, a multi-dimensional vector space including a plurality of embedding vectors each corresponding respectively to one of the second plurality of words; and cross-reference each of the second plurality of words with its corresponding embedding vector to produce an extended vocabulary corresponding to the received vocabulary. || 11. A method for use by a system including a computing platform having processing hardware and a system memory storing a software code, the method comprising: receiving, by the software code executed by the processing hardware, a vocabulary including a first plurality of words; identifying, from among the first plurality of words by the software code executed by the processing hardware, a second plurality of words for extending the received vocabulary; pairing, by the software code executed by the processing hardware, each word of the second plurality of words with every other word included among the second plurality of words to provide a plurality of word pairs; outputting, by the software code executed by the processing hardware the plurality of word pairs to a vocabulary administrator; receive, by the software code executed by the processing hardware, a plurality of word pair characterizations identifying each of the plurality of word pairs as one of similar, dissimilar, or neither similar nor dissimilar; configuring, by the software code executed by the processing hardware based on the plurality of word pair characterizations, a multi-dimensional vector space including a plurality of embedding vectors each corresponding respectively to one of the second plurality of words; and cross-referencing, by the software code executed by the processing hardware, each of the second plurality of words with its corresponding embedding vector to produce an extended vocabulary corresponding to the received vocabulary.