A system includes a computing platform having a hardware processor, and a system memory storing a software code and a content labeling predictive model. The hardware processor is configured to execute the software code to scan a database to identify content assets stored in the database, parse metadata stored in the database to identify labels associated with the content assets, and generate a graph by creating multiple first links linking each of the content assets to its corresponding label or labels. The hardware processor is configured to further execute the software code to train, using the graph, the content labeling predictive model, to identify, using the trained content labeling predictive model, multiple second links among the content assets and the labels, and to annotate the content assets based on the second links.
BACKGROUND (1) The archive of content assets saved and stored after being utilized one or a relatively few times in the course of movie production is vast, complex, and difficult or effectively impossible to search. The vastness of such an archive can easily be attributed to the sheer number of movies created over the several decades of movie making. The complexity of such an archive may be less apparent, but arises from the diversity or heterogeneity of the content collected together. For example, movie making may require the use of content assets in the form of images, both still and video, three-dimensional (3D) models, and textures, to name a few examples. Furthermore, each exemplary media type may be stored using multiple file formats, which may differ across the different media types. That is to say, for instance, video may be stored in multiple different file formats, each of which is different from the file formats used to store textures or 3D models. (2) Due to the heterogeneity of the archived content assets described above, making effective reuse of those assets poses a substantial challenge. For example, identifying archived content typically required a manual search by a human archivist. Moreover, because the archived content assets are often very sparsely labeled by metadata, the manual search may become a manual inspection of each of thousands of content assets, making the search process costly, inefficient, and in most instances completely impracticable. Conse
1. A system comprising: a computing platform including a hardware processor and a system memory storing a software code and a content labeling predictive model; the hardware processor configured to execute the software code to: scan a database to identify a plurality of content assets stored in the database; parse metadata stored in the database to identify a plurality of labels associated with the plurality of identified content assets; generate graph data by creating a plurality of first links linking each one of the plurality of identified content assets to each of the plurality of identified labels associated with a corresponding one of the plurality of identified content assets; use at least one classifier network to generate training data for the content labeling predictive model using the graph data; train, using a training dataset including the training data generated by the at least one classifier network, the content labeling predictive model; identify, using the trained content labeling predictive model, a plurality of second links among the plurality of identified content assets and the plurality of identified labels; and annotate the plurality of identified content assets based on the plurality of identified second links. ||
7. A system comprising: a computing platform including a hardware processor and a system memory storing a software code and a content labeling predictive model; the hardware processor configured to execute the software code to: scan a database to identify a plurality of content assets stored in the database; parse metadata stored in the database to identify a plurality of labels associated with the plurality of identified content assets; generate graph data by creating a plurality of first links linking each one of the plurality of identified content assets to each of the plurality of identified labels associated with a corresponding one of the plurality of identified content assets; cluster the graph data in an unsupervised learning process; route the graph data to at least one classifier network based on the clustering; use the at least one classifier network to generate training data for the content labeling predictive model using the graph data; and train, using a training dataset including the training data generated by the at least one classifier network, the content labeling predictive model. ||
8. A method comprising: scanning a database to identify a plurality of content assets stored in the database; parsing metadata stored in the database to identify a plurality of labels associated with the plurality of identified content assets; generating a graph by creating a plurality of first links linking each one of the plurality of identified content assets to each of the plurality of identified labels associated with a corresponding one of the plurality of identified content assets; training, based on the graph, a content labeling predictive model; identifying, using the trained content labeling predictive model, a plurality of second links among the plurality of identified content assets and the plurality of identified labels; and annotating the plurality of identified content assets based on the plurality of identified second links. ||
12. A method comprising: scanning a database to identify a plurality of content assets stored in the database; parsing metadata stored in the database to identify a plurality of labels associated with the plurality of identified content assets; generating a graph by creating a plurality of first links linking each one of the plurality of identified content assets to each of the plurality of identified labels associated with a corresponding one of the plurality of identified content assets; clustering nodes of the graph in an unsupervised learning process to generate clusters of nodes; labeling the clusters of nodes using one or more classification networks to generate cluster labels; updating the graph by linking the clusters of nodes to the cluster labels; and training, based on the graph, a content labeling predictive model. ||
13. A system comprising: a computing platform including a hardware processor and a system memory storing a search engine and a graph of a plurality of content assets; the hardware processor configured to execute the search engine to: receive a search term as an input from a user of the system; traverse the graph using the search term to discover a content asset corresponding to the search term, wherein the discovered content asset is one of a plurality of content assets annotated using a content labeling predictive model; and output a search result identifying the discovered content; wherein the plurality of content assets are annotated by: scanning a database to identify the plurality of content assets stored in the database; parsing metadata stored in the database to identify a plurality of labels associated with the plurality of identified content assets; generating graph data by creating a plurality of first links linking each one of the plurality of identified content assets to each of the plurality of identified labels associated with a corresponding one of the plurality of identified content assets; training, based on the graph data, the content labeling predictive model; identifying, using the trained content labeling predictive model, a plurality of second links among the plurality of identified content assets and the plurality of identified labels; and annotating the plurality of identified content assets based on the plurality of identified second links.