ConceptioArchiveGoogle Patents
Google Patentsopen access

Using large language models to update data in mapping systems and applications — Nvidia Corporation (US20240419902A1)

Nvidia Corporation · Google Patents
Google Patents · Patents · License: Open Access
Open Source ↗
patent, google patents, intellectual property, US20240419902A1, Nvidia Corporation, Shuang Wu, en, 2024

ABSTRACT

Abstract

Approaches presented herein provide for the identification of differences between local map data, for a region of a physical environment, and observation or perception data generated by one or more machines or other such sources. In at least one embodiment, sensors on an ego machine can capture sensor data for a region in which the ego machine is located, and a language model on the ego machine can compare this sensor data, or perception data generated using the sensor data, against the local map data. The language model can generate a tokenized description of identified differences, in a domain-specific language. The tokenized description can be transmitted to a map management service that can compare these differences against differences identified by other machines, for example, to determine whether to update and redistribute at least a portion of the map data.

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63/521,627, filed Jun. 16, 2023, entitled “USING LANGUAGE MODELS FOR MAPPING IN AUTONOMOUS SYSTEMS AND APPLICATIONS,” the full disclosure of which is hereby incorporated in its entirety for all purposes.

BACKGROUND

There are various operations—such as may relate to autonomous or semi-autonomous navigation and robotic simulation—where it can be desirable to generate or reconstruct a realistic digital and/or virtual environment that complies with real-world rules and constraints. As an example, maps—such as high definition (HD) maps—are widely relied upon for semi-autonomous and autonomous operations. Autonomous and semi-autonomous vehicles and machines may rely on these maps, as well as real time sensor data, for navigation, localization, path or route planning, and/or other operations. In order to ensure that the maps are accurate and updated to account for any changes, map management systems can collect data captured by sensors of various vehicles driving along various routes and can analyze that data to attempt to determine whether a change to the map data might be warranted. Such a process is typically time consuming and complicated, as different vehicles can provide data in different formats, of different types, and/or with different accuracies, precisions, or confidence levels. This can result in the map data taking a sufficiently long time to update, which may be undesirable for operations such as autonomous vehicle navigation. Further, such an approach can require the collection and processing of a significant amount of high-precision data, which can be expensive in terms of computing and network resources.

BRIEF DESCRIPTION OF THE DRAWINGS

Various embodiments in accordance with the present disclosure will be described with reference to the drawings, in which:

FIG. 1 illustrates example data produced at stages of an environment reconstruction process, according to at least one embodiment;

FIGS. 2 A and 2 B illustrates an example environment reconstruction pipeline, along with an example tokenized description generated using such a pipeline, according to at least one embodiment;

FIG. 3 A illustrates an example pipeline for generating a tokenized description of differences identified between sensor data and map data for at least a portion of an environment, and determining whether to update the map data based in part on those differences, according to at least one embodiment;

FIG. 3 B illustrates an example set of components of a machine for identifying differences between local map and perception data, according to at least one embodiment;

FIG. 3 C illustrates an example set of components of a machine for identifying differences between local map and sensor data using a trained language model, according to at least one embodiment;

FIG. 3 D illustrates an example network-based system to determine whether to update map data based in part upon difference information received from multiple machines, according to at least one embodiment;

FIG. 4 A illustrates an example process for generating tokenized descriptions for differences identified between local map data and perception data, according to at least one embodiment;

FIG. 4 B illustrates an example process for generating tokenized descriptions for differences identified between local map data and a set of observations using a trained language model, according to at least one embodiment;

FIG. 4 C illustrates an example process to determine whether to update map data based in part upon differences identified by multiple machines or other such sources, according to at least one embodiment;

FIG. 5 A illustrates an example map graph, according to at least one embodiment;

FIG. 5 B illustrates an example landmark analysis system, according to at least one embodiment;

FIG. 5 C illustrates an example tokenized text string, according to at least one embodiment;

FIG. 5 D illustrates an example lane graph, according to at least one embodiment;

FIG. 5 E illustrates an example architecture for determining an output state, according to at least one embodiment;

FIG. 5 F illustrates an example image of an intersection in an example map, according to at least one embodiment;

FIG. 5 G illustrates an example process for generating a text string representation of an environment, according to at least one embodiment;

FIG. 5 H illustrates an example process for generating a tokenized text string representation of a physical environment, according to at least one embodiment

FIG. 6 illustrates components of a distributed system that can be used to update map data based in part on a tokenized description generated for an environment, according to at least one embodiment;

FIG. 7 A illustrates inference and/or training logic, according to at least one embodiment;

FIG. 7 B illustrates inference and/or training logic, according to at least one embodiment;

FIG. 8 illustrates an example data center system, according to at least one embodiment;

FIG. 9 illustrates a computer system, according to at least one embodiment;

FIG. 10 illustrates a computer system, according to at least one embodiment;

FIG. 11 illustrates at least portions of a graphics processor, according to one or more embodiments;

FIG. 12 illustrates at least portions of a graphics processor, according to one or more embodiments;

FIG. 13 is an example data flow diagram for an advanced computing pipeline, in accordance with at least one embodiment;

FIG. 14 is a system diagram for an example system for training, adapting, instantiating and deploying machine learning models in an advanced computing pipeline, in accordance with at least one embodiment;

FIGS. 15 A and 15 B illustrate a data flow diagram for a process to train a machine learning model, as well as client-server architecture to enhance annotation tools with pre-trained annotation models, in accordance with at least one embodiment;

FIG. 16 A illustrates an example of an autonomous vehicle, according to at least one embodiment;

FIG. 16 B illustrates an example of camera locations and fields of view for the autonomous vehicle of FIG. 16 A , according to at least one embodiment;

FIG. 16 C is a block diagram illustrating an example system architecture for the autonomous vehicle of FIG. 16 A , according to at least one embodiment; and

FIG. 16 D is a diagram illustrating a system for communication between cloud-based server(s) and the autonomous vehicle of FIG. 16 A , according to at least one embodiment.

DETAILED DESCRIPTION

In the following description, various embodiments will be described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will also be apparent to one skilled in the art that the embodiments may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified in order not to obscure the embodiment being described.

The systems and methods described herein may be used by, without limitation, non-autonomous vehicles or machines, semi-autonomous or autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS), one or more in-vehicle infotainment systems, one or more emergency vehicle detection systems), piloted and un-piloted robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying vessels, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, trains, underwater craft, remotely operated vehicles such as drones, and/or other vehicle types. Further, the systems and methods described herein may be used for a variety of purposes, by way of example and without limitation, for machine control, machine locomotion, machine driving, synthetic data generation, generative Al, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environment simulation, data center processing, conversational Al, light transport simulation (e.g., ray-tracing, path tracing, etc.), collaborative content creation for 3D assets, generative Al, cloud computing, and/or any other suitable applications.

Disclosed embodiments may be comprised in a variety of different systems such as automotive systems (e.g., an in-vehicle infotainment system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aerial systems, medial systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational Al operations, systems implementing one or more language models—such as large language models (LLMs), systems for performing generative Al operations (e.g., using one or more language models), systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems implemented at least partially using cloud computing resources, and/or other types of systems.

Approaches in accordance with various illustrative embodiments provide for the generation of tokenized representations of one or more regions or domains within a physical environment. In particular, various embodiments can use a large language model (LLM), or other generative artificial intelligence (AI)-based approach, to generate a tokenized description (or other text-based representation) of a region or domain. A language model can be trained to represent a region based on not only low-level primitives determinable from captured sensor data, for example, but also aspects such as the semantics, topology, and geometry related to those primitives, as well as the relationships between objects determinable using those primitives. These tokenized representations can correspond to, or be used to generate, feature vectors, embeddings, or points in a latent space, among other such options, representative of the respective regions or domains. The feature vectors in aggregate can be used to represent an entire environment or set of regions.

In at least one embodiment, a language model can be used to identify differences between the local map data, for a region of a physical environment, and observation or perception data obtained for the region, such as may be obtained or generated in real time by an ego machine (such as an ego vehicle). A set of observations can be obtained for an environment, where those observations may correspond to sensor data captured by one or more sensors of similar or different types. In at least one embodiment, the observations can be analyzed by a perception module to generate a set of perception data, where the perception data may relate to objects identified in the environment, as well as determined or inferred aspects of those objects. A current and/or reference location in the environment can also be determined, which can be used to identify local map data relevant to that location. The localized mapping data can be analyzed together with the observations (e.g., sensor data or perception data) so that corresponding objects or features in the map data and observations are identified and correlated. A trained language model can analyze the objects and/or features (e.g., embeddings or feature vectors) in the map data and observation data to attempt to identify differences that may warrant updates to the map data. This may include identifying any or all such differences, or differences that satisfy at least one selection criterion, among other such options. A language model can attempt to generate a single, fused representation of the environment based at least in part on the map data and the perception data, where that representation or another tokenized representation can include tokens that include information about the identified d

CROSS-REFERENCE TO RELATED APPLICATIONS

This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63/521,627, filed Jun. 16, 2023, entitled “USING LANGUAGE MODELS FOR MAPPING IN AUTONOMOUS SYSTEMS AND APPLICATIONS,” the full disclosure of which is hereby incorporated in its entirety for all purposes.

BACKGROUND

There are various operations—such as may relate to autonomous or semi-autonomous navigation and robotic simulation—where it can be desirable to generate or reconstruct a realistic digital and/or virtual environment that complies with real-world rules and constraints. As an example, maps—such as high definition (HD) maps—are widely relied upon for semi-autonomous and autonomous operations. Autonomous and semi-autonomous vehicles and machines may rely on these maps, as well as real time sensor data, for navigation, localization, path or route planning, and/or other operations. In order to ensure that the maps are accurate and updated to account for any changes, map management systems can collect data captured by sensors of various vehicles driving along various routes and can analyze that data to attempt to determine whether a change to the map data might be warranted. Such a process is typically time consuming and complicated, as different vehicles can provide data in different formats, of different types, and/or with different accuracies, precisions, or confidence levels. This can result in the map data taking a sufficiently long time to update, which may be undesirable for operations such as autonomous vehicle navigation. Further, such an approach can require the collection and processing of a significant amount of high-precision data, which can be expensive in terms of computing and network resources.

BRIEF DESCRIPTION OF THE DRAWINGS

Various embodiments in accordance with the present disclosure will be described with reference to the drawings, in which:

FIG. 1 illustrates example data produced at stages of an environment reconstruction process, according to at least one embodiment;

FIGS. 2 A and 2 B illustrates an example environment reconstruction pipeline, along with an example tokenized description generated using such a pipeline, according to at least one embodiment;

FIG. 3 A illustrates an example pipeline for generating a tokenized description of differences identified between sensor data and map data for at least a portion of an environment, and determining whether to update the map data based in part on those differences, according to at least one embodiment;

FIG. 3 B illustrates an example set of components of a machine for identifying differences between local map and perception data, according to at least one embodiment;

FIG. 3 C illustrates an example set of components of a machine for identifying differences between local map and sensor data using a trained language model, according to at least one embodiment;

FIG. 3 D illustrates an example network-based system to determine whether to update map data based in part upon difference information received from multiple machines, according to at least one embodiment;

FIG. 4 A illustrates an example process for generating tokenized descriptions for differences identified between local map data and perception data, according to at least one embodiment;

FIG. 4 B illustrates an example process for generating tokenized descriptions for differences identified between local map data and a set of observations using a trained language model, according to at least one embodiment;

FIG. 4 C illustrates an example process to determine whether to update map data based in part upon differences identified by multiple machines or other such sources, according to at least one embodiment;

FIG. 5 A illustrates an example map graph, according to at least one embodiment;

FIG. 5 B illustrates an example landmark analysis system, according to at least one embodiment;

FIG. 5 C illustrates an example tokenized text string, according to at least one embodiment;

FIG. 5 D illustrates an example lane graph, according to at least one embodiment;

FIG. 5 E illustrates an example architecture for determining an output state, according to at least one embodiment;

FIG. 5 F illustrates an example image of an intersection in an example map, according to at least one embodiment;

FIG. 5 G illustrates an example process for generating a text string representation of an environment, according to at least one embodiment;

FIG. 5 H illustrates an example process for generating a tokenized text string representation of a physical environment, according to at least one embodiment

FIG. 6 illustrates components of a distributed system that can be used to update map data based in part on a tokenized description generated for an environment, according to at least one embodiment;

FIG. 7 A illustrates inference and/or training logic, according to at least one embodiment;

FIG. 7 B illustrates inference and/or training logic, according to at least one embodiment;

FIG. 8 illustrates an example data center system, according to at least one embodiment;

FIG. 9 illustrates a computer system, according to at least one embodiment;

FIG. 10 illustrates a computer system, according to at least one embodiment;

FIG. 11 illustrates at least portions of a graphics processor, according to one or more embodiments;

FIG. 12 illustrates at least portions of a graphics processor, according to one or more embodiments;

FIG. 13 is an example data flow diagram for an advanced computing pipeline, in accordance with at least one embodiment;

FIG. 14 is a system diagram for an example system for training, adapting, instantiating and deploying machine learning models in an advanced computing pipeline, in accordance with at least one embodiment;

FIGS. 15 A and 15 B illustrate a data flow diagram for a process to train a machine learning model, as well as client-server architecture to enhance annotation tools with pre-trained annotation models, in accordance with at least one embodiment;

FIG. 16 A illustrates an example of an autonomous vehicle, according to at least one embodiment;

FIG. 16 B illustrates an example of camera locations and fields of view for the autonomous vehicle of FIG. 16 A , according to at least one embodiment;

FIG. 16 C is a block diagram illustrating an example system architecture for the autonomous vehicle of FIG. 16 A , according to at least one embodiment; and

FIG. 16 D is a diagram illustrating a system for communication between cloud-based server(s) and the autonomous vehicle of FIG. 16 A , according to at least one embodiment.

DETAILED DESCRIPTION

In the following description, various embodiments will be described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will also be apparent to one skilled in the art that the embodiments may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified in order not to obscure the embodiment being described.

The systems and methods described herein may be used by, without limitation, non-autonomous vehicles or machines, semi-autonomous or autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS), one or more in-vehicle infotainment systems, one or more emergency vehicle detection systems), piloted and un-piloted robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying vessels, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, trains, underwater craft, remotely operated vehicles such as drones, and/or other vehicle types. Further, the systems and methods described herein may be used for a variety of purposes, by way of example and without limitation, for machine control, machine locomotion, machine driving, synthetic data generation, generative Al, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environment simulation, data center processing, conversational Al, light transport simulation (e.g., ray-tracing, path tracing, etc.), collaborative content creation for 3D assets, generative Al, cloud computing, and/or any other suitable applications.

Disclosed embodiments may be comprised in a variety of different systems such as automotive systems (e.g., an in-vehicle infotainment system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aerial systems, medial systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational Al operations, systems implementing one or more language models—such as large language models (LLMs), systems for performing generative Al operations (e.g., using one or more language models), systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems implemented at least partially using cloud computing resources, and/or other types of systems.

Approaches in accordance with various illustrative embodiments provide for the generation of tokenized representations of one or more regions or domains within a physical environment. In particular, various embodiments can use a large language model (LLM), or other generative artificial intelligence (AI)-based approach, to generate a tokenized description (or other text-based representation) of a region or domain. A language model can be trained to represent a region based on not only low-level primitives determinable from captured sensor data, for example, but also aspects such as the semantics, topology, and geometry related to those primitives, as well as the relationships between objects determinable using those primitives. These tokenized representations can correspond to, or be used to generate, feature vectors, embeddings, or points in a latent space, among other such options, representative of the respective regions or domains. The feature vectors in aggregate can be used to represent an entire environment or set of regions.

In at least one embodiment, a language model can be used to identify differences between the local map data, for a region of a physical environment, and observation or perception data obtained for the region, such as may be obtained or generated in real time by an ego machine (such as an ego vehicle). A set of observations can be obtained for an environment, where those observations may correspond to sensor data captured by one or more sensors of similar or different types. In at least one embodiment, the observations can be analyzed by a perception module to generate a set of perception data, where the perception data may relate to objects identified in the environment, as well as determined or inferred aspects of those objects. A current and/or reference location in the environment can also be determined, which can be used to identify local map data relevant to that location. The localized mapping data can be analyzed together with the observations (e.g., sensor data or perception data) so that corresponding objects or features in the map data and observations are identified and correlated. A trained language model can analyze the objects and/or features (e.g., embeddings or feature vectors) in the map data and observation data to attempt to identify differences that may warrant updates to the map data. This may include identifying any or all such differences, or differences that satisfy at least one selection criterion, among other such options. A language model can attempt to generate a single, fused representation of the environment based at least in part on the map data and the perception data, where that representation or another tokenized representation can include tokens that include information about the identified differences. The model can use its domain-specific learning, as well as semantic, relationship, topology, geometry, and other information provided with, or determinable from, the map and perception data, to attempt to infer a consistent representation of the environment and then identify differences or inconsistencies in that representation. The tokenized description of at least the identified differences can correspond to a string of text-based tokens written in a domain-specific language. The tokenized description can be a compact and discrete representation of the environment, which is lightweight enough to be processed in real time or near real time but robust enough to include the necessary information for making decisions relevant to the target operation(s). The compactness of the tokenized description can be improved in at least one embodiment by training the language model to generate tokens only differences for those objects, or relevant aspects of those objects, that are determined to be important for a given task, operation, or domain. The ability to infer a consistent representation from map and perception data allows for useful tokenized descriptions to be generated even where the map and/or perception data may be unavailable, incomplete, inaccurate, or otherwise unreliable. In at least one embodiment, tokenized differences identified by multiple machines (e.g., vehicles or other such sources) can be transmitted to a map management service or other such recipient, which may be hosted across at least one network, as may use a set of cloud-based resources. The map management service can attempt to aggregate and correlate this difference data, and attempt to come to a consensus, with at least a minimum level of confidence, as to whether one or more updates should be performed with respect to the map data. If an update is to be performed, that update can be performed and then the updated map data, or information about the updates to the map, can be propagated or otherwise made available to at least the relevant machines or other such recipients.

The updates or updated map data may also be provided using at least one tokenized representation, which may be in the same domain-specific language relevant to a specific type of operation.

In at least one embodiment, a language model can be used to generate representations of operational design domains (ODDs)—such as intersections—using a tokenized representation, written in a language such as Road Topology Language (RTL). An embedding can be generated for each such ODD, allowing each ODD to be represented by, for example, a point in an n-dimensional latent space. Similar ODDs, such as similar intersections, will have similar embeddings. Generated embeddings can capture semantic, geometric, topological, and/or other information for the ODDs.

In at least one embodiment such a language model can generate a representation of an environment that complies with real world rules and constructs, and that accounts for omissions or errors in the input data to be used to generate the representation. Objects in an environment can be represented using individual tokens or token sequences, optionally with token descriptors providing semantics and other information related to these tokens. A text-based representation can be a one-dimensional string of these tokens and token descriptors, which can encapsulate the important spatial information and semantics of an environment. An advantage of such a text-based description is that it can be discrete and compact, allowing for quick processing, search, updating, and other such operations. A generated text-based representation of an environment can be used to generate a number of other types of representations useful for various operations or tasks, such as may include birds-eye view maps, high definition (HD) maps, or 3D virtual environments, among other such options.

In at least one embodiment, generative Al can be used to provide a semantic understanding of an environment based at least in part on sensor data captured for an environment. This sensor data can be processed and fed to a trained generative Al model (such as a large language model or “LLM”), for example, which can output a textual description of an environment in a structured textual format, such as in a Road Topology Language (RTL). A text string in RTL can provide a tokenized representation of a map or graph of an environment. The generative Al can be trained in such a way as to be able to fill in gaps or correct errors in the sensor data based on a semantic understanding of the objects or elements in the environment. The language model can receive input including semantic, location, and/or geometric information determined for an environment, such as by processing sensor data (e.g., image or LIDAR data) captured for an environment, and can update the textual representation of a scene as the environment changes due to movement or other such occurrences. The language model may also take other inputs as well, such as prior maps or context information (indicating things like weather, time of day, season, urban/rural region, geographic location, etc.). The input data can be represented by embeddings, feature vectors, or points in a latent space, which allows for relatively simple searching for similar environments. In this way, quick determinations of actions to be taken in an environment can be made by determining which actions were taken in similar environments, particularly when there may be insufficient data available for a current environment or situation to make a high confidence decision as to an action to be taken. The ability to determine what others have done in similar environments can help a system to function in a way similar to how a human uses “intuition” in a given situation even when there may be data missing, such as where snow may have obscured the lines along a road but the human can infer where to drive based on other information available in the environment. Such an approach can be used for a wide variety of geospatial information processing and autonomous driving tasks (such as map building, map editing, map-based navigation, planning and driving) by representing those tasks as document manipulation tasks. A generative Al once trained can also be used to generate realistic simulation environments that comply with real world rules, such as may be useful for testing autonomous vehicles or robots, or other such machines. A machine as used herein can include any appropriate physical (or at least partially virtual) device, system, or component that is able to process data to perform one or more actions, such as may include one or more physical actions in a real world environment. Such an approach can also be used to correct or update noisy or partial environment graphs or maps. The generative Al model might take the sensor data directly as input or might receive input that is generated from the sensor data in one or more stages of a pipeline, such as stages to extract features and generate embeddings of those features in a latent space that can be provided as input to the generative model.

Approaches in accordance with various illustrative embodiments can provide for the use of language models for mapping in autonomous or semi-autonomous systems and applications. Systems and methods are disclosed that use one or more language models (e.g., LLMs) to perform various mapping operations—such as map building, map editing, map-based navigation, routing, planning, and perception, error checking, data cleaning, and data validation, among others. For example, a deep learning model—such as an LLM—may encapsulate domain knowledge about how road networks and/or objects are structured. By training an LLM to predict structure and attributes of a graph described in a domain specific language (DSL)—such as RTL—the LLM learns to establish correct relationships among objects on the road. The RTL may express road, object, and/or other map-related information (e.g., by modeling relationships among lane elements and other map features) using language, such that the LLM learns to interpret the RTL—in addition to natural or conversational language—to generate outputs. An automated process may be implemented to convert existing map information to the RTL, and to convert outputs of the LLM from RTL to a suitable map format (e.g., a format for an HD map deployed in a production vehicle). Such an LLM may be used to solve various challenging problems related to mapping—such as identifying or correcting mistakes or gaps in maps, creating maps from a photo or video stream of road data, creating maps from aerial or satellite images, and/or creating maps from text descriptions. Once created, the maps can be used for various tasks, such as for autonomous vehicles (AV) or autonomous systems, semi-autonomous vehicles or systems (e.g., for advanced driver assistance systems (ADAS)), simulation systems (e.g., for developing or testing/validating AV/ADAS algorithms or for creating training data for AV/ADAS perception), and/or the like.

Variations of this and other such functionality can be used as well within the scope of the various embodiments as would be apparent to one of ordinary skill in the art in light of the teachings and suggestions contained herein.

FIG. 1 illustrates an example data processing flow that can be implemented in an environment representation and/or reconstruction system in accordance with at least one embodiment. In this example, sensor data 104 (or other raw data captured or representative of an environment) is obtained with respect to a specific environment 102 . The environment can be any appropriate physical environment, such as an indoor or outdoor environment that may include any number of different types of objects or elements. The sensor data can include data captured or obtained using any of a number of different types of sensors, as may include cameras, LIDAR systems, radars, sonic sensors, distance sensors, and the like. Additional data may be obtained that relates to the environment 102 as well in various embodiments, as may relate to basic map data, contextual data, motion data, or other such data, which may also be obtained for virtual, augmented, or enhanced environments. In this example, the sensor data 104 (and any other available and useful data) can be used to generate an initial representation 106 of the environment 102 . In at least one embodiment, this may include a point cloud representation of the environment 102 generated by analyzing and aggregating the sensor data 104 that may have been captured by multiple sensors in order to generate a single, n-dimensional (e.g., 2D, 3D, or 4D) representation of the environment. Other initial representations can be generated as well, as may depend at least in part upon the type of sensor data provided. If image data is provided, the image data may be analyzed to attempt to determine feature and depth information, which can be combined from multiple images from different viewpoints to attempt to generate at least a 3D representation of the environment 102 , or at least objects and shapes within that environment.

This initial representation 106 of the environment 102 can be analyzed to attempt to determine specific aspects 108 of the environment. For example, a point cloud can be analyzed to attempt to determine the categories (or types) of objects represented in the environment, as may relate to roadways, traffic signs, sidewalks, buildings, and the like. The representation can also be analyzed to attempt to determine the locations of these objects in the environment, as may be defined using a set of 3D coordinates relative to a determined origin location. The initial representation 106 can also be analyzed to attempt to determine various relationships between these objects, such as where a crosswalk crosses specific lanes or where a stop sign is associated with a specific lane and indicates an expected behavior. Once these determined aspects 108 are obtained, these aspects can be used to generate an object-based representation 110 of the environment 102 . Various other types of representations can be generated as well within the scope of various embodiments. As illustrated, the object-based representation 110 will not be a comprehensive description of the environment 102 in this example, but will instead focus on the types of objects or features of the environment that are potentially relevant to a particular task. For autonomous driving, for example, the object-based representation may include objects such as road lanes, crosswalks, intersections, and the like, but may not include objects that may not be directly relevant to driving, as may include buildings, billboards, mailboxes, and other such objects, except to the extent those objects may be relevant to a specific operation or task. In this example, the object-based representation 110 also does not include vehicles, pedestrians, or other movable objects that will only be in specific locations in the environment 102 at specific times, but any or all of these and other such objects could be included in the representation as well within the scope of various embodiments.

From this object-based representation, an object graph 112 can be generated that provides a different representation of the environment 102 . An advantage of the object graph 112 is that it is relatively lightweight, and can be used to compactly describe aspects of the environment 102 that are important for a particular task or operation. For example, such an object graph 112 could be provided to a map generator in order to generate an HD map (or other such map or representation) that can be provided to an autonomous vehicle to make navigation decisions. Such an object graph 112 can also be provided as input to an environment generator that can generate a realistic 3D virtual environment that can be used for tasks such as robotic simulation or digital world recreation. A large number of object graphs can be stored to represent a number of different environments, which can require significantly less memory or storage capacity than sensor data, such as a large number of high resolution images. Such object graphs can also be analyzed quickly to allow for real-time operations, such as autonomous navigation or control.

A challenge with existing approaches to generate such representations is that there is a limited ability to perform automated geospatial information processing, particularly using an algorithm framework that is sufficiently generic to support a wide variety of use cases. Existing solutions typically have task-specific designs that cannot easily adapt to new task requirements, contain built-in assumptions that might not always hold in real-world situations, and do not make effective use of available data and human input. Existing approaches are also limited in their ability to learn from large amounts of diverse data that can be relevant to these different tasks or use cases. Many existing solutions depend heavily on domain expertise and manually-designed logic or rules in various steps of the processing pipeline. These attempted solutions are difficult to accurately complete and improve, and require manual effort to moderate the results and make them correct. Improvements in these systems are costly and generally offer smaller and smaller performance gains for the effort spent.

Approaches in accordance with at least one embodiment can provide a versatile approach to processing information about such an environment 102 , as may include geospatial and semantic information. In at least one embodiment, a deep learning model can be used that encapsulates domain-specific (or agnostic) knowledge about how objects in an environment are structured and related. An example deep learning model is a large language model (LLM) that can be trained to generate a textual description of an environment that retains semantic understanding of an environment in addition to providing information about the categories and locations of objects in the environment. In at least one embodiment, an LLM can generate a tokenized text string as a representation of an environment, where objects in the environment are represented as tokens in the string. There can also be a set of token descriptors in the string, and associated with specific tokens, that provide semantic and/or relationship information with respect to the various tokens of the string. In addition to generating a compact yet thorough representation of an environment, for example, an advantage of using a model such as an LLM is that the LLM can fill in gaps in the sensor data or otherwise make corrections where needed to provide a more accurate representation of the environment. For example, training an LLM to predict the next token in the text string (corresponding to a next object in an object graph, for example) can help the LLM to learn to establish correct relationships between objects in the environment. This can include, for example, identifying or correcting mistakes or gaps in environment representations, creating environment representations (e.g., maps or object graphs) from a photo or video stream of environment data, creating environment representations from aerial or satellite images, and creating environment descriptions from textual descriptions, among other such tasks.

In at least one embodiment, a language model-based approach can be used that can allow model training on large-scale existing environment representations, such as maps, making data-driven performance improvements easier and more scalable with respect to domain expertise. A training approach can be used that can specifically teach the LLM to identify the next token in the graph. In at least one embodiment, an LLM can generate a deep underlying representation of how objects and/or networks in the environment are connected or related, as well as a model of the graph data already presented as input to the LLM. In at least one embodiment, various tasks in geospatial information processing can be unified under a shared formulation, such that the same algorithmic models can be re-used without extra engineering effort. Processing efficiency can be further improved through replacing manual labor with machine learning model-based automation. A large language model can be trained on vast amounts of environment data so that it can automate various tasks such as missing element detection, inaccurate element correction, and inference of relationships among elements, among other such tasks. Each of these can be achieved without heavily depending on human expertise to explicitly design for, and can be improved continuously with additional training data. Such a model can leverage existing environment (e.g., map) data without requiring additional data curation and labeling cost. The model can be trained in a task-agnostic way so that the model can be extended to other use cases without significant additional effort. These representations can include, or be used to generate, high quality maps useful for tasks such as those related to an advanced driver assistance system (ADAS), autonomous vehicle (AV), unmanned aerial vehicle (UAV) or simulation system, such as may be useful for developing or testing/validating AV/ADAS/UAV algorithms or creating training data for AV/ADAS/UAV perception.

Approaches in accordance with at least one embodiment attempt to improve, optimize, or at least control the way in which an environment is perceived. In various existing systems, perception of an environment is relatively primitive and based around rules for detected objects. For example, an existing system might analyze a captured image to identify the location of roadway lanes and lane markers, but do not have any concept of what the lines on the roadways mean, or how those lines relate to nearby road signs or traffic lights. An existing system might recognize the objects and use the locations of those objects to generate a map reflecting those objects. The system might attempt to determine relationships and apply rules to these objects to ensure the placement makes sense and determine any relationships, but this is typically done during post-processing when most other data has already been discarded. Applying rules based on detected objects means that it can be difficult to detect gaps, errors, or omissions that might otherwise be detected if the relationships and semantic meanings of various objects in a seen were known and used in the process of generating the representation of the environment.

Further, a rules-based approach is harder to scale in many instances.

An approach in accordance with at least one embodiment can obtain and apply such knowledge earlier in the process. As mentioned, a large language model can take input relating to the semantics, location, and relationship between various objects in an environment, and can use this information to determine based on its learning how to generate a realistic environment representation based on this input that can make up for the fact that the input data may be somewhat incomplete or erroneous. By representing the environment through text, a language model can apply its learnings to determine how to structure the representation to ensure realism and completeness, and fill in gaps in the input data based on what it has learned from similar situations. A language model has the advantage of taking text as input, rather than images or other large instances of sensor data, which can be processed relatively quickly during training. This allows a generative model to be trained using millions or even billions of such documents, with self-supervision, which provides for better understanding of behavior and relationships, as well as which behavior and/or relationships apply to a given environment or situation. By converting an object- or feature-based representation into a language representation, for example, this text-based representation can be used to train a language model to understand the various correlations between categories of objects and their relative locations, including ways that may be difficult to enumerate comprehensively. Attempting to capture all the relevant real-world correlations, relationships, and other semantic aspects would be extremely difficult to do using only explicit rules as would be required for various existing systems.

In the example of FIG. 1 , a language model could take as input an object-based representation 110 and generate what is essentially a tokenized text string representation of the object graph 112 . In other embodiments, the language model might be able to take other inputs that would allow for at least some steps in this generation pipeline to be eliminated as separate steps performed by separate processes or components. For example, an LLM could be trained to take in a set of determined aspects 108 (e.g., semantics, topology, or geometry information for an environment or objects in that environment) in text format and generate a tokenized text string representative of the object graph 112 without ever having to generate an object-based representation. Similarly, in some embodiments an LLM can take as input the internal representation 106 , or even the sensor data 104 , without the need for separate intermediate representations. For example, a model (as part of the LLM or a separate model) can analyze the sensor data 104 for the environment and encode features of the sensor data into a latent space (or other embedding). The LLM can then take a feature vector as input that is a function of these individual latent space encodings, and can directly generate the tokenized text string representation of the environment. The features extracted can include semantic, relationship, and geometry features, among other such options. Encoding such features in a latent space can prevent this information from being discarded early in the generation process, and allow for more accurate representations or reconstructions to be generated.

In at least one embodiment, the tokenized text string can include a sequence of tokens, where each token represents an object in the environment. The string also can include a set of “token descriptors” that provide some semantic context or other useful information for a given token. The tokens can also be in a specific sequence, which not only can be useful in generating an object graph from the text string, but also allows semantic learning to be applied to the sequence of tokens as an LLM might typically do for the words of a sentence. A number of languages can be used to represent such an environment, as long as the language is able to provide the representation as a sequential notation of discrete tokens. In at least one embodiment, a custom language might be used that includes specific tokens and token descriptors that can accurately and compactly represent a specific type of environment. For example, a road topology language (RTL) might be used that includes terminology and syntax useful for representing map data for environments including roadways. A unified, sequential, tokenized text representation can be used to model an object graph, and an object graph can be quickly generated from such a sequential tokenized text string in a way that is consistently repeatable. A language model can be trained to understand and “speak” in at least one specific language, such as RTL. As a trained LLM will know how to manipulate or fill in a sentence in natural language, so can an LLM learn to fill in a text string in a structured representation language. The LLM can also infer relationships between objects based on its understanding of the language. The LLM can then generate a unified text representation of an environment that can include information that was not present or determinable from the input alone but that allows the environment to be more realistic and to comply with real world rules and/or constraints. These may include, for example, local traffic rules or ordinances, customs, and abilities of objects in the environment, among other such options. The language model can be trained to learn the semantics and syntax of the language, as well as the reasoning behind the semantics and syntax, including the physical concepts behind various object relationships. Instead of considering lane boundaries as lines in space, an LLM can consider the boundaries as associated with lanes of a roadway that come with various requirements, traffic rule or behaviors, and associated objects.

A language model trained to generate a representation using such a language can be used in at least one embodiment to describe the physical layout of an environment, such as may be useful for generating high quality maps. A model can generate text to describe other aspects of an environment as well, as may include characters, animals, vehicles, or other objects and elements that might move or change position or pose over time, and that might only be in an environment for a limited period of time. For example, a text string might be generated that provides a representation including a map view that illustrates where a vehicle can navigate, and also including representations of pedestrians, other vehicles, buildings, or other types of objects or entities that may be important for navigation or other such tasks. If a language model is able to generate a presentation that accurately describes aspects of the environment including nearby vehicles and pedestrians, for example, then navigation decisions may be able to be made using this representation without a separate need to identify such objects and provide that as additional input to a navigation or control system. An example perception map or representation can be generated that may include anything or everything in an environment that can be perceived using the available sensor data (or other such data) along with understanding of the physical rules or relationships for such an environment.

FIG. 2 A illustrates an example pipeline 200 that can be used to generate a text-based representation of an environment in accordance with at least one embodiment. Rather than requiring at least some amount of manual interaction, such an approach can automatically generate a representation from a variety of different types of input data. In this example, a capture device 202 can include, or be associated with, one or more sensors

204 , 206 that can capture or generate information about an environment 208 . The capture device can include any device, system, or component that is able to obtain sensor data from one or more sensors and either process that sensor data or transmit that sensor data for processing, as may include a desktop computer, a smart phone, a vehicle with data processing capability, or a robotic assembly, among other such options. The sensors can include any appropriate type of sensor that is able to capture or generate useful information about an environment, including sensors such as cameras, infrared (IR) sensors, ultrasonic sensors, depth sensors, LIDAR systems, radar systems, or other such sensors or data capture elements. The environment 208 can include an environment in which the capture device 202 is located, or that is within a capture distance of one or more sensors

204 , 206 .

In this example, the capture device 202 can provide the sensor data to be analyzed by a feature extraction module 210 . As mentioned, the feature extraction can be performed as part of a large language model 212 or by a separate model or algorithm, among other such options. In this example, the feature extraction module 210 can include an encoder that can extract features from the various instances of sensor data and encode those features as embeddings or points in a latent space 214 . The environment 208 in at least one embodiment can be represented by a set of embeddings or points in latent space, which may then be represented by one or more feature vectors corresponding to those individual embeddings. The latent space 214 may be an n-dimensional latent space, where each environment (or state of an environment) can correspond to a point (or vector) in the n-dimensional latent space.

In this example, at least one feature vector representing the point in the n-dimensional space can be provided as input to a large language model 212 . Various other types of embeddings or representations can be used as well within the scope of various embodiments. In at least one embodiment, each object in the environment can be represented by a token in a text string to be generated, as well as an embedding, feature vector, or point in an n-dimensional latent space, as discussed previously. Such a feature vector or embedding can specify not only the type of object, but can also represent various features of that object that can help to encode, for example, semantic, geographic, and/or topological information for that object.

The language model can use this input to generate a tokenized text string that is representative of the environment. In this example, the language model might receive other input as well that may help to generate a more accurate representation. For example, the language might receive a prior or partial map or environment rep

CLAIMS

Claims ( 20 )

What is claimed is:

1 . A method, comprising:

obtaining a set of observations corresponding to a region of a physical environment; identifying local map data corresponding to the region; generating, based at least on a trained language model processing data representative of the local map data and at least a subset of the set of observations, a tokenized description indicating one or more differences between the local map data and the set of observations; and determining, based at least on the tokenized description, whether one or more updates are to be performed with respect to the local map data based on the one or more differences.

2 . The method of claim 1 , wherein the set of observations includes at least one of sensor data, captured using one or more sensors in the region, or perception data generated using at least the sensor data.

3 . The method of claim 1 , wherein the tokenized description is compared with additional tokenized descriptions received that correspond to the region in order to determine, with at least a minimum level of confidence, whether to perform the one or more updates with respect to the local map data.

4 . The method of claim 1 , wherein the tokenized description includes one or more text-based tokens specific to the one or more differences, the one or more text-based tokens including at least one of a type of difference, a delta indicating an extent of a difference, or a confidence value in a difference determination.

5 . The method of claim 1 , wherein the set of observations are determined using an ego machine operating in, or proximate to, the region of the physical environment, and wherein the trained language model is located on the ego machine, the ego machine to transmit the tokenized description across at least one network to a system to determine whether to perform the one or more updates.

6 . The method of claim 1 , wherein potential differences are analyzed for at least two levels of granularity, starting at a higher level of granularity.

7 . The method of claim 1 , wherein the tokenized description further includes one or more recommended changes to the map data.

8 . The method of claim 1 , further comprising:

receiving information for one or more updates to the local map data; and storing the updated map data for use in at least one of future operation or future difference determinations.

9 . The method of claim 1 , wherein the tokenized description is written in a road topology language (RTL) or other domain specific language (DSL).

10 . The method of claim 1 , wherein the tokenized description is determined based at least on at least one of semantic, topological, geometric, kinematic, or relational information of features in the set of observations.

11 . A processor, including one or more logical units to:

generate a set of observations corresponding to a region of a physical environment; identify local map data corresponding to the region; and generate, based at least on a large language model (LLM) processing data corresponding to the local map data and at least a subset of the set of observations, a tokenized description indicating one or more differences identified between the local map data and the set of observations, wherein the tokenized description is used to determine whether to perform one or more updates to the local map data.

12 . The processor of claim 11 , wherein the set of observations includes at least one of sensor data, captured using one or more sensors in the region, or perception data generated using at least the sensor data.

13 . The processor of claim 11 , wherein the tokenized description is compared with additional tokenized descriptions received that correspond to the region in order to determine, with at least a minimum level of confidence, whether to perform the one or more updates with respect to the local map data.

14 . The processor of claim 11 , wherein the tokenized description includes one or more text-based tokens specific to the one or more differences, the one or more text-based tokens including at least one of a type of difference, a delta indicating an extent of a difference, or a confidence value in a difference determination.

15 . The processor of claim 11 , wherein the set of observations are determined on an ego machine operating in, or proximate to, the region of the physical environment, and wherein the LLM is located on the ego machine, the ego machine to transmit the tokenized description across at least one network to a system to determine whether to perform the one or more updates.

16 . The processor of claim 11 , wherein the processor is comprised in at least one of:

a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system for performing generative Al operations using a large language model (LLM); a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs); a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for performing generative operations using a language model (LM); a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources.

17 . A system comprising:

one or more processors to determine one or more updates to map data based at least on one or more differences between the map data and a set of observations for the region, the one or more differences being identified based at least on a language model processing the map data and data corresponding to the set of observations.

18 . The system of claim 17 , wherein the set of observations includes at least one of sensor data, captured using one or more sensors in the region, or perception data generated using at least the sensor data.

19 . The system of claim 17 , wherein the map data, the set of observations, and the one or more differences are represented in a domain specific language (DSL) corresponding to a mapping domain.

20 . The system of claim 17 , wherein the system comprises at least one of:

a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system for performing generative Al operations using a large language model (LLM); a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs); a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for performing generative operations using a language model (LM); a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources.

US18/417,105

2023-06-16

2024-01-19

Using large language models to update data in mapping systems and applications

Pending

US20240419902A1

( en )

Priority Applications (1)

Application Number

Priority Date

Filing Date

Title

US18/417,105

US20240419902A1

( en )

2023-06-16

2024-01-19

Using large language models to update data in mapping systems and applications

Applications Claiming Priority (2)

Application Number

Priority Date

Filing Date

Title

US202363521627P

2023-06-16

2023-06-16

US18/417,105

US20240419902A1

( en )

2023-06-16

2024-01-19

Using large language models to update data in mapping systems and applications

Publications (1)

Publication Number

Publication Date

US20240419902A1

true

US20240419902A1 ( en )

2024-12-19

Family

ID=93844623

Family Applications (7)

Application Number

Title

Priority Date

Filing Date

US18/472,941

Pending

US20240419903A1

( en )

2023-06-16

2023-09-22

Processing sensor data using language models in map generation systems and applications

US18/474,591

Pending

US20240419904A1

( en )

2023-06-16

2023-09-26

Using language models to verify map data in map generation systems and applications

US18/483,089

Pending

US20240419905A1

( en )

2023-06-16

2023-10-09

Training machine learning models using captured human reasoning

US18/500,426

Pending

US20240419906A1

( en )

2023-06-16

2023-11-02

Generating higher resolution map data using language models

US18/502,747

Pending

US20240419907A1

( en )

2023-06-16

2023-11-06

Using large language models for similarity determinations in content generation systems and applications

US18/409,018

Pending

US20240418515A1

( en )

2023-06-16

2024-01-10

Using a language model to localize and route plan for navigation systems and applications

US18/417,105

Pending

US20240419902A1

( en )

2023-06-16

2024-01-19

Using large language models to update data in mapping systems and applications

Family Applications Before (6)

Application Number

Title

Priority Date

Filing Date

US18/472,941

Pending

US20240419903A1

( en )

2023-06-16

2023-09-22

Processing sensor data using language models in map generation systems and applications

US18/474,591

Pending

US20240419904A1

( en )

2023-06-16

2023-09-26

Using language models to verify map data in map generation systems and applications

US18/483,089

Pending

US20240419905A1

( en )

2023-06-16

2023-10-09

Training machine learning models using captured human reasoning

US18/500,426

Pending

US20240419906A1

( en )

2023-06-16

2023-11-02

Generating higher resolution map data using language models

US18/502,747

Pending

US20240419907A1

( en )

2023-06-16

2023-11-06

Using large language models for similarity determinations in content generation systems and applications

US18/409,018

Pending

US20240418515A1

( en )

2023-06-16

2024-01-10

Using a language model to localize and route plan for navigation systems and applications

Country Status (1)

Country

Link

US

( 7 )

US20240419903A1

( en )

Cited By (4)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20240425050A1

( en )

*

2023-06-23

2024-12-26

GM Global Technology Operations LLC

Probabilistic driving behavior modeling system for a vehicle

US20250094908A1

( en )

*

2023-09-20

2025-03-20

International Business Machines Corporation

Evaluating impact of data feature deletion on associated policies

US12435989B2

( en )

*

2021-07-30

2025-10-07

Beijing Tusen Zhitu Technology Co., Ltd.

Semantic map and point cloud map construction method, apparatus and storage medium

US20250341405A1

( en )

*

2024-05-01

2025-11-06

GM Global Technology Operations LLC

System and method of actor-based map attribute generation

Families Citing this family (30)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

IL265818A

( en )

*

2019-04-02

2020-10-28

Ception Tech Ltd

System and method for determining the position and orientation of an object in space

US20250141755A1

( en )

*

2020-05-28

2025-05-01

Leatron Llc

Methods and systems for object-aware fuzzy processing based on analogies

US20240419903A1

( en )

*

2023-06-16

2024-12-19

Nvidia Corporation

Processing sensor data using language models in map generation systems and applications

WO2025014661A1

( en )

*

2023-07-07

2025-01-16

Palantir Technologies Inc.

Systems and methods for user interactive georegistration

US20250029287A1

( en )

*

2023-07-19

2025-01-23

Capital One Services, Llc

Generating images based on generated clusters

US20250036871A1

( en )

*

2023-07-25

2025-01-30

Dashlane SAS

Systems and Methods for Analysis of Hypertext Markup Language

US20250093164A1

( en )

*

2023-09-15

2025-03-20

Google Llc

Foundational Models for Semantic Routing

US20250103052A1

( en )

*

2023-09-26

2025-03-27

Boston Dynamics, Inc.

Dynamic performance of actions by a mobile robot based on sensor data and a site model

US12554924B2

( en )

*

2023-10-09

2026-02-17

Quabbin Patent Holdings, Inc.

Computer-implemented methods and systems for generative text painting

WO2025080599A1

( en )

2023-10-09

2025-04-17

Quabbin Patent Holdings

Computer-implemented methods and systems for dynamic prompt generation and integration with large language models for document revision

US12513176B2

( en )

*

2023-10-26

2025-12-30

A10 Networks Inc

Large language model based intelligent malicious packet detection

US20250145176A1

( en )

*

2023-11-02

2025-05-08

Nec Laboratories America, Inc.

Llm-based hybrid planner for autonomous driving

TWI905569B

( en )

*

2023-11-14

2025-11-21

財團法人資訊工業策進會

Novel multisensory decision generating device and method

US20250181835A1

( en )

*

2023-11-30

2025-06-05

Intuit Inc.

Indirect lookup using semantic matching and a large language model

US12602549B1

( en )

*

2025-02-12

2026-04-14

AtomBeam Technologies Inc.

Persistent cognitive machine with curated long term memory

US12386832B1

( en )

*

2024-01-26

2025-08-12

Salesforce, Inc.

Synthetic data generation for query plans

US20250247303A1

( en )

*

2024-01-31

2025-07-31

Microsoft Technology Licensing, Llc

Cloud architect

US20250334415A1

( en )

*

2024-04-26

2025-10-30

Zoox, Inc.

Generating local graph data

US12608584B2

( en )

*

2024-06-10

2026-04-21

Aaru, Inc.

Agent-based modeler using multimodal input

US20260057108A1

( en )

*

2024-08-26

2026-02-26

Ford Global Technologies, Llc

Vehicle based anonymization of localization vehicle data

US20260105000A1

( en )

*

2024-10-16

2026-04-16

Dell Products L.P.

Model and query server for local inferencing and training with generative models

US12561346B1

( en )

*

2024-11-06

2026-02-24

Sap Se

LLM-powered microservice registration and discovery

US12373935B1

( en )

*

2025-01-21

2025-07-29

Uveye Ltd.

Generating interactive vehicle inspection interfaces using multi-model artificial intelligence and anchor-based spatial tracking

US12585882B1

( en )

*

2025-02-12

2026-03-24

Atobeam Technologies Inc.

Evolutionary thought caching for multi-stage language model systems

US12517941B1

( en )

*

2025-04-04

2026-01-06

Poma Ai Gmbh

Retrieval-augmented generation for large language models

US12602207B1

( en )

*

2025-05-13

2026-04-14

Moonshot AI Inc

Method and system for classifying content of web pages using machine learning techniques

CN120780424B

( en )

*

2025-06-24

2026-05-08

深圳市爱德数智科技股份有限公司

A task collaboration management method and system based on offline applications

CN120949974B

( en )

*

2025-10-16

2025-12-26

中国科学院空天信息创新研究院

Map interaction system and method based on large language model and MCP protocol

CN121390817B

( en )

*

2025-12-24

2026-04-28

浙江农林大学

A Geospatial Analysis Intelligent Agent System and Method Based on a Large Model

CN121478874B

( en )

*

2026-01-07

2026-04-28

吉奥时空信息技术股份有限公司

A method, system, and storage medium for generating spatiotemporal narrative maps

Citations (19)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20090191901A1

( en )

*

1994-06-24

2009-07-30

Behr David A

Electronic navigation system and method

US20140236472A1

( en )

*

2011-12-29

2014-08-21

Barbara Rosario

Navigation systems and associated methods

US9424461B1

( en )

*

2013-06-27

2016-08-23

Amazon Technologies, Inc.

Object recognition for three-dimensional bodies

US20170371861A1

( en )

*

2016-06-24

2017-12-28

Mind Lakes, Llc

Architecture and processes for computer learning and understanding

US20180158157A1

( en )

*

2016-12-02

2018-06-07

Bank Of America Corporation

Geo-targeted Property Analysis Using Augmented Reality User Devices

US20190095428A1

( en )

*

2017-09-26

2019-03-28

Hitachi, Ltd.

Information processing apparatus, dialogue processing method, and dialogue system

US20190251759A1

( en )

*

2016-06-30

2019-08-15

The Car Force Inc.

Vehicle data aggregation and analysis platform providing dealership service provider dashboard

US20190271559A1

( en )

*

2018-03-02

2019-09-05

DeepMap Inc.

Visualization of high definition map data

US20210162995A1

( en )

*

2018-08-14

2021-06-03

Mobileye Vision Technologies Ltd.

Navigation relative to pedestrians at crosswalks

US20220180056A1

( en )

*

2020-12-09

2022-06-09

Here Global B.V.

Method and apparatus for translation of a natural language query to a service execution language

US20220335074A1

( en )

*

2020-04-24

2022-10-20

Baidu Online Network Technology (Beijing) Co., Ltd.

Method and apparatus of establishing similarity model for retrieving geographic location

US20240161520A1

( en )

*

2022-11-10

2024-05-16

Salesforce, Inc.

Systems and methods for a vision-language pretraining framework

US20240210194A1

( en )

*

2022-05-02

2024-06-27

Google Llc

Determining places and routes through natural conversation

US20240354491A1

( en )

*

2023-04-24

2024-10-24

Yahoo Assets Llc

Computerized systems and methods for an electronic inbox digest

US20240395261A1

( en )

*

2023-05-23

2024-11-28

Xiaomin Li

Virtual assistant with adaptive personality traits

US20240414337A1

( en )

*

2023-06-08

2024-12-12

Hitachi, Ltd.

Adaptive image compression for connected vehicles

US20240420418A1

( en )

*

2023-06-16

2024-12-19

Nvidia Corporation

Using language models in autonomous and semi-autonomous systems and applications

US20240419903A1

( en )

*

2023-06-16

2024-12-19

Nvidia Corporation

Processing sensor data using language models in map generation systems and applications

US20250200283A1

( en )

*

2023-06-16

2025-06-19

Nvidia Corporation

Using large language models to augment perception data in environment reconstruction systems and applications

Family Cites Families (32)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US7606714B2

( en )

*

2003-02-11

2009-10-20

Microsoft Corporation

Natural language classification within an automated response system

US20110054899A1

( en )

*

2007-03-07

2011-03-03

Phillips Michael S

Command and control utilizing content information in a mobile voice-to-speech application

US9189959B2

( en )

*

2012-06-27

2015-11-17

International Business Machines Corporation

Navigation system providing a super detail mode of operation to assist user's driving

US10002177B1

( en )

*

2013-09-16

2018-06-19

Amazon Technologies, Inc.

Crowdsourced analysis of decontextualized data

DE112016000308T5

( en )

*

2015-01-09

2017-10-19

Harman International Industries, Incorporated

Techniques for adjusting the level of detail of driving instructions

US9519643B1

( en )

*

2015-06-15

2016-12-13

Microsoft Technology Licensing, Llc

Machine map label translation

US20210297453A1

( en )

*

2015-10-28

2021-09-23

Qomplx, Inc.

Pathfinding in two and three-dimensional spaces using an automated planning service

US9612123B1

( en )

*

2015-11-04

2017-04-04

Zoox, Inc.

Adaptive mapping to navigate autonomous vehicles responsive to physical environment changes

US10831202B1

( en )

*

2017-09-01

2020-11-10

Zoox, Inc.

Onboard use of scenario description language

US11194994B2

( en )

*

2017-12-20

2021-12-07

X Development Llc

Semantic zone separation for map generation

US10789288B1

( en )

*

2018-05-17

2020-09-29

Shutterstock, Inc.

Relational model based natural language querying to identify object relationships in scene

US10984780B2

( en )

*

2018-05-21

2021-04-20

Apple Inc.

Global semantic word embeddings using bi-directional recurrent neural networks

US10795793B1

( en )

*

2018-11-19

2020-10-06

Intuit Inc.

Method and system for simulating system failures using domain-specific language constructs

US10955841B2

( en )

*

2018-12-28

2021-03-23

At&T Intellectual Property I, L.P.

Autonomous vehicle sensor security system

US10965712B2

( en )

*

2019-04-15

2021-03-30

Qualys, Inc.

Domain specific language for defending against a threat-actor and adversarial tactics, techniques, and procedures

US11157784B2

( en )

*

2019-05-08

2021-10-26

GM Global Technology Operations LLC

Explainable learning system and methods for autonomous driving

EP3907679B1

( en )

*

2020-05-08

2023-09-13

Accenture Global Solutions Limited

Enhanced robot fleet navigation and sequencing

US11907863B2

( en )

*

2020-07-24

2024-02-20

International Business Machines Corporation

Natural language enrichment using action explanations

EP4200717A2

( en )

*

2020-08-24

2023-06-28

Unlikely Artificial Intelligence Limited

A computer implemented method for the automated analysis or use of data

US12045288B1

( en )

*

2020-09-24

2024-07-23

Amazon Technologies, Inc.

Natural language selection of objects in image data

US12099832B2

( en )

*

2020-11-06

2024-09-24

Universite De Grenoble Alpes

Internet of things (IoT) capability platform

US12001798B2

( en )

*

2021-01-13

2024-06-04

Salesforce, Inc.

Generation of training data for machine learning based models for named entity recognition for natural language processing

US12381876B2

( en )

*

2021-04-22

2025-08-05

Microsoft Technology Licensing, Llc

Anomaly-based mitigation of access request risk

WO2022258671A2

( en )

*

2021-06-08

2022-12-15

Five AI Limited

Support tools for autonomous vehicle testing

US12340611B2

( en )

*

2021-10-13

2025-06-24

Samsung Electronics Co., Ltd.

Method and electronic device for intelligently sharing content

US20230394855A1

( en )

*

2022-06-01

2023-12-07

Microsoft Technology Licensing, Llc

Image paragraph generator

US20240212265A1

( en )

*

2022-12-21

2024-06-27

Meta Platforms Technologies, Llc

Generative VR World Creation from Natural Language

US20240267779A1

( en )

*

2023-02-03

2024-08-08

Tusimple, Inc.

Autonomous vehicle communication gateway manager

US12469505B2

( en )

*

2023-02-14

2025-11-11

The Nielsen Company (Us), Llc

Use of steganographic information as basis to process a voice command

US20240303496A1

( en )

*

2023-03-09

2024-09-12

Adobe Inc.

Exploiting domain-specific language characteristics for language model pretraining

US20240330534A1

( en )

*

2023-03-30

2024-10-03

Beta Air, Llc

Apparatus and methods for battery model generation

US20240385744A1

( en )

*

2023-05-17

2024-11-21

Jeremey White

MindGallery: AI Powered Digital Art Display with Vocal Command & Touchscreen Interface

2023

2023-09-22

US

US18/472,941

patent/US20240419903A1/en

active

Pending

2023-09-26

US

US18/474,591

patent/US20240419904A1/en

active

Pending

2023-10-09

US

US18/483,089

patent/US20240419905A1/en

active

Pending

2023-11-02

US

US18/500,426

patent/US20240419906A1/en

active

Pending

2023-11-06

US

US18/502,747

patent/US20240419907A1/en

active

Pending

2024

2024-01-10

US

US18/409,018

patent/US20240418515A1/en

active

Pending

2024-01-19

US

US18/417,105

patent/US20240419902A1/en

active

Pending

Patent Citations (24)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20090191901A1

( en )

*

1994-06-24

2009-07-30

Behr David A

Electronic navigation system and method

US20140236472A1

( en )

*

2011-12-29

2014-08-21

Barbara Rosario

Navigation systems and associated methods

US9424461B1

( en )

*

2013-06-27

2016-08-23

Amazon Technologies, Inc.

Object recognition for three-dimensional bodies

US20170371861A1

( en )

*

2016-06-24

2017-12-28

Mind Lakes, Llc

Architecture and processes for computer learning and understanding

US20190251759A1

( en )

*

2016-06-30

2019-08-15

The Car Force Inc.

Vehicle data aggregation and analysis platform providing dealership service provider dashboard

US20180158157A1

( en )

*

2016-12-02

2018-06-07

Bank Of America Corporation

Geo-targeted Property Analysis Using Augmented Reality User Devices

US20190095428A1

( en )

*

2017-09-26

2019-03-28

Hitachi, Ltd.

Information processing apparatus, dialogue processing method, and dialogue system

US20190271559A1

( en )

*

2018-03-02

2019-09-05

DeepMap Inc.

Visualization of high definition map data

US20210162995A1

( en )

*

2018-08-14

2021-06-03

Mobileye Vision Technologies Ltd.

Navigation relative to pedestrians at crosswalks

US20220335074A1

( en )

*

2020-04-24

2022-10-20

Baidu Online Network Technology (Beijing) Co., Ltd.

Method and apparatus of establishing similarity model for retrieving geographic location

US20220180056A1

( en )

*

2020-12-09

2022-06-09

Here Global B.V.

Method and apparatus for translation of a natural language query to a service execution language

US20240210194A1

( en )

*

2022-05-02

2024-06-27

Google Llc

Determining places and routes through natural conversation

US20240160853A1

( en )

*

2022-11-10

2024-05-16

Salesforce, Inc.

Systems and methods for a vision-language pretraining framework

US20240161520A1

( en )

*

2022-11-10

2024-05-16

Salesforce, Inc.

Systems and methods for a vision-language pretraining framework

US20240354491A1

( en )

*

2023-04-24

2024-10-24

Yahoo Assets Llc

Computerized systems and met

Related documents

Record · ID 653956
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.