ConceptioArchiveGoogle Patents
Google Patentsopen access

Systems and methods for depth estimation by learning triangulation and … — Magic Leap, Inc. (US11948320B2)

Magic Leap, Inc. · Google Patents
Google Patents · Patents · License: Open Access
Open Source ↗
magicleap
patent, google patents, intellectual property, US11948320B2, Magic Leap, Inc., Ayan Tuhinendu SINHA, en, 2024

ABSTRACT

Abstract

Systems and methods for estimating depths of features in a scene or environment surrounding a user of a spatial computing system, such as a virtual reality, augmented reality or mixed reality (collectively, cross reality) system, in an end-to-end process. The estimated depths can be utilized by a spatial computing system, for example, to provide an accurate and effective 3D cross reality experience.

Description

CROSS REFERENCE TO RELATED APPLICATIONS

The present application claims benefit under 35 U.S.C. § 119 to U.S. Provisional Patent Application Ser. No. 62/985,773 filed on Mar. 5, 2020, entitled “SYSTEMS AND METHODS FOR DEPTH ESTIMATION BY LEARNING TRIANGULATION AND DENSIFICATION OF SPARSE POINTS FOR MULTI-VIEW STEREO,” which is hereby incorporated by reference into the present application in its entirety.

FIELD OF THE INVENTION

The present invention is related to computing, learning network configurations, and connected mobile computing systems, methods, and configurations, and more specifically to systems and methods for estimating depths of features in a scene from multi-view images, which estimated depths may be used in mobile computing systems, methods, and configurations featuring at least one wearable component configured for virtual and/or augmented reality operation.

BACKGROUND

Modern computing and display technologies have facilitated the development of systems for so called “virtual reality” (“VR”), “augmented reality” (“AR”), and/or “mixed reality” (“MR”) environments or experiences, referred to collectively as “cross-reality” (“XR”) environments or experiences. This can be done by presenting computer-generated imagery to a user through a head-mounted display. This imagery creates a sensory experience which immerses the user in a simulated environment. This data may describe, for example, virtual objects that may be rendered in a way that users' sense or perceive as a part of a physical world and can interact with the virtual objects. The user may experience these virtual objects as a result of the data being rendered and presented through a user interface device, such as, for example, a head-mounted display device. The data may be displayed to the user to see, or may control audio that is played for the user to hear, or may control a tactile (or haptic) interface, enabling the user to experience touch sensations that the user senses or perceives as feeling the virtual object.

XR systems may be useful for many applications, spanning the fields of scientific visualization, medical training, engineering design and prototyping, tele-manipulation and tele-presence, and personal entertainment. VR systems typically involve presentation of digital or virtual image information without transparency to actual real-world visual input.

AR systems generally supplement a real-world environment with simulated elements. For example, AR systems may provide a user with a view of a surrounding real-world environment via a head-mounted display. Computer-generated imagery can also be presented on the head-mounted display to enhance the surrounding real-world environment. This computer-generated imagery can include elements which are contextually-related to the surrounding real-world environment. Such elements can include simulated text, images, objects, and the like. MR systems also introduce simulated objects into a real-world environment, but these objects typically feature a greater degree of interactivity than in AR systems.

AR/MR scenarios often include presentation of virtual image elements in relationship to real-world objects. For example, an AR/MR scene is depicted wherein a user of an AR/MR technology sees a real-world scene featuring the environment surrounding the user, including structures, objects, etc. In addition to these features, the user of the AR/MR technology perceives that they “see” computer generated features (i.e., virtual object), even though such features do not exist in the real-world environment. Accordingly, AR and MR, in contrast to VR, include one or more virtual objects in relation to real objects of the physical world. The virtual objects also interact with the real world objects, such that the AR/MR system may also be termed a “spatial computing” system in relation to the system's interaction with the 3D world surrounding the user. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment in using the XR system, and also opens the door for a variety of applications that present realistic and readily understandable information about how the physical world might be altered.

The visualization center of the brain gains valuable perception information from the motion of both eyes and components thereof relative to each other. Vergence movements (i.e., rolling movements of the pupils toward or away from each other to converge the lines of sight of the eyes to fixate upon an object) of the two eyes relative to each other are closely associated with accommodation (or focusing) of the lenses of the eyes. Under normal conditions, accommodating the eyes, or changing the focus of the lenses of the eyes, to focus upon an object at a different distance will automatically cause a matching change in vergence to the same distance, under a relationship known as the “accommodation-vergence reflex.” Likewise, a change in vergence will trigger a matching change in accommodation, under normal conditions. Working against this reflex, as do most conventional stereoscopic VR/AR/MR configurations, is known to produce eye fatigue, headaches, or other forms of discomfort in users.

Stereoscopic wearable glasses generally feature two displays—one for the left eye and one for the right eye—that are configured to display images with slightly different element presentation such that a three-dimensional perspective is perceived by the human visual system. Such configurations have been found to be uncomfortable for many users due to a mismatch between vergence and accommodation (“vergence-accommodation conflict”) which must be overcome to perceive the images in three dimensions. Indeed, some users are not able to tolerate stereoscopic configurations. These limitations apply to VR, AR, and MR systems. Accordingly, most conventional VR/AR/MR systems are not optimally suited for presenting a rich, binocular, three-dimensional experience in a manner that will be comfortable and maximally useful to the user, in part because prior systems fail to address some of the fundamental aspects of the human perception system, including the vergence-accommodation conflict.

Various systems and methods have been disclosed for addressing the vergence-accommodation conflict. For example, U.S. Utility patent application Ser. No. 14/555,585 discloses VR/AR/MR systems and methods that address the vergence-accommodation conflict by projecting light at the eyes of a user using one or more light-guiding optical elements such that the light and images rendered by the light appear to originate from multiple depth planes. All patent applications, patents, publications, and other references referred to herein are hereby incorporation by reference in their entireties, and for all purposes. The light-guiding optical elements are designed to in-couple virtual light corresponding to digital or virtual objects, propagate it by total internal reflection (“TIR”), and then out-couple the virtual light to display the virtual objects to the user's eyes. In AR/MR systems, the light-guiding optical elements are also designed to be transparent to light from (e.g., reflecting off of) actual real-world objects. Therefore, portions of the light-guiding optical elements are designed to reflect virtual light for propagation via TIR while being transparent to real-world light from real-world objects in AR/MR systems.

AR/MR scenarios often include interactions between virtual objects and a real-world physical environment. Similarly, some VR scenarios include interactions between completely virtual objects and other virtual objects. Delineating objects in the physical environment facilitates interactions with virtual objects by defining the metes and bounds of those interactions (e.g., by defining the extent of a particular structure or object in the physical environment). For instance, if an AR/MR scenario includes a virtual object (e.g., a tentacle or a fist) extending from a particular object in the physical environment, defining the extent of the object in three dimensions allows the AR/MR system to present a more realistic AR/MR scenario. Conversely, if the extent of objects is not defined or inaccurately defined, artifacts and errors will occur in the displayed images. For instance, a virtual object may appear to extend partially or entirely from midair adjacent an object instead of from the surface of the object. As another example, if an AR/MR scenario includes a virtual character walking on a particular horizontal surface in a physical environment, inaccurately defining the extent of the surface may result in the virtual character appearing to walk off of the surface without falling, and instead floating in midair.

Accordingly, depth sensing of scenes, such as a surrounding environment, are useful for in a wide range of applications, ranging from cross reality systems to autonomous driving. Estimating depth of scenes can be broadly divided into classes: active and passive sensing. Active sensing techniques include LiDAR, structured-light and time-of-flight (ToF) cameras, whereas depth estimation using a monocular camera or stereopsis of an array of cameras is termed passive sensing. Active sensors are currently the de-facto standard of applications requiring depth sensing due to good accuracy and low latency in varied environments. (see [Ref. 44]). Numbered references in brackets (“[Ref. ##]”) refer to the reference list appended below; each of these references is incorporated by reference in its entirety herein.

However, active sensors have their own of limitation. LiDARs are prohibitively expensive and provide sparse measurements. Structured-light and ToF depth cameras have limited range and completeness due to the physics of light transport. Furthermore, they are power hungry and inhibit mobility critical for AR/VR applications on wearables. Consequently, computer vision researchers have pursued passive sensing techniques as a ubiquitous, cost-effective and energy-efficient alternative to active sensors. (See [Ref. 30]).

Passive depth sensing using stereo cameras requires a large baseline and careful calibration for accurate depth estimation. (See [Ref. 3]). A large baseline is infeasible for mobile devices like phones and wearables. An alternative is to use multi-view stereo (MVS) techniques for a moving monocular camera to estimate depth. MVS generally refers to the problem of reconstructing 3D scene structure from multiple images with known camera poses and intrinsics. (See [Ref 14]). The unconstrained nature of camera motion alleviates the baseline limitation of stereo-rigs, and the algorithm benefits from multiple observations of the same scene from continuously varying viewpoints. (See [Ref. 17]). However, camera motion also makes depth estimation more challenging relative to rigid stereo-rigs due to pose uncertainty and added complexity of motion artifacts. Most MVS approaches involve building a 3D cost volume, usually with a plane sweep stereo approach. (See [Refs. 41,18]). Accurate depth estimation using MVS relies on 3D convolutions on the cost volume, which is both memory as well as computationally expensive, scaling cubically with the resolution. Furthermore, redundant compute is added by ignoring useful image-level properties such as interest points and their descriptors, which are a necessary precursor to camera pose estimation, and hence, any MVS technique. This increases the overall cost and energy requirements for passive sensing.

Passive sensing using a single image is fundamentally unreliable due to scale ambiguity in 2D images. Deep learning based monocular depth estimation approaches formulate the problem as depth regression (see [Refs. 10,11]) and have reduced the performance gap to those of active sensors (see [Refs. 26,24]), but still far from being practical. Recently, sparse-to-dense depth estimation approaches have been proposed to remove the scale ambiguity and improve robustness of monocular depth estimation. (See [Ref. 30]. Indeed, recent sparse-to-dense approaches with less than 0.5% depth samples have accuracy comparable to active sensors, with higher range and completeness. (See [Ref 6]. However, these approaches assume accurate or seed depth samples from an active sensor which is limiting. The alternative is to use the sparse 3D landmarks output from the best performing algorithms for Simultaneous Localization and Mapping (SLAM) (see [Ref 31]) or Visual Inertial Odometry (VIO) (see [Ref. 33]). However, using depth evaluated from these sparse landmarks in lieu of depth from active sensors, significantly degrades performance. (See [Ref 43]). This is not surprising as the learnt sparse-to-dense network ignores potentially useful cues, structured noise and biases present in SLAM or VIO algorithm.

Sparse feature based methods are standard for SLAM or VIO techniques due to their high speed and accuracy. The detect-then-describe approach is the most common approach to sparse feature extraction, wherein, interest points are detected and then described for a patch around the point. The descriptor encapsulates higher level information, which is missed by typical low-level interest points such as corners, blobs, etc. Prior to the deep learning revolution, classical systems like SIFT (see [Ref 28] and ORB (see [Ref 37] were ubiquitously used as descriptors for feature matching for low level vision tasks. Deep neural networks directly optimizing for the objective at hand have now replaced these hand engineered features across a wide array of applications. However, such an end-to-end network has remained elusive for SLAM (see [Ref. 32] due to the components being non-differentiable. General purpose descriptors learned by methods such as SuperPoint (see [Ref 9], LIFT (see [Ref 42]), and GIFT (see [Ref 27] aim to bridge the gap towards differentiable SLAM.

MVS approaches either directly reconstruct a 3D volume or output a depth map which can be flexibly used for 3D reconstruction or other applications. Methods of reconstructing 3D volumes (see [Ref 41, 5] are restricted to small spaces or isolated objects either due to the high memory load of operating in a 3D voxelized space (see [Refs. 35, 39], or due to the difficulty of learning point representations in complex environments (see [Ref. 34]). The use of multi-view images captured in indoor environments has progressed lately starting with DeepMVS (see [Ref. 18]) which proposed a learned patch matching approach. MVDepthNet (see [Ref 40]), and DPSNet (see [Ref 19] build a cost volume for depth estimation. Recently, GP-MVSNet (see [Ref 17]) built upon MVDepthNet to coherently fuse temporal information using Gaussian processes. All these methods utilize the plane sweep algorithm during some stage of depth estimation, resulting in an accuracy vs efficiency trade-off.

Sparse-to-dense depth estimation has also recently emerged as a way to supplement active depth sensors due to their range limitations when operating on a power budget, and to fill in depth in hard to detect regions such as dark or reflective objects. One approach was proposed by Ma et. al (see [Ref 30], which was followed by Chen et. al. (see [Ref. 6, 43]) which introduced innovations in the representation and network architecture. A convolutional spatial propagation module is proposed in [Ref. 7] to in-fill the missing depth values. Recently, self-supervised approaches (see [Refs. 13, 12]) have been explored for the sparse-to-dense problem. (See [Ref. 29]).

It can be seen that multi-view stereo (MVS) represents an advantageous middle approach between the accuracy of active depth sensing and the practicality of monocular depth estimation. Cost volume based approaches employing 3D convolutional neural networks (CNNs) have considerably improved the accuracy of MVS systems. However, this accuracy comes at a high computational cost which impedes practical adoption.

Accordingly, there is a need for improved systems and methods for depth estimation of a scene which does not depend on costly and ineffective active depth sensing, and improves upon the efficiency and/or accuracy of prior passive depth sensing techniques. In addition, the systems and methods for depth estimation should be implementable in XR systems having displays which are lightweight, low-cost, have a small form-factor, have a wide virtual image field of view, and are as transparent as possible.

SUMMARY

The embodiments disclosed herein are directed to systems and methods for estimating depths of features in a scene or environment surrounding a user of a spatial computing system, such as an XR system, in an end-to-end process. The estimated depths can be utilized by a spatial computing system, for example, to provide an accurate and effective 3D XR experience. The resulting 3D XR experience is displayable in a rich, binocular, three-dimensional experience that is comfortable and maximally useful to the user, in part because it can present images in a manner which addresses some of the fundamental aspects of the human perception system, such as the vergence-accommodation mismatch. For instance, the estimated depths may be used to generate a 3D reconstruction having accurate depth data enabling the 3D images to be displayed in multiple focal planes. The 3D reconstruction may also enable accurate management of interactions between virtual objects, other virtual objects, and/or real world objects.

<div id="p-0022" num="0021" class="descripti

CROSS REFERENCE TO RELATED APPLICATIONS

The present application claims benefit under 35 U.S.C. § 119 to U.S. Provisional Patent Application Ser. No. 62/985,773 filed on Mar. 5, 2020, entitled “SYSTEMS AND METHODS FOR DEPTH ESTIMATION BY LEARNING TRIANGULATION AND DENSIFICATION OF SPARSE POINTS FOR MULTI-VIEW STEREO,” which is hereby incorporated by reference into the present application in its entirety.

FIELD OF THE INVENTION

The present invention is related to computing, learning network configurations, and connected mobile computing systems, methods, and configurations, and more specifically to systems and methods for estimating depths of features in a scene from multi-view images, which estimated depths may be used in mobile computing systems, methods, and configurations featuring at least one wearable component configured for virtual and/or augmented reality operation.

BACKGROUND

Modern computing and display technologies have facilitated the development of systems for so called “virtual reality” (“VR”), “augmented reality” (“AR”), and/or “mixed reality” (“MR”) environments or experiences, referred to collectively as “cross-reality” (“XR”) environments or experiences. This can be done by presenting computer-generated imagery to a user through a head-mounted display. This imagery creates a sensory experience which immerses the user in a simulated environment. This data may describe, for example, virtual objects that may be rendered in a way that users&#39; sense or perceive as a part of a physical world and can interact with the virtual objects. The user may experience these virtual objects as a result of the data being rendered and presented through a user interface device, such as, for example, a head-mounted display device. The data may be displayed to the user to see, or may control audio that is played for the user to hear, or may control a tactile (or haptic) interface, enabling the user to experience touch sensations that the user senses or perceives as feeling the virtual object.

XR systems may be useful for many applications, spanning the fields of scientific visualization, medical training, engineering design and prototyping, tele-manipulation and tele-presence, and personal entertainment. VR systems typically involve presentation of digital or virtual image information without transparency to actual real-world visual input.

AR systems generally supplement a real-world environment with simulated elements. For example, AR systems may provide a user with a view of a surrounding real-world environment via a head-mounted display. Computer-generated imagery can also be presented on the head-mounted display to enhance the surrounding real-world environment. This computer-generated imagery can include elements which are contextually-related to the surrounding real-world environment. Such elements can include simulated text, images, objects, and the like. MR systems also introduce simulated objects into a real-world environment, but these objects typically feature a greater degree of interactivity than in AR systems.

AR/MR scenarios often include presentation of virtual image elements in relationship to real-world objects. For example, an AR/MR scene is depicted wherein a user of an AR/MR technology sees a real-world scene featuring the environment surrounding the user, including structures, objects, etc. In addition to these features, the user of the AR/MR technology perceives that they “see” computer generated features (i.e., virtual object), even though such features do not exist in the real-world environment. Accordingly, AR and MR, in contrast to VR, include one or more virtual objects in relation to real objects of the physical world. The virtual objects also interact with the real world objects, such that the AR/MR system may also be termed a “spatial computing” system in relation to the system&#39;s interaction with the 3D world surrounding the user. The experience of virtual objects interacting with real objects greatly enhances the user&#39;s enjoyment in using the XR system, and also opens the door for a variety of applications that present realistic and readily understandable information about how the physical world might be altered.

The visualization center of the brain gains valuable perception information from the motion of both eyes and components thereof relative to each other. Vergence movements (i.e., rolling movements of the pupils toward or away from each other to converge the lines of sight of the eyes to fixate upon an object) of the two eyes relative to each other are closely associated with accommodation (or focusing) of the lenses of the eyes. Under normal conditions, accommodating the eyes, or changing the focus of the lenses of the eyes, to focus upon an object at a different distance will automatically cause a matching change in vergence to the same distance, under a relationship known as the “accommodation-vergence reflex.” Likewise, a change in vergence will trigger a matching change in accommodation, under normal conditions. Working against this reflex, as do most conventional stereoscopic VR/AR/MR configurations, is known to produce eye fatigue, headaches, or other forms of discomfort in users.

Stereoscopic wearable glasses generally feature two displays—one for the left eye and one for the right eye—that are configured to display images with slightly different element presentation such that a three-dimensional perspective is perceived by the human visual system. Such configurations have been found to be uncomfortable for many users due to a mismatch between vergence and accommodation (“vergence-accommodation conflict”) which must be overcome to perceive the images in three dimensions. Indeed, some users are not able to tolerate stereoscopic configurations. These limitations apply to VR, AR, and MR systems. Accordingly, most conventional VR/AR/MR systems are not optimally suited for presenting a rich, binocular, three-dimensional experience in a manner that will be comfortable and maximally useful to the user, in part because prior systems fail to address some of the fundamental aspects of the human perception system, including the vergence-accommodation conflict.

Various systems and methods have been disclosed for addressing the vergence-accommodation conflict. For example, U.S. Utility patent application Ser. No. 14/555,585 discloses VR/AR/MR systems and methods that address the vergence-accommodation conflict by projecting light at the eyes of a user using one or more light-guiding optical elements such that the light and images rendered by the light appear to originate from multiple depth planes. All patent applications, patents, publications, and other references referred to herein are hereby incorporation by reference in their entireties, and for all purposes. The light-guiding optical elements are designed to in-couple virtual light corresponding to digital or virtual objects, propagate it by total internal reflection (“TIR”), and then out-couple the virtual light to display the virtual objects to the user&#39;s eyes. In AR/MR systems, the light-guiding optical elements are also designed to be transparent to light from (e.g., reflecting off of) actual real-world objects. Therefore, portions of the light-guiding optical elements are designed to reflect virtual light for propagation via TIR while being transparent to real-world light from real-world objects in AR/MR systems.

AR/MR scenarios often include interactions between virtual objects and a real-world physical environment. Similarly, some VR scenarios include interactions between completely virtual objects and other virtual objects. Delineating objects in the physical environment facilitates interactions with virtual objects by defining the metes and bounds of those interactions (e.g., by defining the extent of a particular structure or object in the physical environment). For instance, if an AR/MR scenario includes a virtual object (e.g., a tentacle or a fist) extending from a particular object in the physical environment, defining the extent of the object in three dimensions allows the AR/MR system to present a more realistic AR/MR scenario. Conversely, if the extent of objects is not defined or inaccurately defined, artifacts and errors will occur in the displayed images. For instance, a virtual object may appear to extend partially or entirely from midair adjacent an object instead of from the surface of the object. As another example, if an AR/MR scenario includes a virtual character walking on a particular horizontal surface in a physical environment, inaccurately defining the extent of the surface may result in the virtual character appearing to walk off of the surface without falling, and instead floating in midair.

Accordingly, depth sensing of scenes, such as a surrounding environment, are useful for in a wide range of applications, ranging from cross reality systems to autonomous driving. Estimating depth of scenes can be broadly divided into classes: active and passive sensing. Active sensing techniques include LiDAR, structured-light and time-of-flight (ToF) cameras, whereas depth estimation using a monocular camera or stereopsis of an array of cameras is termed passive sensing. Active sensors are currently the de-facto standard of applications requiring depth sensing due to good accuracy and low latency in varied environments. (see [Ref. 44]). Numbered references in brackets (“[Ref. ##]”) refer to the reference list appended below; each of these references is incorporated by reference in its entirety herein.

However, active sensors have their own of limitation. LiDARs are prohibitively expensive and provide sparse measurements. Structured-light and ToF depth cameras have limited range and completeness due to the physics of light transport. Furthermore, they are power hungry and inhibit mobility critical for AR/VR applications on wearables. Consequently, computer vision researchers have pursued passive sensing techniques as a ubiquitous, cost-effective and energy-efficient alternative to active sensors. (See [Ref. 30]).

Passive depth sensing using stereo cameras requires a large baseline and careful calibration for accurate depth estimation. (See [Ref. 3]). A large baseline is infeasible for mobile devices like phones and wearables. An alternative is to use multi-view stereo (MVS) techniques for a moving monocular camera to estimate depth. MVS generally refers to the problem of reconstructing 3D scene structure from multiple images with known camera poses and intrinsics. (See [Ref 14]). The unconstrained nature of camera motion alleviates the baseline limitation of stereo-rigs, and the algorithm benefits from multiple observations of the same scene from continuously varying viewpoints. (See [Ref. 17]). However, camera motion also makes depth estimation more challenging relative to rigid stereo-rigs due to pose uncertainty and added complexity of motion artifacts. Most MVS approaches involve building a 3D cost volume, usually with a plane sweep stereo approach. (See [Refs. 41,18]). Accurate depth estimation using MVS relies on 3D convolutions on the cost volume, which is both memory as well as computationally expensive, scaling cubically with the resolution. Furthermore, redundant compute is added by ignoring useful image-level properties such as interest points and their descriptors, which are a necessary precursor to camera pose estimation, and hence, any MVS technique. This increases the overall cost and energy requirements for passive sensing.

Passive sensing using a single image is fundamentally unreliable due to scale ambiguity in 2D images. Deep learning based monocular depth estimation approaches formulate the problem as depth regression (see [Refs. 10,11]) and have reduced the performance gap to those of active sensors (see [Refs. 26,24]), but still far from being practical. Recently, sparse-to-dense depth estimation approaches have been proposed to remove the scale ambiguity and improve robustness of monocular depth estimation. (See [Ref. 30]. Indeed, recent sparse-to-dense approaches with less than 0.5% depth samples have accuracy comparable to active sensors, with higher range and completeness. (See [Ref 6]. However, these approaches assume accurate or seed depth samples from an active sensor which is limiting. The alternative is to use the sparse 3D landmarks output from the best performing algorithms for Simultaneous Localization and Mapping (SLAM) (see [Ref 31]) or Visual Inertial Odometry (VIO) (see [Ref. 33]). However, using depth evaluated from these sparse landmarks in lieu of depth from active sensors, significantly degrades performance. (See [Ref 43]). This is not surprising as the learnt sparse-to-dense network ignores potentially useful cues, structured noise and biases present in SLAM or VIO algorithm.

Sparse feature based methods are standard for SLAM or VIO techniques due to their high speed and accuracy. The detect-then-describe approach is the most common approach to sparse feature extraction, wherein, interest points are detected and then described for a patch around the point. The descriptor encapsulates higher level information, which is missed by typical low-level interest points such as corners, blobs, etc. Prior to the deep learning revolution, classical systems like SIFT (see [Ref 28] and ORB (see [Ref 37] were ubiquitously used as descriptors for feature matching for low level vision tasks. Deep neural networks directly optimizing for the objective at hand have now replaced these hand engineered features across a wide array of applications. However, such an end-to-end network has remained elusive for SLAM (see [Ref. 32] due to the components being non-differentiable. General purpose descriptors learned by methods such as SuperPoint (see [Ref 9], LIFT (see [Ref 42]), and GIFT (see [Ref 27] aim to bridge the gap towards differentiable SLAM.

MVS approaches either directly reconstruct a 3D volume or output a depth map which can be flexibly used for 3D reconstruction or other applications. Methods of reconstructing 3D volumes (see [Ref 41, 5] are restricted to small spaces or isolated objects either due to the high memory load of operating in a 3D voxelized space (see [Refs. 35, 39], or due to the difficulty of learning point representations in complex environments (see [Ref. 34]). The use of multi-view images captured in indoor environments has progressed lately starting with DeepMVS (see [Ref. 18]) which proposed a learned patch matching approach. MVDepthNet (see [Ref 40]), and DPSNet (see [Ref 19] build a cost volume for depth estimation. Recently, GP-MVSNet (see [Ref 17]) built upon MVDepthNet to coherently fuse temporal information using Gaussian processes. All these methods utilize the plane sweep algorithm during some stage of depth estimation, resulting in an accuracy vs efficiency trade-off.

Sparse-to-dense depth estimation has also recently emerged as a way to supplement active depth sensors due to their range limitations when operating on a power budget, and to fill in depth in hard to detect regions such as dark or reflective objects. One approach was proposed by Ma et. al (see [Ref 30], which was followed by Chen et. al. (see [Ref. 6, 43]) which introduced innovations in the representation and network architecture. A convolutional spatial propagation module is proposed in [Ref. 7] to in-fill the missing depth values. Recently, self-supervised approaches (see [Refs. 13, 12]) have been explored for the sparse-to-dense problem. (See [Ref. 29]).

It can be seen that multi-view stereo (MVS) represents an advantageous middle approach between the accuracy of active depth sensing and the practicality of monocular depth estimation. Cost volume based approaches employing 3D convolutional neural networks (CNNs) have considerably improved the accuracy of MVS systems. However, this accuracy comes at a high computational cost which impedes practical adoption.

Accordingly, there is a need for improved systems and methods for depth estimation of a scene which does not depend on costly and ineffective active depth sensing, and improves upon the efficiency and/or accuracy of prior passive depth sensing techniques. In addition, the systems and methods for depth estimation should be implementable in XR systems having displays which are lightweight, low-cost, have a small form-factor, have a wide virtual image field of view, and are as transparent as possible.

SUMMARY

The embodiments disclosed herein are directed to systems and methods for estimating depths of features in a scene or environment surrounding a user of a spatial computing system, such as an XR system, in an end-to-end process. The estimated depths can be utilized by a spatial computing system, for example, to provide an accurate and effective 3D XR experience. The resulting 3D XR experience is displayable in a rich, binocular, three-dimensional experience that is comfortable and maximally useful to the user, in part because it can present images in a manner which addresses some of the fundamental aspects of the human perception system, such as the vergence-accommodation mismatch. For instance, the estimated depths may be used to generate a 3D reconstruction having accurate depth data enabling the 3D images to be displayed in multiple focal planes. The 3D reconstruction may also enable accurate management of interactions between virtual objects, other virtual objects, and/or real world objects.

Accordingly, one embodiment is directed to a method for estimating depth of features in a scene from multi-view images. First, multi-view images are obtained, including an anchor image of the scene and a set of reference images of the scene. This may be accomplished by one or more suitable cameras, such as cameras of an XR system. The anchor image and reference images are passed through a shared RGB encoder and descriptor decoder which (1) outputs a respective descriptor field of descriptors for the anchor image and each reference image, (ii) detects interest points in the anchor image in conjunction with relative poses to determine a search space in the reference images from alternate view-points, and (iii) outputs intermediate feature maps. The respective descriptors are sampled in the search space of each reference image to determine descriptors in the search space and matching the identified descriptors with descriptors for the interest points in the anchor image. The matched descriptors are referred to as matched keypoints. The matched keypoints are triangulated using singular value decomposition (SVD) to output 3D points. The 3D points are passed through a sparse depth encoder to create a sparse depth image from the 3D points and output feature maps. A depth decoder then generates a dense depth image based on the output feature maps for the sparse depth encoder and the intermediate feature maps from the RGB encoder.

In another aspect of the method, the shared RGB encoder and descriptor decoder may comprise two encoders including an RGB image encoder and a sparse depth image encoder, and three decoders including an interest point detection encoder, a descriptor decoder, and a dense depth prediction encoder.

In still another aspect, the shared RGB encoder and descriptor decoder may be a fully-convolutional neural network configured to operate on a full resolution of the anchor image and transaction images.

In yet another aspect, the method may further comprise feeding the feature maps from the RGB encoder into a first task-specific decoder head to determine weights for the detecting of interest points in the anchor image and outputting interest point descriptions.

In yet another aspect of the method, the descriptor decoder may comprise a U-Net like architecture to fuse fine and course level image information for matching the identified descriptors with descriptors for the interest points.

In another aspect of the method, the search space may be constrained to a respective epipolar line in the reference images plus a fixed offset on either side of the epipolar line, and within a feasible depth sensing range along the epipolar line.

In still another aspect of the method, bilinear sampling may be used by the shared RGB encoder and descriptor decoder to output the respective descriptors at desired points in the descriptor field.

In another aspect of the method, the step of triangulating the matched keypoints comprises estimating respective two dimensional (2D) positions of the interest points by computing a softmax across spatial axes to output cross-correlation maps; performing a soft-argmax operation to calculate the 2D position of joints as a center of mass of corresponding cross-correlation maps; performing a linear algebraic triangulation from the 2D estimates; and using a singular value decomposition (SVD) to output 3D points.

Another disclosed embodiment is directed to a cross reality (XR) system which is configured to estimate depths, and utilized such depths as described herein. The cross reality system comprises a head-mounted display device having a display system. For example, the head-mounted display may have a pair of near-eye displays in an eyeglasses-like structure. A computing system is in operable communication with the head-mounted display. A plurality of camera sensors are in operable communication with the computing system. The computing system is configured to estimate depths of features in a scene from a plurality of multi-view images captured by the camera sensors any of the methods described above. In additional aspects of the cross reality system, the process may include any one or more of the additional aspects of the cross reality system described above. For instance, the process may include obtaining a multi-view images, including an anchor image of the scene and a set of reference images of a scene within a field of view of the camera sensors from the camera sensors; passing the anchor image and reference images through a shared RGB encoder and descriptor decoder which (1) outputs a respective descriptor field of descriptors for the anchor image and each reference image, (ii) detects interest points in the anchor image in conjunction with relative poses to determine a search space in the reference images from alternate view-points, and (iii) outputs intermediate feature maps; sampling the respective descriptors in the search space of each reference image to determine descriptors in the search space and matching the identified descriptors with descriptors for the interest points in the anchor image, such matched descriptors referred to as matched keypoints; triangulating the matched keypoints using singular value decomposition (SVD) to output 3D points; passing the 3D points through a sparse depth encoder to create a sparse depth image from the 3D points and output feature maps; and a depth decoder generating a dense depth image based on the output feature maps for the sparse depth encoder and the intermediate feature maps from the RGB encoder.

BRIEF DESCRIPTION OF THE DRAWINGS

This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

The drawings illustrate the design and utility of preferred embodiments of the present disclosure, in which similar elements are referred to by common reference numerals. In order to better appreciate how the above-recited and other advantages and objects of the present disclosure are obtained, a more particular description of the present disclosure briefly described above will be rendered by reference to specific embodiments thereof, which are illustrated in the accompanying drawings. Understanding that these drawings depict only typical embodiments of the disclosure and are not therefore to be considered limiting of its scope, the disclosure will be described and explained with additional specificity and detail through the use of the accompanying drawings.

FIG. 1 is a schematic diagram of an exemplary cross reality system for providing a cross reality experience, according to one embodiment.

FIG. 2 is a schematic diagram of a method for depth estimation of a scene, according to one embodiment.

FIG. 3 is a block diagram of the architecture of a shared RGB encoder and descriptor decoder used in the method of FIG. 2 , according to one embodiment.

FIG. 4 illustrates a process for restricting the range of the search space using epipolar sampling and depth range sampling, as used in the method of FIG. 2 , according to one embodiment.

FIG. 5 is a block diagram illustrating the architecture for a key-point network, as used in the method of FIG. 2 , according to one embodiment.

FIG. 6 illustrates a qualitative comparison between an example of the method of FIG. 2 and various other different methods.

FIG. 7 shows sample 3D reconstructions of the scene from the estimated depth maps I the example of the method of FIG. 2 , described herein.

FIG. 8 shows a Table 1 having a comparison of the performance of different descriptors on ScanNet.

FIG. 9 shows a Table 2 having a comparison of the performance of depth estimation on ScanNet.

FIG. 10 shows a Table 3 having a comparison of the performance of depth estimation on ScanNet for different numbers of images.

FIG. 11 shows a Table 4 having a comparison of depth estimation on Sun3D.

FIG. 12 sets forth an equation for a process for the descriptor of each interest point being convolved with the descriptor field along its corresponding epipolar line for each image view-point as used in the method of FIG. 2 , according to one embodiment.

FIGS. 13 - 16 set forth equations for a process for an algebraic triangulation to obtain 3D points as used in the method of FIG. 2 , according to one embodiment.

DETAILED DESCRIPTION

The following describes various embodiments of systems and methods for estimating depths of features in a scene or environment surrounding a user of a spatial computing system, such as an XR system, in an end-to-end process. The various embodiments are described in detail with reference to the drawings, which are provided as illustrative examples of the disclosure to enable those skilled in the art to practice the disclosure. Notably, the figures and the examples below are not meant to limit the scope of the present disclosure. Where certain elements of the present disclosure may be partially or fully implemented using known components (or methods or processes), only those portions of such known components (or methods or processes) that are necessary for an understanding of the present disclosure will be described, and the detailed descriptions of other portions of such known components (or methods or processes) will be omitted so as not to obscure the disclosure. Further, various embodiments encompass present and future known equivalents to the components referred to herein by way of illustration.

Furthermore, the systems and methods for estimating depths of features in a scene or environment surrounding a user of a spatial computing system may also be implemented independently of XR systems, and the embodiments depicted herein are described in relation to XR systems for illustrative purposes only.

Referring to FIG. 1 , an exemplary XR system 100 according to one embodiment is illustrated. The XR system 100 includes a head-mounted display device 2 (also referred to as a head worn viewing component 2 ), a hand-held controller 4 (also referred to as a hand-held controller component 4 ), and an interconnected auxiliary computing system or controller 6 (also referred to as an interconnected auxiliary computing system or controller component 6 ) which may be configured to be worn as a belt pack or the like on the user. Each of these components are in operable communication (i.e., operatively coupled) to each other and to other connected resources 8 (such as cloud computing or cloud storage resources) via wired or

wireless communication connections

10 , 12 , 14 , 16 , 17 , 18 , such as those specified by IEEE 802.11, Bluetooth®, and other connectivity standards and configurations. The head-mounted display device 2 includes two depicted optical elements 20 through which the user may see the world around them along with video images and visual components produced by the associated system components, including a pair of image sources (e.g., micro-display panels) and viewing optics for displaying computer generated images on the optical elements 20 , for an augmented reality experience. In the illustrated embodiment, the head-mounted display device 2 and pair of image sources are lightweight, low-cost, have a small form-factor, have a wide virtual image field of view, and are as transparent as possible. As illustrated in FIG. 1 , the XR system 100 also includes various sensors configured to provide information pertaining to the environment around the user, including but not limited to various camera type sensors

22 , 24 , 26 (such as monochrome, color/RGB, and/or thermal), depth camera sensors 28 , and/or sound sensors 30 (such as microphones).

In addition, it is desirable that the XR system 100 is configured to present virtual image information in multiple focal planes (for example, two or more) in order to be practical for a wide variety of use-cases without exceeding an acceptable allowance for vergence-accommodation mismatch. U.S. patent application Ser. Nos. 14/555,585, 14/690,401, 14/331,218, 15/481,255, 62/627,155, 62/518,539, 16/229,532, 16/155,564, 15/413,284, 16/020,541, 62,702,322, 62/206,765, 15,597,694, 16/221,065, 15/968,673, and 62/682,788, each of which is incorporated by reference herein in its entirety, describe various aspects of the XR system 100 and its components in more detail.

In various embodiments a user wears an augmented reality system such as the XR system 100 depicted in FIG. 1 , which may also be termed a “spatial computing” system in relation to such system&#39;s interaction with the three dimensional world around the user when operated. The

cameras

22 , 24 , 26 and computing system 6 are configured to map the environment around the user, and/or to create a “mesh” of such environment, comprising various points representative of the geometry of various objects within the environment around the user, such as walls, floors, chairs, and the like. The spatial computing system may be configured to map or mesh the environment around the user, and to run or operate software, such as that available from Magic Leap, Inc., of Plantation, Florida, which may be configured to utilize the map or mesh of the room to assist the user in placing, manipulating, visualizing, creating, and modifying various objects and elements in the three-dimensional space around the user. As shown in FIG. 1 , the XR system 100 may also be operatively coupled to additional connected resources 8 , such as other computing systems, by cloud or other connectivity configurations.

It is understood that the methods, systems and configurations described herein are broadly applicable to various scenarios outside of the realm of wearable spatial computing such as the XR system 100 , subject to the appropriate sensors and associated data being available.

In contrast to prior systems and methods for depth estimation of scenes, the presently disclosed systems and methods learn the sparse 3D landmarks in conjunction with the sparse to dense formulation in an end-to-end manner so as to (a) remove dependence on a cost volume in the MVS technique, thus, significantly reducing compute, (b) complement camera pose estimation using sparse VIO or SLAM by reusing detected interest points and descriptors, (c) utilize geometry-based MVS concepts to guide the algorithm and improve the interpretability, and (d) benefit from the accuracy and efficiency of sparse-to-dense techniques. The network in the present systems and methods is a multitask model (see [Ref 22]), comprised of an encoder-decoder structure composed of two encoders, one for RGB image and one for sparse depth image, and three decoders: one interest point detection, one for descriptors and one for the dense depth prediction. A differentiable module is also utilized that efficiently triangulates points using geometric priors and forms the critical link between the interest point decoder, descriptor decoder, and the sparse depth encoder enabling end-to-end training.

These methods and configurations are broadly applicable to various scenarios outside of realm of wearable spatial computing, subject to the appropriate sensors and associated data being available.

One of the challenges in spatial computing relates to the utilization of data captured by various operatively coupled sensors (such as

elements

22 , 24 , 26 , 28 of the system of FIG. 1 ) of the XR system 100 in making determinations useful and/or critical to the user, such as in computer vision and/or object recognition challenges that may, for example, relate to the three-dimensional world around a user. Disclosed herein are methods and systems for generating a 3D reconstruction of a scene, such as the 3D environment surrounding the user of the XR system 100 , using only RGB images, such as the RGB images from the

cameras

22 , 24 , and 26 , without using depth data from the depth sensors 28 .

In contrast to previous methods of depth estimation of scenes, such as indoor environments, the present disclosure introduces an approach for depth estimation by learning triangulation and densification of sparse points for multi-view stereo. Distinct from cost volume approaches, the presently discloses systems and methods utilize an efficient depth estimation approach by first (a) detecting and evaluating descriptors for interest points, then (b) learning to match and triangulate a small set of interest points, and finally densifying this sparse set of 3D points using CNNs. An end-to-end network efficiently performs all three steps within a deep learning framework and trained with intermediate 2D image and 3D geometric supervision, along with depth supervision. Crucially, the first step of the presently disclosed method complements pose estimation using interest point detection and descriptor learning. The present methods are shown to produce state-of-the-art results on depth estimation with lower compute for different scene lengths. Furthermore, this method generalizes to newer environments and the descriptors output by the network compare favorably to strong baselines.

In the present disclosed method, the sparse 3D landmarks are learned in conjunction with the sparse to dense formulation in an end-to-end manner so as to (a) remove the dependence on a cost volume as in the MVS technique, thus, significantly reducing computational costs, (b) complement camera pose estimation using sparse VIO or SLAM by reusing detected interest points and descriptors, (c) utilize geometry-based MVS concepts to guide the algorithm and improve the interpretability, and (d) benefit from the accuracy and efficiency of sparse-to-dense techniques. The network used in the method is a multitask model (e.g., see [Ref 22]), comprised of an encoder-decoder structure composed of two encoders, one for RGB image and one for sparse depth image, and three decoders: one interest point detection, one for descriptors and one for the dense depth prediction. The method also utilizes a differentiable module that efficiently triangulates points using geometric priors and forms the critical link between the interest point decoder, descriptor decoder, and the sparse depth encoder enabling end-to-end training.

One embodiment of a method 110 , as well as a system 110 , for depth estimation of a scene is can be broadly sub-divided into three steps as illustrated in the schematic diagram of FIG. 2 . The method 110 can be broadly sub-divided into three steps as illustrated FIG. 2 . In the first step 112 , the target or anchor image 114 and the multi-view images 116 are passed through a shared RGB encoder and descriptor decoder 118 (including an RGB image encoder 119 , a detector decoder 121 , and a descriptor decoder 123 ) to output a descriptor field 120 for each image

114 , 116 . Interest points 122 are also detected for the target or the anchor image 114 . In the second step 124 , the interest points 122 in the anchor image 114 in conjunction with the relative poses 126 are used to determine the search space in the reference images 116 from alternate view-points. Descriptors 132 are sampled in the search space using an epipolar sampler 127 and point sampler 129 , respectively, to output sampled descriptors 128 and are matched by a soft matcher 130 with descriptors 128 for the interest points 122 . Then, the matched keypoints 134 are triangulated using SVD using a triangulation module 136 to output 3D points 138 . The output 3D points 138 are used by a sparse depth encoder 140 to create a sparse depth image. In the third and final step 142 , the output feature maps for the sparse depth encoder 140 and intermediate feature maps from the RGB encoder 119 are collectively used to inform the depth decoder 144 and output a dense depth image 146 . Each of the three steps are described in greater detail below.

As described above, the shared RGB encoder and descriptor decoder 118 is composed of two encoders, the RGB image encoder 119 and the sparse depth image encoder 140 , and three decoders, the detector decoder 121 (also referred to as the interest point detector decoder 121 ), the descriptor decoder 123 , and the dense depth decoder 144 (also referred to as dense depth predictor decoder 144 ). In one embodiment, the shared RGB encoder and descriptor decoder 118 may comprise a SuperPoint-like (see [Ref. 9]) formulation of a fully-convolutional neural network architecture which operates on a full-resolution image and produces interest point detection accompanied by fixed length descriptors. The model has a single, shared encoder to process and reduce the input image dimensionality. The feature maps from the RGB encoder 119 feed into two task-specific decoder “heads”, which learn weights for interest point detection and interest point description. This joint formulation of interest point detection and description in SuperPoint enables sharing compute for the detection and description tasks, as well as the downstream task of depth estimation. However, SuperPoint was trained on grayscale images with focus on interest point detection and description for continuous pose estimation on high frame rate video streams, and hence, has a relatively shallow encoder. On the contrary, the present method is interested in image sequences with sufficient baseline, and consequently longer intervals between subsequent frames. Furthermore, SuperPoint&#39;s shallow backbone suitable for sparse point analysis has limited capacity for our downstream task of dense depth estimation. Hence, the shallow backbone is replaced with a ResNet-50 (see [Ref. 16]) encoder which balances efficiency and performance. The output resolution of the interest point detector decoder 121 is identical to that of SuperPoint. In order to fuse fine and coarse level image information critical for point matching, the method 110 may utilize a U-Net (see [Ref. 36]) like architecture for the descriptor decoder 123 . The descriptor decoder 123 outputs an N- dimensional descriptor tensor 120 at ⅛th the image resolution, similar to SuperPoint. This architecture is illustrated in FIG. 3 . The interest point detector network is trained by distilling the output of the original SuperPoint network and the descriptors are trained by the matching formulation described below.

The previous step provides interest points for the anchor image and descriptors for all images, i.e., the anchor image and full set of reference images. The next step 124 of the method 110 includes point matching and triangulation. A naïve approach would be to match descriptors of the interest points 122 sampled from the descriptor field 120 of the anchor image 114 to all possible positions in each reference image 116 . However, this is computationally prohibitive. Hence, the method 110 invokes geometrical constraints to restrict the search space and improve efficiency. Using concepts from multi-view geometry, the method e100 only searches along the epipolar line in the reference images (see [ FIG. 14 ]). The epipolar line is determined using the fundamental matrix, F, using the relation xFx T =0, where x is the set of points in the image. The matched point is guaranteed to lie on the epipolar li

CLAIMS

Claims ( 16 )

What is claimed is:

1. A method for estimating depth of features in a scene from multi-view images, the method comprising:

obtaining multi-view images, including an anchor image of the scene and a set of reference images of the scene;

passing the anchor image and reference images through a shared RGB encoder and descriptor decoder which (1) outputs a respective descriptor field of descriptors for the anchor image and each reference image, (ii) detects sparse interest points in the anchor image in conjunction with relative poses to determine a search space in the reference images from alternate view-points, and (iii) outputs intermediate feature maps, wherein in the descriptor decoder utilizes a learning architecture configured to train descriptors;

sampling the respective descriptors in the search space of each reference image to determine descriptors in the search space and matching the identified descriptors with descriptors for the sparse interest points in the anchor image, such matched descriptors referred to as matched keypoints;

triangulating the matched keypoints using singular value decomposition (SVD) to output 3D points;

passing the 3D points through a sparse depth encoder to create a sparse depth image from the 3D points and output feature maps; and

a depth decoder generating a dense depth image based on the output feature maps for the sparse depth encoder and the intermediate feature maps from the RGB encoder.

2. The method of claim 1 , wherein the shared RGB encoder and descriptor decoder comprises two encoders including an RGB image encoder and a sparse depth image encoder, and three decoders including an interest point detection encoder, a descriptor decoder, and a dense depth prediction encoder.

3. The method of claim 1 , wherein the shared RGB encoder and descriptor decoder is a fully-convolutional neural network configured to operate on a full resolution of the anchor image and transaction images.

4. The method of claim 1 , further comprising:

feeding the feature maps from the RGB encoder into a first task-specific decoder head to determine weights for the detecting of the sparse interest points in the anchor image and outputting interest point descriptions.

5. The method of claim 1 , wherein the descriptor decoder comprises a U-Net like architecture to fuse fine and course level image information for matching the identified descriptors with descriptors for the sparse interest points.

6. The method of claim 1 , wherein the search space is constrained to a respective epipolar line in the reference images plus a fixed offset on either side of the epipolar line, and within a feasible depth sensing range along the epipolar line.

7. The method of claim 1 , wherein bilinear sampling is used by the shared RGB encoder and descriptor decoder to output the respective descriptors at desired points in the descriptor field.

8. The method of claim 1 , wherein the step of triangulating the matched keypoints comprises:

estimating respective two dimensional (2D) positions of the sparse interest points by computing a softmax across spatial axes to output cross-correlation maps;

performing a soft-argmax operation to calculate the 2D position of joints as a center of mass of corresponding cross-correlation maps;

performing a linear algebraic triangulation from the 2D estimates; and

using a singular value decomposition (SVD) to output 3D points.

9. A cross reality system, comprising:

a head-mounted display device having a display system;

a computing system in operable communication with the head-mounted display;

a plurality of camera sensors in operable communication with the computing system;

wherein the computing system is configured to estimate depths of features in a scene from a plurality of multi-view images captured by the camera sensors by a process comprising:

obtaining a multi-view images, including an anchor image of the scene and a set of reference images of a scene within a field of view of the camera sensors from the camera sensors;

passing the anchor image and reference images through a shared RGB encoder and descriptor decoder which (1) outputs a respective descriptor field of descriptors for the anchor image and each reference image, (ii) detects sparse interest points in the anchor image in conjunction with relative poses to determine a search space in the reference images from alternate view-points, and (iii) outputs intermediate feature maps, wherein in the descriptor decoder utilizes a learning architecture configured to train descriptors;

sampling the respective descriptors in the search space of each reference image to determine descriptors in the search space and matching the identified descriptors with descriptors for the sparse interest points in the anchor image, such matched descriptors referred to as matched keypoints;

triangulating the matched keypoints using singular value decomposition (SVD) to output 3D points;

passing the 3D points through a sparse depth encoder to create a sparse depth image from the 3D points and output feature maps; and

a depth decoder generating a dense depth image based on the output feature maps for the sparse depth encoder and the intermediate feature maps from the RGB encoder.

10. The cross reality system of claim 9 , wherein the shared RGB encoder and descriptor decoder comprises two encoders including an RGB image encoder and a sparse depth image encoder, and three decoders including an interest point detection encoder, a descriptor decoder, and a dense depth prediction encoder.

11. The cross reality system of claim 9 , wherein the shared RGB encoder and descriptor decoder is a fully-convolutional neural network configured to operate on a full resolution of the anchor image and transaction images.

12. The cross reality system of claim 9 , wherein the process for estimating depths of features in a scene from a plurality of multi-view images captured by the camera sensors further comprises:

feeding the feature maps from the RGB encoder into a first task-specific decoder head to determine weights for the detecting of the sparse interest points in the anchor image and outputting interest point descriptions.

13. The cross reality system of claim 9 , wherein the descriptor decoder comprises a U-Net like architecture to fuse fine and course level image information for matching the identified descriptors with descriptors for the sparse interest points.

14. The cross reality system of claim 9 , wherein the search space is constrained to a respective epipolar line in the reference images plus a fixed offset on either side of the epipolar line, and within a feasible depth sensing range along the epipolar line.

15. The cross reality system of claim 9 , wherein bilinear sampling is used by the shared RGB encoder and descriptor decoder to output the respective descriptors at desired points in the descriptor field.

16. The cross reality system of claim 9 , wherein the step of triangulating the matched keypoints comprises:

estimating respective two dimensional (2D) positions of the sparse interest points by computing a softmax across spatial axes to output cross-correlation maps;

performing a soft-argmax operation to calculate the 2D position of joints as a center of mass of corresponding cross-correlation maps;

performing a linear algebraic triangulation from the 2D estimates; and

using a singular value decomposition (SVD) to output 3D points.

US17/194,117

2020-03-05

2021-03-05

Systems and methods for depth estimation by learning triangulation and densification of sparse points for multi-view stereo

Active

2041-06-07

US11948320B2

( en )

Priority Applications (1)

Application Number

Priority Date

Filing Date

Title

US17/194,117

US11948320B2

( en )

2020-03-05

2021-03-05

Systems and methods for depth estimation by learning triangulation and densification of sparse points for multi-view stereo

Applications Claiming Priority (2)

Application Number

Priority Date

Filing Date

Title

US202062985773P

2020-03-05

2020-03-05

US17/194,117

US11948320B2

( en )

2020-03-05

2021-03-05

Systems and methods for depth estimation by learning triangulation and densification of sparse points for multi-view stereo

Publications (2)

Publication Number

Publication Date

US20210279904A1

US20210279904A1 ( en )

2021-09-09

US11948320B2

true

US11948320B2 ( en )

2024-04-02

Family

ID=77554899

Family Applications (1)

Application Number

Title

Priority Date

Filing Date

US17/194,117

Active

2041-06-07

US11948320B2

( en )

2020-03-05

2021-03-05

Systems and methods for depth estimation by learning triangulation and densification of sparse points for multi-view stereo

Country Status (5)

Country

Link

US

( 1 )

US11948320B2

( en )

EP

( 1 )

EP4115145A4

( en )

JP

( 1 )

JP7640569B2

( en )

CN

( 1 )

CN115210532A

( en )

WO

( 1 )

WO2021178919A1

( en )

Cited By (2)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20220108420A1

( en )

*

2021-12-14

2022-04-07

Intel Corporation

Method and system of efficient image rendering for near-eye light field displays

US12315182B2

( en )

*

2022-02-02

2025-05-27

Rapyuta Robotics Co., Ltd.

Apparatus and a method for estimating depth of a scene

Families Citing this family (22)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20200137380A1

( en )

*

2018-10-31

2020-04-30

Intel Corporation

Multi-plane display image synthesis mechanism

BR112022023671A2

( en )

*

2020-06-01

2022-12-20

Qualcomm Inc

METHODS AND APPARATUS FOR OCCLUSION MANIPULATION TECHNIQUES

US11481871B2

( en )

*

2021-03-12

2022-10-25

Samsung Electronics Co., Ltd.

Image-guided depth propagation for space-warping images

US12211202B2

( en )

*

2021-10-13

2025-01-28

GE Precision Healthcare LLC

Self-supervised representation learning paradigm for medical images

US12086965B2

( en )

*

2021-11-05

2024-09-10

Adobe Inc.

Image reprojection and multi-image inpainting based on geometric depth parameters

CN114972451B

( en )

*

2021-12-06

2025-03-04

东华理工大学

A remote sensing image registration method based on rotation-invariant SuperGlue matching

CN114170316B

( en )

*

2021-12-13

2025-05-06

杭州电子科技大学

A pose estimation optimization method based on surround viewpoints

CN114820745B

( en )

*

2021-12-13

2024-09-13

南瑞集团有限公司

Monocular vision depth estimation system, method, computer device and computer readable storage medium

CN114332510B

( en )

*

2022-01-04

2024-03-22

安徽大学

Hierarchical image matching method

CN114742794B

( en )

*

2022-04-02

2024-08-27

北京信息科技大学

A temporary road detection method and system based on triangulation

CN114913287B

( en )

*

2022-04-07

2023-08-22

北京拙河科技有限公司

Three-dimensional human body model reconstruction method and system

WO2023225235A1

( en )

*

2022-05-19

2023-11-23

Innopeak Technology, Inc.

Method for predicting depth map via multi-view stereo system, electronic apparatus and storage medium

US12340530B2

( en )

2022-05-27

2025-06-24

Toyota Research Institute, Inc.

Photometric cost volumes for self-supervised depth estimation

CN115619892B

( en )

*

2022-06-06

2026-03-24

北京机电工程研究所

A Real-Time Dense Mapping Method for UAVs Based on Visual SLAM and Deep Learning

CN115222889A

( en )

*

2022-07-19

2022-10-21

深圳万兴软件有限公司

3D reconstruction method and device based on multi-view image and related equipment

CN115330935A

( en )

*

2022-08-02

2022-11-11

广东顺德工业设计研究院(广东顺德创新设计研究院)

Three-dimensional reconstruction method and system based on deep learning

CN115496857B

( en )

*

2022-09-22

2026-03-24

淮阴工学院

A Deep Learning-Based Multi-View 3D Reconstruction Method

CN116071504B

( en )

*

2023-03-06

2023-06-09

安徽大学

Multi-view three-dimensional reconstruction method for high-resolution image

CN117033769B

( en )

*

2023-07-07

2026-01-13

中国平安人寿保险股份有限公司

Front-end component retrieval method, device, equipment and storage medium

CN116934829B

( en )

*

2023-09-15

2023-12-12

天津云圣智能科技有限责任公司

Unmanned aerial vehicle target depth estimation method and device, storage medium and electronic equipment

CN119991966B

( en )

*

2025-02-17

2025-10-28

北京航空航天大学

Multi-view three-dimensional reconstruction method based on epipolar transducer and MVSNet network enhanced by implicit neural optimization

CN120411345B

( en )

*

2025-03-13

2025-10-17

中国科学院计算技术研究所

A multi-view stereo reconstruction method based on segmentation-driven and edge-aligned deformation

Citations (20)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20120163672A1

( en )

2010-12-22

2012-06-28

David Mckinnon

Depth Estimate Determination, Systems and Methods

US20140049612A1

( en )

2011-10-11

2014-02-20

Panasonic Corporation

Image processing device, imaging device, and image processing method

US20140192154A1

( en )

2011-08-09

2014-07-10

Samsung Electronics Co., Ltd.

Method and device for encoding a depth map of multi viewpoint video data, and method and device for decoding the encoded depth map

US20150199825A1

( en )

*

2014-01-13

2015-07-16

Transgaming Inc.

Method and system for expediting bilinear filtering

US20150237329A1

( en )

2013-03-15

2015-08-20

Pelican Imaging Corporation

Systems and Methods for Estimating Depth Using Ad Hoc Stereo Array Cameras

US20150262412A1

( en )

*

2014-03-17

2015-09-17

Qualcomm Incorporated

Augmented reality lighting with dynamic geometry

US20150269723A1

( en )

*

2014-03-18

2015-09-24

Arizona Board Of Regents On Behalf Of Arizona State University

Stereo vision measurement system and method

US20160148079A1

( en )

2014-11-21

2016-05-26

Adobe Systems Incorporated

Object detection using cascaded convolutional neural networks

US9619748B1

( en )

2002-09-30

2017-04-11

Michael Lamport Commons

Intelligent control with hierarchical stacked neural networks

US20180176545A1

( en )

2016-11-25

2018-06-21

Nokia Technologies Oy

Virtual reality display

US20180365532A1

( en )

2017-06-20

2018-12-20

Nvidia Corporation

Semi-supervised learning for landmark localization

US20190051056A1

( en )

*

2017-08-11

2019-02-14

Sri International

Augmenting reality using semantic segmentation

US20190108683A1

( en )

2016-04-01

2019-04-11

Pcms Holdings, Inc.

Apparatus and method for supporting interactive augmented reality functionalities

US20190130275A1

( en )

2017-10-26

2019-05-02

Magic Leap, Inc.

Gradient normalization systems and methods for adaptive loss balancing in deep multitask networks

US10304193B1

( en )

*

2018-08-17

2019-05-28

12 Sigma Technologies

Image segmentation and object detection using fully convolutional neural network

US20190278983A1

( en )

*

2018-03-12

2019-09-12

Nvidia Corporation

Three-dimensional (3d) pose estimation from a monocular camera

US10682108B1

( en )

*

2019-07-16

2020-06-16

The University Of North Carolina At Chapel Hill

Methods, systems, and computer readable media for three-dimensional (3D) reconstruction of colonoscopic surfaces for determining missing regions

US20210142497A1

( en )

*

2019-11-12

2021-05-13

Geomagical Labs, Inc.

Method and system for scene image modification

US20210150747A1

( en )

*

2019-11-14

2021-05-20

Samsung Electronics Co., Ltd.

Depth image generation method and device

US20210237774A1

( en )

*

2020-01-31

2021-08-05

Toyota Research Institute, Inc.

Self-supervised 3d keypoint learning for monocular visual odometry

Family Cites Families (2)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20180242017A1

( en )

*

2017-02-22

2018-08-23

Twitter, Inc.

Transcoding video

DK201870700A1

( en )

*

2018-06-20

2020-01-14

Aptiv Technologies Limited

Over-the-air (ota) mobility services platform

2021

2021-03-05

JP

JP2022552548A

patent/JP7640569B2/en

active

Active

2021-03-05

US

US17/194,117

patent/US11948320B2/en

active

Active

2021-03-05

WO

PCT/US2021/021239

patent/WO2021178919A1/en

not_active

Ceased

2021-03-05

CN

CN202180017832.2A

patent/CN115210532A/en

active

Pending

2021-03-05

EP

EP21764896.3A

patent/EP4115145A4/en

active

Pending

Patent Citations (20)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US9619748B1

( en )

2002-09-30

2017-04-11

Michael Lamport Commons

Intelligent control with hierarchical stacked neural networks

US20120163672A1

( en )

2010-12-22

2012-06-28

David Mckinnon

Depth Estimate Determination, Systems and Methods

US20140192154A1

( en )

2011-08-09

2014-07-10

Samsung Electronics Co., Ltd.

Method and device for encoding a depth map of multi viewpoint video data, and method and device for decoding the encoded depth map

US20140049612A1

( en )

2011-10-11

2014-02-20

Panasonic Corporation

Image processing device, imaging device, and image processing method

US20150237329A1

( en )

2013-03-15

2015-08-20

Pelican Imaging Corporation

Systems and Methods for Estimating Depth Using Ad Hoc Stereo Array Cameras

US20150199825A1

( en )

*

2014-01-13

2015-07-16

Transgaming Inc.

Method and system for expediting bilinear filtering

US20150262412A1

( en )

*

2014-03-17

2015-09-17

Qualcomm Incorporated

Augmented reality lighting with dynamic geometry

US20150269723A1

( en )

*

2014-03-18

2015-09-24

Arizona Board Of Regents On Behalf Of Arizona State University

Stereo vision measurement system and method

US20160148079A1

( en )

2014-11-21

2016-05-26

Adobe Systems Incorporated

Object detection using cascaded convolutional neural networks

US20190108683A1

( en )

2016-04-01

2019-04-11

Pcms Holdings, Inc.

Apparatus and method for supporting interactive augmented reality functionalities

US20180176545A1

( en )

2016-11-25

2018-06-21

Nokia Technologies Oy

Virtual reality display

US20180365532A1

( en )

2017-06-20

2018-12-20

Nvidia Corporation

Semi-supervised learning for landmark localization

US20190051056A1

( en )

*

2017-08-11

2019-02-14

Sri International

Augmenting reality using semantic segmentation

US20190130275A1

( en )

2017-10-26

2019-05-02

Magic Leap, Inc.

Gradient normalization systems and methods for adaptive loss balancing in deep multitask networks

US20190278983A1

( en )

*

2018-03-12

2019-09-12

Nvidia Corporation

Three-dimensional (3d) pose estimation from a monocular camera

US10304193B1

( en )

*

2018-08-17

2019-05-28

12 Sigma Technologies

Image segmentation and object detection using fully convolutional neural network

US10682108B1

( en )

*

2019-07-16

2020-06-16

The University Of North Carolina At Chapel Hill

Methods, systems, and computer readable media for three-dimensional (3D) reconstruction of colonoscopic surfaces for determining missing regions

US20210142497A1

( en )

*

2019-11-12

2021-05-13

Geomagical Labs, Inc.

Method and system for scene image modification

US20210150747A1

( en )

*

2019-11-14

2021-05-20

Samsung Electronics Co., Ltd.

Depth image generation method and device

US20210237774A1

( en )

*

2020-01-31

2021-08-05

Toyota Research Institute, Inc.

Self-supervised 3d keypoint learning for monocular visual odometry

Non-Patent Citations (16)

* Cited by examiner, † Cited by third party

Title

Amendment After Final for U.S. Appl. No. 16/879,736 dated Jul. 29, 2022.

Chen et al. " Fully Convolutional Neural Network with Augmented Atrous Spatial Pyramid Pool and Fully Connected Fusion Path for High Resolution Remote Sensing Image Segmentation. " In: Appl. Sci .; May 1, 2019, [online] [retrieved on Jul. 30, 2020 (Jul. 30, 2020)] Retrieved from the Internet &lt;URL: https://www.mdpi.com/2076-3417/9/9/1816&gt;, entire document.

Extended European Search Report for EP Patent Appln. No. 20809006.8 dated Aug. 11, 2022.

Extended European Search Report for EP Patent Appln. No. 21764896.3 dated Jun. 30, 2023.

Final Office Action for U.S. Appl. No. 16/879,736 dated Apr. 29, 2022.

Final Office Action for U.S. Appl. No. 16/879,736 dated Jan. 12, 2023.

Fisher Yu et al.: " Deep Layer Aggregation " arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jul. 20, 2017 (Jul. 20, 2017), XP081001885, * Section 3 and 4; figures 2,3,4 *.

Foreign NOA for JP Patent Appln. No. 2021-568891 dated Oct. 13, 2023.

Foreign Response for EP Patent Appln. No. 20809006.8 dated Mar. 9, 2023.

Non-Final Office Action for U.S. Appl. No. 16/879,736 dated Sep. 15, 2022.

Notice of Allowance for U.S. Appl. No. 16/879,736 dated Mar. 31, 2023.

Notice of Allowance for U.S. Appl. No. 16/879,736 dated May 30, 2023.

PCT International Search Report and Written Opinion for International Appln. No. PCT/US20/33885 (Attorney Docket No. ML-0875WO), Applicant Magic Leap, Inc., dated Aug. 31, 2020 (10 pages).

PCT International Search Report and Written Opinion for International Appln. No. PCT/US21/21239, Applicant Magic Leap, Inc., dated May 26, 2021 (16 pages).

Rajeev Ranjan et al.: " HyperFace: A Deep Multi-Task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition ", May 11, 2016 (May 11, 2016), XP055552344, DOI: 10.1109/TPAMI.2017.2781233 Retrieved from the Internet: URL:https://arxiv.org/pdf/1603.01249v2 [retrieved on Feb. 6, 2019] * 3, 4 and 5; figures 2, 6 *.

Wende (On Improving the Performance of Multi-threaded CUDA Applications with Concurrent Kemel Execution by Kemel Reordering, 2012 Symposium on Application Accelerators in High Performance Computing) (Year: 2012).

Cited By (3)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20220108420A1

( en )

*

2021-12-14

2022-04-07

Intel Corporation

Method and system of efficient image rendering for near-eye light field displays

US12400290B2

( en )

*

2021-12-14

2025-08-26

Intel Corporation

Method and system of efficient image rendering for near-eye light field displays

US12315182B2

( en )

*

2022-02-02

2025-05-27

Rapyuta Robotics Co., Ltd.

Apparatus and a method for estimating depth of a scene

Also Published As

Publication number

Publication date

US20210279904A1

( en )

2021-09-09

WO2021178919A1

( en )

2021-09-10

EP4115145A4

( en )

2023-08-02

JP7640569B2

( en )

2025-03-05

JP2023515669A

( en )

2023-04-13

EP4115145A1

( en )

2023-01-11

CN115210532A

( en )

2022-10-18

Similar Documents

Publication

Publication Date

Title

US20210279904A1

( en )

2021-09-09

Systems and methods for depth estimation by learning triangulation and densification of sparse points for multi-view stereo

US11694387B2

( en )

2023-07-04

Systems and methods for end to end scene reconstruction from multiview images

Sinha et al.

2020

Deltas: Depth estimation by learning triangulation and densification of sparse points

Zioulis et al.

2018

Omnidepth: Dense depth estimation for indoors spherical panoramas

EP3942529B1

( en )

2025-10-01

Predicting three-dimensional articulated and target object pose

US11238606B2

( en )

2022-02-01

Method and system for performing simultaneous localization and mapping using convolutional image transformation

US11698529B2

( en )

2023-07-11

Systems and methods for distributing a neural network across multiple computing devices

Herrera-Granda et al.

2024

Monocular visual SLAM, visual odometry, and structure from motion methods applied to 3D reconstruction: A comprehensive survey

Niu et al.

2024

Overview of image-based 3D reconstruction technology

Zhou et al.

2022

A lightweight hand gesture recognition in complex backgrounds

Tenze et al.

2024

altiro3d: scene representation from single image and novel view synthesis

Yuan et al.

2025

Self-supervised monocular depth estimation with depth-motion prior for pseudo-lidar

Dao et al.

2022

FastMDE: A fast CNN architecture for monocular depth estimation at high resolution

He et al.

2024

Single image depth estimation using improved U-Net and edge-guide loss

Meng et al.

2019

Un-VDNet: unsupervised network for visual odometry and depth estimation

Sinha et al.

2020

Depth estimation by learning triangulation and densification of sparse points for multi-view stereo

Jiao et al.

2026

Large-kernel spatially parallel feature fusion for monocular 3D perception in autonomous driving

Liu et al.

2018

Template-based 3d reconstruction of non-rigid deformable object from monocular video

Sekkati et al.

2023

Depth Learning Methods for Bridges Inspection Using UAV

Yang et al.

2025

High-Precision Depth Estimation Networks Using Low-Resolution Depth and RGB Image Sensors for Low Cost MR Glasses

Ameen et al.

2026

Beyond RGB-D: Perception of Glass, Mirrors, and See-Through Scenes for Robotics

Sengupta

2020

Visual tracking of deformable objects with RGB-D camera

Thukral et al.

2024

Converting 2D Images to Point Cloud Using Depth Estimation

Fang

2023

Weakly Supervised Learning of Cross-Modal Depth Estimation Using RGB-D Sensors

Zhou et al.

2025

Transparent Object Depth Completion with Stereo Image Guidance

Legal Events

Date

Code

Title

Description

2021-03-05

FEPP

Fee payment procedure

Free format text : ENTITY STATUS SET TO UNDISCOUNTED (ORIGINAL EVENT CODE: BIG.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY

2021-03-10

AS

Assignment

Owner name : MAGIC LEAP, INC., FLORIDA

Free format text : ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNOR:SINHA, AYAN TUHINENDU;REEL/FRAME:055550/0264

Effective date : 20200727

2021-06-11

FEPP

Fee payment procedure

Free format text : PETITION RELATED TO MAINTENANCE FEES GRANTED (ORIGINAL EVENT CODE: PTGR); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY

2021-08-22

STPP

Information on status: patent application and granting procedure in general

Free format text : DOCKETED NEW CASE - READY FOR EXAMINATION

2022-05-24

AS

Assignment

Owner name : CITIBANK, N.A., AS COLLATERAL AGENT, NEW YORK

Free format text : SECURITY INTEREST;ASSIGNORS:MOLECULAR IMPRINTS, INC.;MENTOR ACQUISITION ONE, LLC;MAGIC LEAP, INC.;REEL/FRAME:060338/0665

Effective date : 20220504

2022-11-07

STPP

Information on status: patent application and granting procedure in general

Free format text : NON FINAL ACTION MAILED

2023-06-30

STPP

Information on status: patent application and granting procedure in general

Free format text : DOCKETED NEW CASE - READY FOR EXAMINATION

2023-07-24

STPP

Information on status: patent application and granting procedure in general

Free format text : NOTICE OF ALLOWANCE MAI

Related documents

Record · ID 607312
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.