ABSTRACT
Abstract
A distributed, online machine learning system is presented. Contemplated systems include many private data servers, each having local private data. Researchers can request that relevant private data servers train implementations of machine learning algorithms on their local private data without requiring de-identification of the private data or without exposing the private data to unauthorized computing systems. The private data servers also generate synthetic or proxy data according to the data distributions of the actual data. The servers then use the proxy data to train proxy models. When the proxy models are sufficiently similar to the trained actual models, the proxy data, proxy model parameters, or other learned knowledge can be transmitted to one or more non-private computing devices. The learned knowledge from many private data servers can then be aggregated into one or more trained global models without exposing private data.
Description
CROSS REFERENCE TO RELATED APPLICATION
This application claims the priority under 35 USC 119 from U.S. Provisional Patent Application Ser. 62/363,697, entitled Distributed Machine Learning Systems, Apparatus, and Methods, filed on Jul. 18, 2016 by Szeto, the contents of which are incorporated by reference in their entirety.
FIELD OF THE INVENTION
The field of the invention is distributed machine learning technologies.
BACKGROUND
The background description includes information that may be useful in understanding the present inventive subject matter. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed inventive subject matter, or that any publication specifically or implicitly referenced is prior art.
With the recent growth of highly accessible and cost-effective machine learning platforms (e.g., Google®'s Artificial Intelligence including TensorFlow, Amazon's Machine Learning, Microsoft's Azure Machine Learning, OpenAI, SciKit-Learn, Matlab, etc.), data analysts have numerous off-the-shelf options available to them for conducting automated analysis of large data sets. Additionally, in parallel to the growth of machine learning platforms, target data sets have also grown in size. For example, Yahoo! has released several large data sets to the public having sizes on the order of terabytes and The Cancer Genome Atlas (TCGA) data portal offers access to massive amounts of clinical information and genomic characterization data. These pre-built data sets are made readily available to data analysts.
Unfortunately, researchers often encounter obstacles when compiling data sets for their in-progress research, especially when attempting to build trained machine learning models capable of generating interesting predictions using in-the-field data. One major obstacle is that researchers often lack access to the data they require. Consider, for example, a scenario where a researcher wishes to build a trained model from patient data where the patient data is stored in multiple hospitals' electronic medical record databases. The researcher would likely not have authorization to access each hospital's patient data due to privacy restrictions or HIPAA compliance. In order to compile a desired data set, the researcher must request the data from the hospital. Assuming the hospital is amenable to the request, the hospital must then de-identify the data to remove references to specific patients before providing the data to the researcher. However, de-identification results in loss of possibly valuable information in the dataset that could be instrumental in training machine learning algorithms, which in turn can provide opportunities for discovering new relationships in the data or provide value predictive properties. Thus, because of the security restrictions, the datasets available to the researcher could lack information. Clearly, researchers would benefit from technologies that could extract learned information or âknowledgeâ while also respecting private or secured information distributed across multiple data stores.
Interestingly, previous efforts associated with analyzing distributed data focus on the nature of machine learning rather than dealing with the technical issues of isolated, private data. For example, U.S. Pat. No. 7,899,225 to Collins et al. titled âSystems and Methods of Clinical State Prediction Utilizing Medical Image Dataâ filed Oct. 26, 2006, describes creating and merging statistical models to create a final multi-dimensional classification space. The statistical models are the mathematical variation models which define the space in which subjects can be represented. Unfortunately, Collins assumes that the system has authorization to access all the data in order to build the predictive models. Collins also fails to provide insights into circumstances where non-centralized data must remain secure or private. Still, it would be useful to be able to combine trained models in some way.
Consider U.S. Pat. No. 8,954,365 to Criminisi et al. titled âDensity Estimation and/or Manifold Learningâ, filed Jun. 21, 2012. Rather than focusing on methods of combining models, Criminisi focuses on simplifying a data set. Criminisi describes a dimensional reduction technique that maps unlabeled data to a lower dimensional space whilst preserving relative distances or other relationships among the unlabeled data points. While useful in reducing computational efforts, such techniques fail to address how to combine models that depend on disparate, private data sets.
Yet another example that attempts to address de-identification of data includes U.S. patent application publication 2014/0222349 to Higgins et al. titled âSystem and Methods for Pharmacogenomic Classificationâ filed Jan. 15, 2014. Higgins describes using surrogate phenotypes that represent clusters in pharmacogenomics populations within de-identified absorption, distribution, metabolism, and excretion (ADME) drug data. The surrogate phonotypes are then used to train learning machines (e.g., a support vector machine) that can then be used for classification of live patient data. Although Higgins provides for building trained learning machines based on surrogate phenotypes, Higgins requires access to de-identified data to build the initial training set. As mentioned previously, de-identified data robs a training dataset of some of its value.
In distributed environments where there can be many entities housing private data, it is not possible to ensure access to large amounts of high quality, de-identified data. This is especially true when a new learning task is launched and no data yet exists that can service the new task. Thus, there remains a considerable need for learning systems that are able to aggregate learned information or knowledge from private data sets in a distributed environment without requiring de-identification of the data before training begins.
All publications identified herein are incorporated by reference to the same extent as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference. Where a definition or use of a term in an incorporated reference is inconsistent or contrary to the definition of that term provided herein, the definition of that term provided herein applies and the definition of that term in the reference does not apply.
In some embodiments, the numbers expressing quantities of ingredients, properties such as concentration, reaction conditions, and so forth, used to describe and claim certain embodiments of the inventive subject matter are to be understood as being modified in some instances by the term âabout.â Accordingly, in some embodiments, the numerical parameters set forth in the written description and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the inventive subject matter are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values presented in some embodiments of the inventive subject matter may contain certain errors necessarily resulting from the standard deviation found in their respective testing measurements.
Unless the context dictates the contrary, all ranges set forth herein should be interpreted as being inclusive of their endpoints and open-ended ranges should be interpreted to include only commercially practical values. Similarly, all lists of values should be considered as inclusive of intermediate values unless the context indicates the contrary.
As used in the description herein and throughout the claims that follow, the meaning of âa,â âan,â and âtheâ includes plural reference unless the context clearly dictates otherwise. Also, as used in the description herein, the meaning of âinâ includes âinâ and âonâ unless the context clearly dictates otherwise.
The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., âsuch asâ) provided with respect to certain embodiments herein is intended merely to better illuminate the inventive subject matter and does not pose a limitation on the scope of the inventive subject matter otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the inventive subject matter.
Groupings of alternative elements or embodiments of the inventive subject matter disclosed herein are not to be construed as limitations. Each group member can be referred to and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group can be included in, or deleted from, a group for reasons of convenience and/or patentability. When any such inclusion or deletion occurs, the specification is herein deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.
SUMMARY
The inventive subject matter provides apparatus, systems, and methods in which distributed, on-line machine learning computers are able to learn information or gain knowledge from private data and distribute the knowledge among peers lacking access to the private data, wherein the distributed knowledge does not include the actual private or restricted features of the local, private data.
For the purposes of this application, it is understood that the term âmachine learningâ refers to artificial intelligence systems configured to learn from data without being explicitly programmed. Such systems are understood to be necessarily rooted in computer technology, and in fact, cannot be implemented or even exist in the absence of computing technology. While machine learning systems utilize various types of statistical analyses, machine learning systems are distinguished from statistical analyses by virtue of the ability to learn without explicit programming and being rooted in computer technology. Thus, the present techniques utilize a distributed data structure that preserves privacy rights while also retaining learnability. Protocols that exchange compressed/learned data, as opposed to raw data, reduces bandwidth overhead.
One aspect of the inventive subject matter includes a distributed machine learning system. In some embodiments, the distributed machine learning system has a plurality of private data servers, possibly operating as peers in a distributed computing environment. Each private data server has access to its own local, private data. The other servers or peers in the system typically lack permission, authority, privilege, or access to others local, private data. Further, each private data server is communicatively coupled with one or more non-private computing devices comprising a global modeling engine; a centralized machine learning computer farm or a different private data server for example. The private data servers are computing devices having one or more processors that are configurable to execute software instructions stored in a non-transitory computer readable memory, where execution of the software instructions gives rise to a modeling engine on the private data server. The modeling engine is configurable to generate one or more trained machine learning models based on the local private data. More specifically, the modeling engine is able to receive model instructions from one or more remote computing devices over a network. The model instructions can be considered as one or more command that instruct the modeling engine to use at least some of the local private data in order to create a trained actual model according to an implementation of a machine learning algorithm (e.g., support vector machine, neural network, decision tree, random forest, deep learning neural network, etc.). The modeling engine creates the trained actual model as a function of the local private data (i.e., a selected or filtered training data set) after any required preprocessing requirements, if any, have been met (e.g., filtering, validating, normalizing, etc.). Once trained, the trained actual model will have one or more actual model parameters or metrics that describe the nature of the trained actual model (e.g., accuracy, accuracy gain, sensitivity, sensitivity gain, performance metrics, weights, learning rate, epochs, kernels, number of nodes, number of layers, etc.). The modeling engine further generates one or more private data distributions from the local private data training set where the private data distributions represent the nature of the local private data used to create the trained model. The modeling engine uses the private data distributions to generate a set of proxy data, which can be considered synthetic data or Monte Carlo data having the same general data distribution characteristics as the local private data, while also lacking the actual private or restricted features of the local, private data. In some cases, Monte Carlo simulations generate deterministic sets of proxy data, by using a seed for a pseudo random number generator. A source for truly random seeds includes those provided by random.org (see URL www.random.org). Private or restricted features of the local private data include, but are not limited to, social security numbers, patient names, addresses or any other personally identifying information, especially information protected under the HIPAA Act. The modeling engine then attempts to validate that the set of proxy data is a reasonable training set stand-in for the local, private data by creating a trained proxy model from the set of proxy data. The resulting trained proxy model is described by one or more proxy model parameters defined according to the same attribute space as the actual model parameters. The modeling engine calculates a similarity score that indicates how similar the trained actual model and the proxy model are to each other as a function of the proxy model parameters and the actual model parameters. Based on the similarity score, the modeling engine can transmit one or more pieces of information related to the trained model, possibly including the set of proxy data or information sufficient to recreate the proxy data, actual model parameters, proxy model parameters, or other features. For example, if the model similarity satisfies a similarity requirement (e.g., compared to a threshold value, etc.), the modeling engine can transmit the set of proxy data to a non-private computing device, which in turn integrates the proxy data in to an aggregated model.
Another aspect of the inventive subject matter includes computer implemented methods of distributed machine learning that respect private data. One embodiment of a method includes a private data server receiving model instructions to create a trained actual model based on at least some local, private data. The model instructions, for example, can include a request to build the trained actual model from an implementation of a machine learning algorithm. A machine learning engine, possibly executing on the private data server, continues by creating the trained actual model according to the model instructions by training the implementation of the machine learning algorithm(s) on relevant local, private data. The resulting trained model comprises one or more actual model parameters that describe the nature of the trained model. Another step of the method includes generating one or more private data distributions that describe the nature of the relevant local, private data. For example, the private data distributions could be represented by a Gaussian distribution, a Poisson distribution, a histogram, a probability distribution, or another type of distribution. From the private data distributions, the machine learning engine can identify or otherwise calculate one or more salient private data features that describe the nature of the private data distributions. Depending upon the type of distribution, example features could include sample data, a mean, a mode, an average, a width, a half-life, a slope, a moment, a histogram, higher order moments, or other types of features. In some, more specific embodiments, the salient private data features could include pr
CROSS REFERENCE TO RELATED APPLICATION
This application claims the priority under 35 USC 119 from U.S. Provisional Patent Application Ser. 62/363,697, entitled Distributed Machine Learning Systems, Apparatus, and Methods, filed on Jul. 18, 2016 by Szeto, the contents of which are incorporated by reference in their entirety.
FIELD OF THE INVENTION
The field of the invention is distributed machine learning technologies.
BACKGROUND
The background description includes information that may be useful in understanding the present inventive subject matter. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed inventive subject matter, or that any publication specifically or implicitly referenced is prior art.
With the recent growth of highly accessible and cost-effective machine learning platforms (e.g., Google®'s Artificial Intelligence including TensorFlow, Amazon's Machine Learning, Microsoft's Azure Machine Learning, OpenAI, SciKit-Learn, Matlab, etc.), data analysts have numerous off-the-shelf options available to them for conducting automated analysis of large data sets. Additionally, in parallel to the growth of machine learning platforms, target data sets have also grown in size. For example, Yahoo! has released several large data sets to the public having sizes on the order of terabytes and The Cancer Genome Atlas (TCGA) data portal offers access to massive amounts of clinical information and genomic characterization data. These pre-built data sets are made readily available to data analysts.
Unfortunately, researchers often encounter obstacles when compiling data sets for their in-progress research, especially when attempting to build trained machine learning models capable of generating interesting predictions using in-the-field data. One major obstacle is that researchers often lack access to the data they require. Consider, for example, a scenario where a researcher wishes to build a trained model from patient data where the patient data is stored in multiple hospitals' electronic medical record databases. The researcher would likely not have authorization to access each hospital's patient data due to privacy restrictions or HIPAA compliance. In order to compile a desired data set, the researcher must request the data from the hospital. Assuming the hospital is amenable to the request, the hospital must then de-identify the data to remove references to specific patients before providing the data to the researcher. However, de-identification results in loss of possibly valuable information in the dataset that could be instrumental in training machine learning algorithms, which in turn can provide opportunities for discovering new relationships in the data or provide value predictive properties. Thus, because of the security restrictions, the datasets available to the researcher could lack information. Clearly, researchers would benefit from technologies that could extract learned information or âknowledgeâ while also respecting private or secured information distributed across multiple data stores.
Interestingly, previous efforts associated with analyzing distributed data focus on the nature of machine learning rather than dealing with the technical issues of isolated, private data. For example, U.S. Pat. No. 7,899,225 to Collins et al. titled âSystems and Methods of Clinical State Prediction Utilizing Medical Image Dataâ filed Oct. 26, 2006, describes creating and merging statistical models to create a final multi-dimensional classification space. The statistical models are the mathematical variation models which define the space in which subjects can be represented. Unfortunately, Collins assumes that the system has authorization to access all the data in order to build the predictive models. Collins also fails to provide insights into circumstances where non-centralized data must remain secure or private. Still, it would be useful to be able to combine trained models in some way.
Consider U.S. Pat. No. 8,954,365 to Criminisi et al. titled âDensity Estimation and/or Manifold Learningâ, filed Jun. 21, 2012. Rather than focusing on methods of combining models, Criminisi focuses on simplifying a data set. Criminisi describes a dimensional reduction technique that maps unlabeled data to a lower dimensional space whilst preserving relative distances or other relationships among the unlabeled data points. While useful in reducing computational efforts, such techniques fail to address how to combine models that depend on disparate, private data sets.
Yet another example that attempts to address de-identification of data includes U.S. patent application publication 2014/0222349 to Higgins et al. titled âSystem and Methods for Pharmacogenomic Classificationâ filed Jan. 15, 2014. Higgins describes using surrogate phenotypes that represent clusters in pharmacogenomics populations within de-identified absorption, distribution, metabolism, and excretion (ADME) drug data. The surrogate phonotypes are then used to train learning machines (e.g., a support vector machine) that can then be used for classification of live patient data. Although Higgins provides for building trained learning machines based on surrogate phenotypes, Higgins requires access to de-identified data to build the initial training set. As mentioned previously, de-identified data robs a training dataset of some of its value.
In distributed environments where there can be many entities housing private data, it is not possible to ensure access to large amounts of high quality, de-identified data. This is especially true when a new learning task is launched and no data yet exists that can service the new task. Thus, there remains a considerable need for learning systems that are able to aggregate learned information or knowledge from private data sets in a distributed environment without requiring de-identification of the data before training begins.
All publications identified herein are incorporated by reference to the same extent as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference. Where a definition or use of a term in an incorporated reference is inconsistent or contrary to the definition of that term provided herein, the definition of that term provided herein applies and the definition of that term in the reference does not apply.
In some embodiments, the numbers expressing quantities of ingredients, properties such as concentration, reaction conditions, and so forth, used to describe and claim certain embodiments of the inventive subject matter are to be understood as being modified in some instances by the term âabout.â Accordingly, in some embodiments, the numerical parameters set forth in the written description and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the inventive subject matter are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable. The numerical values presented in some embodiments of the inventive subject matter may contain certain errors necessarily resulting from the standard deviation found in their respective testing measurements.
Unless the context dictates the contrary, all ranges set forth herein should be interpreted as being inclusive of their endpoints and open-ended ranges should be interpreted to include only commercially practical values. Similarly, all lists of values should be considered as inclusive of intermediate values unless the context indicates the contrary.
As used in the description herein and throughout the claims that follow, the meaning of âa,â âan,â and âtheâ includes plural reference unless the context clearly dictates otherwise. Also, as used in the description herein, the meaning of âinâ includes âinâ and âonâ unless the context clearly dictates otherwise.
The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., âsuch asâ) provided with respect to certain embodiments herein is intended merely to better illuminate the inventive subject matter and does not pose a limitation on the scope of the inventive subject matter otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the inventive subject matter.
Groupings of alternative elements or embodiments of the inventive subject matter disclosed herein are not to be construed as limitations. Each group member can be referred to and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group can be included in, or deleted from, a group for reasons of convenience and/or patentability. When any such inclusion or deletion occurs, the specification is herein deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.
SUMMARY
The inventive subject matter provides apparatus, systems, and methods in which distributed, on-line machine learning computers are able to learn information or gain knowledge from private data and distribute the knowledge among peers lacking access to the private data, wherein the distributed knowledge does not include the actual private or restricted features of the local, private data.
For the purposes of this application, it is understood that the term âmachine learningâ refers to artificial intelligence systems configured to learn from data without being explicitly programmed. Such systems are understood to be necessarily rooted in computer technology, and in fact, cannot be implemented or even exist in the absence of computing technology. While machine learning systems utilize various types of statistical analyses, machine learning systems are distinguished from statistical analyses by virtue of the ability to learn without explicit programming and being rooted in computer technology. Thus, the present techniques utilize a distributed data structure that preserves privacy rights while also retaining learnability. Protocols that exchange compressed/learned data, as opposed to raw data, reduces bandwidth overhead.
One aspect of the inventive subject matter includes a distributed machine learning system. In some embodiments, the distributed machine learning system has a plurality of private data servers, possibly operating as peers in a distributed computing environment. Each private data server has access to its own local, private data. The other servers or peers in the system typically lack permission, authority, privilege, or access to others local, private data. Further, each private data server is communicatively coupled with one or more non-private computing devices comprising a global modeling engine; a centralized machine learning computer farm or a different private data server for example. The private data servers are computing devices having one or more processors that are configurable to execute software instructions stored in a non-transitory computer readable memory, where execution of the software instructions gives rise to a modeling engine on the private data server. The modeling engine is configurable to generate one or more trained machine learning models based on the local private data. More specifically, the modeling engine is able to receive model instructions from one or more remote computing devices over a network. The model instructions can be considered as one or more command that instruct the modeling engine to use at least some of the local private data in order to create a trained actual model according to an implementation of a machine learning algorithm (e.g., support vector machine, neural network, decision tree, random forest, deep learning neural network, etc.). The modeling engine creates the trained actual model as a function of the local private data (i.e., a selected or filtered training data set) after any required preprocessing requirements, if any, have been met (e.g., filtering, validating, normalizing, etc.). Once trained, the trained actual model will have one or more actual model parameters or metrics that describe the nature of the trained actual model (e.g., accuracy, accuracy gain, sensitivity, sensitivity gain, performance metrics, weights, learning rate, epochs, kernels, number of nodes, number of layers, etc.). The modeling engine further generates one or more private data distributions from the local private data training set where the private data distributions represent the nature of the local private data used to create the trained model. The modeling engine uses the private data distributions to generate a set of proxy data, which can be considered synthetic data or Monte Carlo data having the same general data distribution characteristics as the local private data, while also lacking the actual private or restricted features of the local, private data. In some cases, Monte Carlo simulations generate deterministic sets of proxy data, by using a seed for a pseudo random number generator. A source for truly random seeds includes those provided by random.org (see URL www.random.org). Private or restricted features of the local private data include, but are not limited to, social security numbers, patient names, addresses or any other personally identifying information, especially information protected under the HIPAA Act. The modeling engine then attempts to validate that the set of proxy data is a reasonable training set stand-in for the local, private data by creating a trained proxy model from the set of proxy data. The resulting trained proxy model is described by one or more proxy model parameters defined according to the same attribute space as the actual model parameters. The modeling engine calculates a similarity score that indicates how similar the trained actual model and the proxy model are to each other as a function of the proxy model parameters and the actual model parameters. Based on the similarity score, the modeling engine can transmit one or more pieces of information related to the trained model, possibly including the set of proxy data or information sufficient to recreate the proxy data, actual model parameters, proxy model parameters, or other features. For example, if the model similarity satisfies a similarity requirement (e.g., compared to a threshold value, etc.), the modeling engine can transmit the set of proxy data to a non-private computing device, which in turn integrates the proxy data in to an aggregated model.
Another aspect of the inventive subject matter includes computer implemented methods of distributed machine learning that respect private data. One embodiment of a method includes a private data server receiving model instructions to create a trained actual model based on at least some local, private data. The model instructions, for example, can include a request to build the trained actual model from an implementation of a machine learning algorithm. A machine learning engine, possibly executing on the private data server, continues by creating the trained actual model according to the model instructions by training the implementation of the machine learning algorithm(s) on relevant local, private data. The resulting trained model comprises one or more actual model parameters that describe the nature of the trained model. Another step of the method includes generating one or more private data distributions that describe the nature of the relevant local, private data. For example, the private data distributions could be represented by a Gaussian distribution, a Poisson distribution, a histogram, a probability distribution, or another type of distribution. From the private data distributions, the machine learning engine can identify or otherwise calculate one or more salient private data features that describe the nature of the private data distributions. Depending upon the type of distribution, example features could include sample data, a mean, a mode, an average, a width, a half-life, a slope, a moment, a histogram, higher order moments, or other types of features. In some, more specific embodiments, the salient private data features could include proxy data. Once the salient features are available, the machine learning engine transmits the salient private data features over a network to a non-private computing device; a central server or global modeling engine, for example, that can integrate the salient private data features with other data sets to create an aggregated model. Thus, multiple private peers are able to share their learned knowledge without exposing their private data.
Various objects, features, aspects and advantages of the inventive subject matter will become more apparent from the following detailed description of preferred embodiments, along with the accompanying drawing figures in which like numerals represent like components.
BRIEF DESCRIPTION OF THE DRAWING
FIG. 1 is an illustration of an example distributed, online machine learning system, according to the embodiments presented herein.
FIG. 2 is an example machine learning modeling engine architecture deployed within a private data server, according to the embodiments presented herein.
FIG. 3 is a flowchart showing generation of proxy training data in preparation for building a proxy trained model, according to the embodiments presented herein.
FIG. 4 is a flowchart showing generation of one or more similarity scores comparing the similarity of a trained actual model with a trained proxy model, according to the embodiments presented herein.
FIG. 5 is an operational flowchart showing an example method of distributed, online machine learning where private data servers generate proxy data capable of replicating the nature of a trained actual model generated on real data where the proxy data is transmitted to a non-private computing device, according to the embodiments presented herein.
FIG. 6 is an operational flowchart showing an example method of distributed, online machine learning where private data servers transmit salient features of aggregated private data to a non-private computing devices, which in turn creates proxy data for integrating into a trained global model, according to the embodiments presented herein.
DETAILED DESCRIPTION
It should be noted that any language directed to a computer or computing device should be read to include any suitable combination of computing devices, including servers, interfaces, systems, appliances, databases, agents, peers, engines, controllers, modules, or other types of computing devices operating individually, collectively, or cooperatively. One of ordinary skill in the art should appreciate that the computing devices comprise one or more processors configured to execute software instructions that are stored on a tangible, non-transitory computer readable storage medium (e.g., hard drive, FPGA, PLA, PLD, solid state drive, RAM, flash, ROM, external drive, memory stick, etc.). The software instructions specifically configure or program the computing device to provide the roles, responsibilities, or other functionality as discussed below with respect to the disclosed apparatus. Further, the disclosed technologies can be embodied as a computer program product that includes a tangible, non-transitory computer readable medium storing the software instructions executable by a processor to perform the disclosed steps or operations associated with implementations of computer-based algorithms, processes, methods, or other instructions. In some embodiments, the various servers, systems, databases, or interfaces exchange data using standardized protocols or algorithms, possibly based on HTTP, HTTPS, AES, public-private key exchanges, web service APIs, known financial transaction protocols, or other electronic information exchanging methods. Data exchanges among devices can be conducted over a packet-switched network, the Internet, LAN, WAN, VPN, or other type of packet switched network; a circuit switched network; cell switched network; or other type of network.
As used in the description herein and throughout the claims that follow, when a system, engine, server, device, module, or other computing element is described as configured to perform or execute functions on data in a memory, the meaning of âconfigured toâ or âprogrammed toâ is defined as one or more processors or cores of the computing element being programmed by a set of software instructions stored in the memory of the computing element to execute the set of functions on target data or data objects stored in the memory.
One should appreciate that the disclosed techniques provide many advantageous technical effects including construction of communication channels among computing devices over a network to exchange machine learning data while respecting data privacy of the underlying raw data. The computing devices are able to exchange âlearnedâ information or knowledge among each other without comprising privacy. More specifically, rather than transmitting private or secured data to remote computing devices, the disclosed private data servers attempt to âlearnâ information automatically about the local private data via computer-based implementations of one or more machine learning algorithms. The learned information is then exchanged with other computers lacking authorization to access the private data. Further, it should be appreciated that the technical effects include computationally building trained proxy models from distributed, private data and their corresponding data distributions.
The focus of the disclosed inventive subject matter is to enable construction or configuration of a computing device to operate on vast quantities of digital data, beyond the capabilities of a human. Although the digital data typically represents various aspects of patient data, it should be appreciated that the digital data is a representation of one or more digital models of the patients, not âthe patientâ itself. By instantiation of such digital models in the memory of the computing devices, the computing devices are able to manage the digital data or models in a manner that provides utility to a user of the computing device that the user would lack without such a tool, especially within a distributed, online machine learning system. Therefore, the inventive subject matter improves or otherwise optimizes distributed machine learning in environments where the computing devices lack access to private data.
The following discussion provides many example embodiments of the inventive subject matter. Although each embodiment represents a single combination of inventive elements, the inventive subject matter is considered to include all possible combinations of the disclosed elements. Thus if one embodiment comprises elements A, B, and C, and a second embodiment comprises elements B and D, then the inventive subject matter is also considered to include other remaining combinations of A, B, C, or D, even if not explicitly disclosed.
As used herein, and unless the context dictates otherwise, the term âcoupled toâ is intended to include both direct coupling (in which two elements that are coupled to each other contact each other) and indirect coupling (in which at least one additional element is located between the two elements). Therefore, the terms âcoupled toâ and âcoupled withâ are used synonymously.
The following discussion is presented from a health care perspective, and more specifically with respect to building trained machine learning models from genomic sequence data associated with cancer patients. However, it is fully contemplated that the architecture described herein can be adapted to other forms of research beyond oncology and can be leveraged wherever raw data is secured or considered private; insurance data, financial data, social media profile data, human capital data, proprietary experimental data, gaming or gambling data, military data, network traffic data, shopping or marketing data, or other types of data for example.
For example, the techniques presented herein can be used as part of a âlearning as a serviceâ business model. In this type of model, the organization having private data (e.g., healthcare data, genomic data, enterprise data, etc.) may generate machine learning models (e.g., trained actual models, trained proxy models, etc.) and other learned information, and may allow other groups (e.g., start-ups, other institutions, other businesses, etc.) to use these models to analyze their own data or study local data upon payment of a fee. For instance, in a healthcare setting, data collected from patients at particular healthcare institutions could be analyzed to create trained actual models and/or trained proxy models using machine learning. Researchers, data analysts, or other entrepreneurs at a different healthcare institution or company could pay a fee (e.g., one time fee, subscription, etc.) to access the models, e.g., to analyze their own data or to study local data. Thus, in this example, a machine learning model is generated based upon internal data relative to system 100 , and can be used to classify external data relative to system 100 .
In still other embodiments, the organization providing machine learning services could receive fees for analyzing data provided by a 3 rd party. Here, researchers, data analysts, or other entrepreneurs at different healthcare institutions could pay a fee to provide data, similar to the local private data and in a form that could be analyzed separately or could be combined with the local private data, to generate a machine learning model (e.g., a trained actual model or a trained proxy model) along with other learned information that can be used to analyze subsequent sets of data provided by the 3 rd Party. Thus, in this example, a machine learning model is generated based upon external data relative to system 100 , and can be used to classify additional external data relative to system 100 .
Other industries in which these types of âlearning as a serviceâ models could be employed include but are not limited to game data, military data, network traffic/security data, software execution data, simulation data, etc.
Machine learning algorithms create models that form conclusions based upon observed data. For supervised learning, a training dataset is fed into a machine learning algorithm. Here, by providing inputs and known outputs as training data, a machine learning system can create a model based upon this training data. Thus the machine learning algorithm generates a mapping function that maps inputs to an output.
In other embodiments, for unsupervised learning, a dataset is fed into a machine learning system, and the machine learning system analyses the data based upon clustering of data points. In this type of analysis, the underlying structure or distribution of the data is used to generate a model reflecting the distribution or structure of the data. This type of analysis is frequently used to detect similarities (e.g., are two images the same), identify anomalies/outliers, or to detect patterns in a set of data.
Semi-supervised models, a hybrid of the previous two approaches, utilize both supervised and unsupervised models to analyze data.
Machine learning models predict an output (e.g., using classification or regression) based upon inputs (without a known output or answer). Prediction may involve mapping inputs into a category (e.g., analyzing an image to determine whether a characteristic of the image is present). In this type of analysis, the output variable takes the form of a class label, identifying group membership. Thus, this approach can be used to select a category (e.g., based on whether an image contains a specified characteristic).
Regression analysis seeks to minimize error between a regression line and the data points used to generate the line. Here, the output variable take the form of a continuous variable (e.g., a line) to predict a continuous response. Thus, regression can be used to analyze numerical data. These techniques are described more fully below. It should be appreciated that regression analysis can occur in one or more dimensions of relevance according to the research task's requirements.
FIG. 1 is an illustration of an example distributed machine learning system 100 . System 100 is configured as a computer-based research tool allowing multiple researchers or data analysts to create trained machine learning models from many private or secured data sources, to which the researchers would not normally have permission or authority to access. In the example shown, a researcher has permission to access a central machine learning hub represented as non-private computing device 130 , possibly executing as a global modeling engine 136 . Non-private computing device 130 can comprise one or more global model servers (e.g., cloud, SaaS, PaaS, IaaS, LaaS, farm, etc.) that offer distributed machine learning services to the researcher. However, data of interest to the researcher resides on one or more of private data servers 124 A, 124 B, through 124 N (collectively referred to as private data servers 124 ) located at one or more entities 120 A through 120 N over network 115 (e.g., wireless network, an intranet, a cellular network, a packet switched network, an ad-hoc network, the Internet, WAN, VPN, LAN, P2P, etc.). Network 115 can include any combination of the aforementioned networks. The entities can include hospital 120 A, clinic 120 B, through laboratory 120 N (collectively referred to as entities 120 ). Each of entity 120 has access to its own local private data 122 A through 122 N (collectively referred to as private data 122 ), possibly stored on a local storage facility (e.g., a RAID system, a file server, a NAS, a SAN, a network accessible storage device, a storage area network device, a local computer readable memory, a hard disk drive, an optical storage device, a tape drive, a tape library, a solid state disk, etc.). Further, each private data server 124 could include one or more of a BAM server, a SAM server, a GAR server, a BAMBAM server, or even a clinical operating system server. Each of private data server 124 has access to its own local private data 122 and has at least one of modeling engine 126 . For the sake of discussion, each of private data server 120 is considered communicatively coupled, via network 115 , to non-private computing device 130 .
Each set of private data 122 is considered private to its corresponding entity 120 . Under this consideration it should be appreciated that the other entities 120 as well as the researcher accessing the modeling services offered by non-private computing device 130 do not have rights, permissions, or other authorization to access another's private data 122 . For further clarity, the term âprivateâ and ânon-privateâ are relative terms describing the relationship among the various pairs of entities and their corresponding data sets. For example, private data server 124 B in clinic 120 B has access to its local private data 122 B, but does not have access to another's private data, e.g., private data 122 N in laboratory 120 N or private data 122 A in hospital 120 A. In other embodiments, private data server 124 N could be considered as a non-private computing device 130 relative to the other entities. Such a consideration is especially important in embodiments where the various private servers 124 are able to communicate with each other directly over network 115 , possibly in a peer-to-peer fashion or via an affiliation, rather than through a central hub. For example, if a medical institute has multiple locations and/or affiliations, e.g., a main hospital, physician offices, clinics, a secondary hospital, a hospital affiliation, each of these entities could have their own private data 122 , private data server 124 and modeling engine 126 , which may all be visible to each other, but not to a different entity.
Given the nature of the system and the requirements that each of entity 120 must keep its private data 122 secured, the researcher is hard pressed to gain access to the large quantities of high quality data necessary to build desirable trained machine learning models. More specifically, the researcher would have to gain authorization from each entity 120 having private data 122 that is of interest. Further, due to various restrictions (e.g., privacy policies, regulations, HIPAA compliance, etc.), each entity 120 might not be permitted to provide requested data to the researcher. Even under the assumption that the researcher is able to obtain permission from all of entities 120 to obtain their relevant private data 122 , entities 120 would still have to de-identify the data sets. Such de-identification can be problematic due to the time required to de-identify the data and due to loss of information, which can impact the researcher's ability to gain knowledge from training machine learning models.
In the ecosystem/system presented in FIG. 1 , the issues associated with privacy restrictions of private data 122 are addressed by focusing on the knowledge gained from a trained machine learning algorithm rather than the raw data itself. Rather than requesting raw data from each of entity 120 , the researcher is able to define a desired machine learning model that he/she wishes to create. The researcher may interface with system 100 through the non-private computing device 130 ; through one of the private data servers 124 , provided that the researcher has been granted access to the private data server; or through a device external to system 100 that can interface with non-private computing device 130 . The programmatic model instructions on how to create the desired model are then submitted to each relevant private data server 124 , which also has a corresponding modeling engine 126 (i.e., 126 A through 126 N). Each local modeling engine 126 accesses its own local private data 122 and creates local trained models according to model instructions created by the researcher. As each modeling engine 126 gains new learned information, the new knowledge is transmitted back to the researcher at non-private computing device 130 once transmission criteria have been met. The new knowledge can then be aggregated into a trained global model via global modeling engine 136 . Examples of knowledge include (see, e.g., FIG. 2 ) but are not limited to proxy data 260 , trained actual models 240 , trained proxy models 270 , proxy model parameters, model similarity scores, or other types of data that have been de-identified. In some embodiments, the global model server 130 analyzes sets of proxy related information (including for example proxy data 260 , proxy data distributions 362 , proxy model parameters 475 , other proxy related data combined with seeds, etc.) to determine whether the proxy related information from one of private data server 124 has the same shape and/or overall properties as the proxy related data from another private data server 124 , prior to combining such information. Proxy related information that is dissimilar may be flagged for manual review to determine whether the underlying private data distribution set is corrupted, has missing data, or contains a substantial number of outliers. In some embodiments, private patient data considered to be outliers are disregarded and excluded from the techniques disclosed herein. For example, a one-class support vector machine (SVM) could be used to identify outliers that might not be consistent with the core, relevant data. In some embodiments, the one-class SVM is constructed by external peers (e.g., non-private computing device 130 , etc.) based on similar data of interest. The one-class SVM can then transmitted to the private data server 124 . Private data sever 124 can then use the externally generated one-class SVM to ensure that the local data of interest is indeed consistent with external data of interest.
Thus, proxy data may be considered as a transformation of raw data into data of a different form that retains the characteristics of the raw data.
New private data 122 is accessible to private data server 124 on an ongoing basis, e.g., as test results become available, as new diagnoses are made, as new patients are added to the system, etc. For relatively small data sets, proxy data 260 or other proxy related information can be regenerated using all or nearly all of the stored private data. For larger data sets, proxy data can be regenerated using only newly added data. New data may be identified through timestamps, location of storage, geostamping, blockchain hashes, etc.
In other embodiments, new private data is incorporated in real time or in near real time into the machine learning system. Thus, as soon as new private data is available, it can be incorporated into the trained actual models and the trained proxy models. In some embodiments, the machine learning models are updated constantly, e.g., using all available private data (old and newly added private data) or only on newly added private data. Additionally, there is no set timeframe that governs machine learning model updates, and thus, certain machine learning models are updated daily, while other models are updated yearly or even on longer timeframes. This flexibility stands in contrast to traditional machine leaning models which rely on bulk processing of all of the data followed by cycles of training and testing.
In some embodiments, each private data server 124 receives the same programmatic model instructions 230 on how to create a desired model. In other embodiments, a private data server may receive a first set of programmatic model instructions to create a first model, and another private data server may receive a second set of programmatic model instructions to create a second model. Thus, the programmatic model instructions provided to each private data server 124 may be the same or different.
As proxy data 260 is generated and relayed to the global model server 130 , the global model server aggregates the data and generates an updated global model. Once the global model is updated, it can be determined whether the updated global model is an improvement over the previous version of the global model. If the updated global model is an improvement (e.g., the predictive accuracy is improved), new parameters may be provided to the private data servers via the updated model instructions 230 . At the private data server 124 , the performance of the trained actual model (e.g., whether the model improves or worsens) can be evaluated to determine whether the models instructions provided by the updated global model result in an improved trained actual model. Parameters associated with various machine learning model versions may be stored so that earlier machine learning models may be later retrieved, if needed.
In still other embodiments, a private data server 124 may receive proxy related information (including for example proxy data 260 , proxy data distributions 362 , proxy model parameters 475 , other proxy related data combined with seeds, etc.) from a peer private data server (a different private data server 124 ). The private data server may generate models based on its own local private data, or based on both its own local private data and the received proxy related information from a peer private data server. If the predictive accuracy of the combined data sets is improved, then the data sets or learned knowledge are combined.
In some embodiments, the information (e.g., machine learning models including trained proxy models, trained actual models, private data distributions, synthetic/proxy data distributions, actual model parameters, proxy model parameters, similarity scores or any other information generated as part of the machine learning process, etc.) can be geostamped (associated with a location or other identifier indicating where the processing occurred), timestamped, or integrated into a blockchain to archive research (see also US20150332283). Blockchains may be configured as sample-specific audit trails. In this example, the blockchain is instantiated as a single stand-alone chain for a single sample and represents the sample's life cycle or audit trail. Additionally, as the system can continuously receive new data in an asynchronous manner, geostamping can help manage inflow of new information (e.g., for a newly added clinic, all data geostamped as being from the clinic would be incorporated into the machine learning system. It is contemplated that any type of data may be geostamped.
FIG. 2 is an illustration of an example architecture including private data server 224 within an entity 220 with respect to its machine learning activities. The example presented in FIG. 2 illustrates the inventive concepts from the perspective of how private data server 224 interacts with a remote computing device and private data 222 . In more preferred embodiments, private data 222 comprises local private healthcare data, or more specifically includes patient-specific data (e.g., name, SSN, normal WGS, tumor WGS, genomic diff objects, a patient identifier, etc.). Entity 220 typically is an institution having private local raw data and subject to restrictions as discussed above. Example entities include hospitals, labs, clinics, pharmacies, insurance companies, oncologist offices, or other entities having locally stored data. Private data server 224 represents a local server, typically located behind a firewall of the entity 220 . Private data server 224 can be embodied as a computer having one or more processors 297 that are configured to execute software instructions 293 stored in memory 290 . Example servers that can be leveraged for the inventive subject matter include Linux® servers, Windows® servers, or other servers.
The private data server 224 provides access to private data 222 on behalf of the stakeholders of entity 220 . In more preferred embodiments, private data server 224 represents a local cache of specific patient data, especially data sets of large sizes. For example, a patient might be undergoing various treatments for cancer or might be participating in a clinical trial. In such a scenario, the patient's data could include one or more genomic sequence data sets where each data set might include hundreds of gigabytes of data. If there are several patients, the total data set could represent many terabytes or more. Example genomic sequence data sets could include a whole genome sequence (WGS), RNA-seq data, whole exome sequence (WES), proteomic data, differences between tissues (e.g., diseased versus matched normal, tumor versus matched normal, one patient versus another, etc.) or other large data sets. Still further, a patient could have more than one genomic sequence data set on file; a tumor WGS as well as a matched normal WGS. One data sets that is particularly interesting includes genomic differences between a tumor sequence and that of a matched normal sometimes referred to as âgenomic diff objectsâ. Such genomic diff objects and their generation are described more fully in U.S. Pat. Nos. 9,652,587 and 9,646,134 to Sanborn et al., both titled âBAMBAM: Parallel comparative Analysis of High Throughput Sequencing Dataâ and filed May 25, 2011 and Nov. 18, 2011, respectively. Another type of data includes inferred proteomic pathways derived from patient samples as described in U.S. patent application publications 2012/0041683 and 2012/0158391 to Vaske et al. both titled âPathway Recognition Algorithm Using Data Integration on Genomic Models (Paradigm)â, filed on Apr. 29, 2011 and Oct. 26, 2011, respectively.
Providing a local cache of such large data sets via private data server 220 is considered advantageous for multiple reasons. The data sets are of such size that it is prohibitive to obtain such datasets easily on-demand or when immediately required. For example, a full WGS of a patient with a 50Ã read could comprise roughly 150 GB of data. Coupled with a similar WGS of a patient's tumor, the data set could easily be over 300 GB data. Naturally this assumes that there is only a single tumor WGS and a single normal WGS. If there are multiple samples taken at different tumor locations or at different times, the data set could easily exceed a Terabyte of data, just for one patient. The time to download such large datasets or access the datasets remotely far exceeds the urgency required when treating the patient in real-time. Thus, the patient and other stakeholders are best served by having local caches of the patient's data. Still, further is it impracticable to move the data in real-time as the patient moves or otherwise engages with various entities. As an alternative to providing cached data, for large data sets that may not fit within caches, mini Monte Carlo simulations that mimic private data can be used. These types of simulations typically utilize a seed, allowing synthetic private data to be generated with a Monte Carlo simulation in a deterministic fashion, based on parameters of the seed and pseudo random number generators. Once a seed is identified that generates the preferred amount of synthetic private data with minimal modification of the data, the seed can then be provided to any private data server 124 , where it is used to regenerate the synthetic private data using the same pseudo random number generators and other algorithms. Synthetic data may be analyzed to ensure that it does not contain identifying features that should be kept private.
In the example shown, software instructions 293 give rise to the capabilities or functionality of modeling engine 226 . Modeling engine 226 uses private data 222 to train one or more implementations of machine learning algorithms 295 . Example sources of implementations of machine learning algorithms include sci-kit learn, Google®'s Artificial Intelligence including TensorFlowâ¢, OpenAIâ¢, Prediction IOâ¢, Shogunâ¢, WEKA, or Mahoutâ¢, Matlab, Amazon's Machine Learning, Microsoft's Azure Machine Learning, and SciKit-Learn, just to name a few. The various elements depicted within the modeling engine 226 represent the interaction of data and various functional modules within modeling engine 226 . Thus, modeling engine 226 is considered a local agent configured to provide an interface to private data 222 as well as a conduit through which remote researchers over network 215 can create a locally trained model within modeling engine 226 . In a very real sense, modeling engine 226 is a transformation module that converts local, private data 222 to knowledge about the data that can be consumed by external computing devices without comprising privacy. Knowledge can include any information produced by the machine learning system that has been de-identified.
Private data server 224 can take on many different forms. In some embodiments, private data server 224 is a computing appliance integrated within the IT infrastructure of entity 220 , a dedicated server having its own storage system for private data 222 for example. Such an approach is considered advantageous in circumstances where private data 222 relates to large data sets that are targeting specific research projects external to entity 220 . For example, the appliance could store patient data that is highly relevant to government or clinical studies. In other embodiments, private data server 224 can include one or more servers owned by and operated by the IT department of entity 220 where the servers include additional software modeling engine applications that can be deployed on the servers of entity 220 .
In the example shown, private data server 224 is illustrated as a computing device configurable to communicate over network 215 . For the sake of discussion, network 215 is considered the Internet. However, network 215 could also include other forms of networks including VPNs, Intranets, WAN, P2P networks, cellular networks, or other forms of network. Private data server 224 is configurable to use one or more protocols to establish connections with remote devices. Example protocols can be leveraged for such communications include HTTP, HTTPS, SSL, SSH, TCP/IP, UDP/IP, FTP, SCP, WSDL, SOAP, or other types of well-known protocols. It should be appreciated that, although such protocols can be leveraged, it is contemplated that the data exchanged among the devices in the ecosystem/system will be further packaged for easy transport and consumption by the computing devices. For example, the various data elements exchange in the system (e.g., model instructions 230 , proxy data 260 , etc.) can be packaged via one or more markup languages (e.g., XML, YAML, JSON, etc.) or other file formats (e.g., HDF5, etc.).
In some embodiments, private data server 224 will be deployed behind network security infrastructure; a firewall, for example. In such cases, a remote computing device will likely be unable to establish a connection with private data server 224 unless a suitable network address translation (NAT) port has been created in the firewall. However, a more preferable approach is to configure private data server 224 , possibly via modeling engine 226 , to reach out through the firewall and establish a communication link with a central modeling server (e.g., non-private computing device 130 of FIG. 1 ). This approach is advantageous because it does not require modification of the firewall. Still, the communication link can be secured through encryption (e.g., HTTPS, SSL, SSH, AES, etc.).
Modeling engine 226 represents an agent operating within private data server 224 and is configurable to create trained machine learning models. In some embodiments, modeling engine 226 can function within a secured virtual machine or secured container that is dedicated to specific research tasks, which allows multiple, disparate researchers to work in parallel while also ensuring that each researcher's efforts remain secure from each other. For example, modeling engine 226 can be implemented via a Docker® container, where each researcher would have a separate instance of their own modeling engine 226 running on private data server 224 . In other embodiments, the modeling engine 226 can be constructed to process many sessions in parallel, where each session can be implemented as separate threads within the operating system (e.g., Linux, Windows, etc.) of private data server 224 .
Once communication links are established among private data server 224 and one or more remote non-private computing devices, modeling engine 226 is ready to offer its services to outside entities; the researcher, for example. Modeling engine 226 receives one or more of model instructions 230 that instruct modeling engine 226 to create a trained actual model 240 as function of at least some of private data 222 . For example, in some embodiments such as a neural net, inputs and other configuration parameters may be provided by model instructions, and the weights of each input determined by the machine learning system. Trained actual model 240 is a trained machine learning model trained from an implementation of machine learning algorithm 295 . After training is complete, trained actual model 240 comprises one or more trained model parameters 245 .
Modeling engine 226 receives model instructions to create a trained actual model 240 from at least some local private data 222 and according to an implementation of machine learning algorithm 295 . Model instructions 230 represents many possible mechanisms by which modeling engine 226 can be configured to gain knowledge from private data 222 and can comprise a local command generated within entity 220 , a remote command sourced over network 215 , an executable file, a protocol command, a selected command from a menu of options, or other types of instructions. Model instructions 230 can vary widely depending on a desired implementation. In some cases, model instructions 230 can include streamed-lined instructions that inform modeling engine 226 on how to create the desired trained models, possibly in the form of a script (e.g., Python, Ruby, JavaScript, etc.). Further, model instructions can include data filters or data selection criteria that define requirements for desired results sets created from private data 222 as well as which machine learning algorithm 295 is to be used. Consider a scenario where a researcher wishes
CLAIMS
Claims ( 22 )
What is claimed is:
1 . A computer-based distributed machine learning system comprising:
at least one private data server storing local private data and having a local modeling agent; and a non-private data server coupled with the at least one private data server over a network, the non-private data server lacking authorized access to the local private data, the non-private data server comprising at least one processor that, upon execution of software instructions stored in a computer readable memory, performs operations of:
transmitting a definition of a machine learning task to the modeling agent of the at least one private data server, the definition of the machine learning task including machine learning model instructions and local private data features;
enabling the modeling agent of the at least one private data server to generate synthetic data capable of reproducing knowledge gained from execution of the machine learning model instructions on at least some of the local private data having the local private data features, wherein the synthetic data is generated based on at least one data distribution of the local private data which includes multiple patient sample points for different patients, the synthetic data includes at least one synthetic data distribution of synthetic patient data which is different than the at least one data distribution of the local private data, and each of the multiple patient sample points includes at least one a symptom, a test result, a provider name, a diagnosis, a current procedural terminology (CPT) code, an international classification of diseases (ICD) code, or a diagnostic and statistical manual of mental disorders (DSM) code;
receiving from the modeling agent of the at least one private data server, first proxy model data representative of the knowledge gained via the synthetic data;
comparing the first proxy model data with second proxy model data received from a second private data server;
in response to the first proxy model data and the second proxy model data having at least one of a same shape and a same overall property, aggregating a combination of the second proxy model data and the first proxy model data representative of knowledge gained via the synthetic data into a global model corresponding to the machine learning task; and
flagging the second proxy model data in response to the second proxy model data having a different shape than the first proxy model data, to determine whether an underlying private data distribution set of the second proxy model data is corrupted, has missing data, or contains multiple outliers.
2 . The system of claim 1 , wherein the first proxy model data comprises compressed learned data.
3 . The system of claim 1 , wherein the first proxy model data comprises lossy compressed learned data.
4 . The system of claim 1 , wherein the first proxy model data comprises the synthetic data.
5 . The system of claim 1 , wherein the first proxy model data comprises proxy model parameters derived from the synthetic data.
6 . The system of claim 5 , wherein the operations further include duplicating the synthetic data according to the proxy model parameters.
7 . The system of claim 6 , wherein the proxy model parameters include a seed for a deterministic function able to generate the synthetic data.
8 . The system of claim 1 , wherein the first proxy model data comprises a trained proxy model trained on the synthetic data.
9 . The system of claim 1 , wherein the synthetic data comprises Monte Carlo data.
10 . The system of claim 1 , wherein the local private data comprises private patient data.
11 . The system of claim 10 , wherein the private patient data comprises at least one of healthcare data or genomic data.
12 . The system of claim 1 , wherein the local private data includes at least one of insurance data, financial data, social media profile data, human capital data, proprietary experimental data, gaming or gambling data, military data, network traffic data, or shopping or marketing data.
13 . The system of claim 1 , wherein the operations further include paying a fee in exchange for accessing the modeling agent of the at least one private data server.
14 . The system of claim 1 , wherein the machine learning model instructions comprise at least one of supervised machine learning instructions, unsupervised machine learning instructions, or machine learning clustering instructions.
15 . The system of claim 1 , wherein the machine learning model instructions comprise at least one of machine learning regression instructions or machine learning classification instructions.
16 . The system of claim 1 , wherein the operations further include receiving private data metadata about the local private data from the local modeling agent.
17 . The system of claim 16 , wherein the operations further include generating the definition of the machine learning task based on the private data metadata.
18 . The system of claim 17 , wherein the private data metadata comprises attribute space information relating to the local private data.
19 . The system of claim 1 , wherein synthetic data is generated such that similarity score between a trained proxy model trained on the synthetic data and a trained actual model trained on at least some of the local private data, is below a threshold.
20 . The system of claim 19 , wherein the first proxy model data is received upon satisfaction of transmission requirements defined based on the similarity score.
21 . A method of computer-based distributed machine learning, the method comprising:
transmitting, by at least one processor of a non-private data server coupled with at least one private data server over a network, a definition of a machine learning task to a modeling agent of at least one private data server,
the at least one private data server storing local private data and having a local modeling engine,
the non-private data server lacking authorized access to the local private data, and
the definition of the machine learning task including machine learning model instructions and local private data features;
enabling the modeling agent of the at least one private data server to generate synthetic data capable of reproducing knowledge gained from execution of the machine learning model instructions on at least some of the local private data having the local private data features, wherein the synthetic data is generated based on at least one data distribution of the local private data which includes multiple patient sample points for different patients, the synthetic data includes at least one synthetic data distribution of synthetic patient data which is different than the at least one data distribution of the local private data, and each of the multiple patient sample points includes at least one a symptom, a test result, a provider name, a diagnosis, a current procedural terminology (CPT) code, an international classification of diseases (ICD) code, or a diagnostic and statistical manual of mental disorders (DSM) code; receiving from the modeling agent of the at least one private data server, first proxy model data representative of the knowledge gained via the synthetic data; comparing the first proxy model data with second proxy model data received from a second private data server; in response to the first proxy model data and the second proxy model data having at least one of a same shape and a same overall property, aggregating a combination of the second proxy model data and the first proxy model data representative of knowledge gained via the synthetic data into a global model corresponding to the machine learning task; and flagging the second proxy model data in response to the second proxy model data having a different shape than the first proxy model data, to determine whether an underlying private data distribution set of the second proxy model data is corrupted, has missing data, or contains multiple outliers.
22 . A non-transitory computer-readable medium comprising computer-executable instructions configured to, when executed by at least one processor, cause the processor to perform operations including:
transmitting, by at least one processor of a non-private data server coupled with at least one private data server over a network, a definition of a machine learning task to a modeling agent of at least one private data server,
the at least one private data server storing local private data and having a local modeling engine,
the non-private data server lacking authorized access to the local private data, and
the definition of the machine learning task including machine learning model instructions and local private data features;
enabling the modeling agent of the at least one private data server to generate synthetic data capable of reproducing knowledge gained from execution of the machine learning model instructions on at least some of the local private data having the local private data features, wherein the synthetic data is generated based on at least one data distribution of the local private data which includes multiple patient sample points for different patients, the synthetic data includes at least one synthetic data distribution of synthetic patient data which is different than the at least one data distribution of the local private data, and each of the multiple patient sample points includes at least one a symptom, a test result, a provider name, a diagnosis, a current procedural terminology (CPT) code, an international classification of diseases (ICD) code, or a diagnostic and statistical manual of mental disorders (DSM) code; receiving from the modeling agent of the at least one private data server, first proxy model data representative of the knowledge gained via the synthetic data; comparing the first proxy model data with second proxy model data received from a second private data server; in response to the first proxy model data and the second proxy model data having at least one of a same shape and a same overall property, aggregating a combination of the second proxy model data and the first proxy model data representative of knowledge gained via the synthetic data into a global model corresponding to the machine learning task; and flagging the second proxy model data in response to the second proxy model data having a different shape than the first proxy model data, to determine whether an underlying private data distribution set of the second proxy model data is corrupted, has missing data, or contains multiple outliers.
US18/137,812
2016-07-18
2023-04-21
Distributed machine learning systems including generation of synthetic data
Active
2037-07-29
US12518214B2
( en )
Priority Applications (2)
Application Number
Priority Date
Filing Date
Title
US18/137,812
US12518214B2
( en )
2016-07-18
2023-04-21
Distributed machine learning systems including generation of synthetic data
US19/391,021
US20260073306A1
( en )
2016-07-18
2025-11-17
Distributed machine learning systems, apparatus, and methods
Applications Claiming Priority (4)
Application Number
Priority Date
Filing Date
Title
US201662363697P
2016-07-18
2016-07-18
US15/651,345
US11461690B2
( en )
2016-07-18
2017-07-17
Distributed machine learning systems, apparatus, and methods
US17/890,953
US11694122B2
( en )
2016-07-18
2022-08-18
Distributed machine learning systems, apparatus, and methods
US18/137,812
US12518214B2
( en )
2016-07-18
2023-04-21
Distributed machine learning systems including generation of synthetic data
Related Parent Applications (1)
Application Number
Title
Priority Date
Filing Date
US17/890,953
Continuation
US11694122B2
( en )
2016-07-18
2022-08-18
Distributed machine learning systems, apparatus, and methods
Related Child Applications (1)
Application Number
Title
Priority Date
Filing Date
US19/391,021
Continuation
US20260073306A1
( en )
2016-07-18
2025-11-17
Distributed machine learning systems, apparatus, and methods
Publications (2)
Publication Number
Publication Date
US20230267375A1
US20230267375A1 ( en )
2023-08-24
US12518214B2
true
US12518214B2 ( en )
2026-01-06
Family
ID=60940619
Family Applications (4)
Application Number
Title
Priority Date
Filing Date
US15/651,345
Active
2038-03-07
US11461690B2
( en )
2016-07-18
2017-07-17
Distributed machine learning systems, apparatus, and methods
US17/890,953
Active
2037-07-17
US11694122B2
( en )
2016-07-18
2022-08-18
Distributed machine learning systems, apparatus, and methods
US18/137,812
Active
2037-07-29
US12518214B2
( en )
2016-07-18
2023-04-21
Distributed machine learning systems including generation of synthetic data
US19/391,021
Pending
US20260073306A1
( en )
2016-07-18
2025-11-17
Distributed machine learning systems, apparatus, and methods
Family Applications Before (2)
Application Number
Title
Priority Date
Filing Date
US15/651,345
Active
2038-03-07
US11461690B2
( en )
2016-07-18
2017-07-17
Distributed machine learning systems, apparatus, and methods
US17/890,953
Active
2037-07-17
US11694122B2
( en )
2016-07-18
2022-08-18
Distributed machine learning systems, apparatus, and methods
Family Applications After (1)
Application Number
Title
Priority Date
Filing Date
US19/391,021
Pending
US20260073306A1
( en )
2016-07-18
2025-11-17
Distributed machine learning systems, apparatus, and methods
Country Status (12)
Country
Link
US
( 4 )
US11461690B2
( en )
EP
( 1 )
EP3485436A4
( en )
JP
( 1 )
JP2019526851A
( en )
KR
( 1 )
KR20190032433A
( en )
CN
( 1 )
CN109716346A
( en )
AU
( 1 )
AU2017300259A1
( en )
CA
( 1 )
CA3031067A1
( en )
IL
( 1 )
IL264281A
( en )
MX
( 1 )
MX2019000713A
( en )
SG
( 1 )
SG11201900220RA
( en )
TW
( 1 )
TW201812646A
( en )
WO
( 1 )
WO2018017467A1
( en )
Families Citing this family (590)
* Cited by examiner, â Cited by third party
Publication number
Priority date
Publication date
Assignee
Title
US9318108B2
( en )
2010-01-18
2016-04-19
Apple Inc.
Intelligent automated assistant
US8977255B2
( en )
2007-04-03
2015-03-10
Apple Inc.
Method and system for operating a multi-function portable electronic device using voice-activation
US8676904B2
( en )
2008-10-02
2014-03-18
Apple Inc.
Electronic devices with voice command and contextual data processing capabilities
US10614913B2
( en )
*
2010-09-01
2020-04-07
Apixio, Inc.
Systems and methods for coding health records using weighted belief networks
US20220253731A1
( en )
*
2010-11-23
2022-08-11
Values Centered Innovation Enablement Services, Pvt. Ltd.
Dynamic blockchain-based process enablement system (pes)
US10057736B2
( en )
2011-06-03
2018-08-21
Apple Inc.
Active transport based notifications
US10417037B2
( en )
2012-05-15
2019-09-17
Apple Inc.
Systems and methods for integrating third party services with a digital assistant
EP4560630A3
( en )
2013-02-07
2025-08-06
Apple Inc.
Voice trigger for a digital assistant
US10170123B2
( en )
2014-05-30
2019-01-01
Apple Inc.
Intelligent assistant for home automation
US9715875B2
( en )
2014-05-30
2017-07-25
Apple Inc.
Reducing the need for manual start/end-pointing and trigger phrases
US9338493B2
( en )
2014-06-30
2016-05-10
Apple Inc.
Intelligent automated assistant for TV user interactions
WO2016018348A1
( en )
*
2014-07-31
2016-02-04
Hewlett-Packard Development Company, L.P.
Event clusters
US9886953B2
( en )
2015-03-08
2018-02-06
Apple Inc.
Virtual assistant activation
US10460227B2
( en )
2015-05-15
2019-10-29
Apple Inc.
Virtual assistant in a communication session
US10671428B2
( en )
2015-09-08
2020-06-02
Apple Inc.
Distributed personal assistant
US10331312B2
( en )
2015-09-08
2019-06-25
Apple Inc.
Intelligent automated assistant in a media environment
US10747498B2
( en )
2015-09-08
2020-08-18
Apple Inc.
Zero latency digital assistant
US11295506B2
( en )
2015-09-16
2022-04-05
Tmrw Foundation Ip S. Ã R.L.
Chip with game engine and ray trace engine
US11587559B2
( en )
2015-09-30
2023-02-21
Apple Inc.
Intelligent device identification
US10691473B2
( en )
2015-11-06
2020-06-23
Apple Inc.
Intelligent automated assistant in a messaging environment
US10534994B1
( en )
*
2015-11-11
2020-01-14
Cadence Design Systems, Inc.
System and method for hyper-parameter analysis for multi-layer computational structures
US9928230B1
( en )
2016-09-29
2018-03-27
Vignet Incorporated
Variable and dynamic adjustments to electronic forms
US11514289B1
( en )
2016-03-09
2022-11-29
Freenome Holdings, Inc.
Generating machine learning models using genetic data
US12223282B2
( en )
2016-06-09
2025-02-11
Apple Inc.
Intelligent automated assistant in a home environment
US10586535B2
( en )
2016-06-10
2020-03-10
Apple Inc.
Intelligent digital assistant in a multi-tasking environment
DK201670540A1
( en )
2016-06-11
2018-01-08
Apple Inc
Application integration with a digital assistant
US12197817B2
( en )
2016-06-11
2025-01-14
Apple Inc.
Intelligent device arbitration and control
TW201812646A
( en )
2016-07-18
2018-04-01
ç¾ååå¦å¥§ç¾å å ¬å¸
Decentralized machine learning system, decentralized machine learning method, and method of generating substitute data
CN109643347A
( en )
*
2016-08-11
2019-04-16
æ¨ç¹å ¬å¸
Detect scripted or other anomalous interactions with social media platforms
US11196800B2
( en )
*
2016-09-26
2021-12-07
Google Llc
Systems and methods for communication efficient distributed mean estimation
US10769549B2
( en )
*
2016-11-21
2020-09-08
Google Llc
Management and evaluation of machine-learned models based on locally logged data
CN110431551A
( en )
*
2016-12-22
2019-11-08
é¾ç¿æéå ¬å¸
Mixed Data Fingerprints with Principal Component Analysis
WO2018125928A1
( en )
2016-12-29
2018-07-05
DeepScale, Inc.
Multi-channel sensor simulation for autonomous control systems
US11204787B2
( en )
2017-01-09
2021-12-21
Apple Inc.
Application integration with a digital assistant
WO2018131409A1
( en )
*
2017-01-13
2018-07-19
Kddiæ ªå¼ä¼ç¤¾
Information processing method, information processing device, and computer-readable storage medium
CN108303264B
( en )
*
2017-01-13
2020-03-20
åä¸ºææ¯æéå ¬å¸
Cloud-based vehicle fault diagnosis method, device and system
CN107977163B
( en )
*
2017-01-24
2019-09-10
è ¾è®¯ç§æï¼æ·±å³ï¼æéå ¬å¸
Shared data recovery method and device
WO2018176000A1
( en )
2017-03-23
2018-09-27
DeepScale, Inc.
Data synthesis for autonomous control systems
US11263275B1
( en )
*
2017-04-03
2022-03-01
Massachusetts Mutual Life Insurance Company
Systems, devices, and methods for parallelized data structure processing
US11252260B2
( en )
*
2017-04-17
2022-02-15
Petuum Inc
Efficient peer-to-peer architecture for distributed machine learning
US20180322411A1
( en )
*
2017-05-04
2018-11-08
Linkedin Corporation
Automatic evaluation and validation of text mining algorithms
DK180048B1
( en )
2017-05-11
2020-02-04
Apple Inc.
MAINTAINING THE DATA PROTECTION OF PERSONAL INFORMATION
DK201770428A1
( en )
2017-05-12
2019-02-18
Apple Inc.
Low-latency intelligent automated assistant
DK179496B1
( en )
2017-05-12
2019-01-15
Apple Inc.
USER-SPECIFIC Acoustic Models
DK201770411A1
( en )
2017-05-15
2018-12-20
Apple Inc.
MULTI-MODAL INTERFACES
US10303715B2
( en )
2017-05-16
2019-05-28
Apple Inc.
Intelligent automated assistant for media exploration
DK179549B1
( en )
2017-05-16
2019-02-12
Apple Inc.
Far-field extension for digital assistant services
US10671349B2
( en )
2017-07-24
2020-06-02
Tesla, Inc.
Accelerated mathematical engine
US11409692B2
( en )
2017-07-24
2022-08-09
Tesla, Inc.
Vector computational unit
US11893393B2
( en )
2017-07-24
2024-02-06
Tesla, Inc.
Computational array microprocessor system with hardware arbiter managing memory requests
US11157441B2
( en )
2017-07-24
2021-10-26
Tesla, Inc.
Computational array microprocessor system using non-consecutive data formatting
CN110019658B
( en )
*
2017-07-31
2023-01-20
è ¾è®¯ç§æï¼æ·±å³ï¼æéå ¬å¸
Method and related device for generating search term
CN109327421A
( en )
*
2017-08-01
2019-02-12
é¿éå·´å·´é墿§è¡æéå ¬å¸
Data encryption, machine learning model training method, device and electronic device
CA3014813A1
( en )
*
2017-08-21
2019-02-21
Royal Bank Of Canada
System and method for reproducible machine learning
US10311368B2
( en )
*
2017-09-12
2019-06-04
Sas Institute Inc.
Analytic system for graphical interpretability of and improvement of machine learning models
US10713535B2
( en )
*
2017-09-15
2020-07-14
NovuMind Limited
Methods and processes of encrypted deep learning services
US20190087542A1
( en )
*
2017-09-21
2019-03-21
EasyMarkit Software Inc.
System and method for cross-region patient data management and communication
US11869237B2
( en )
*
2017-09-29
2024-01-09
Sony Interactive Entertainment Inc.
Modular hierarchical vision system of an autonomous personal companion
US11341429B1
( en )
*
2017-10-11
2022-05-24
Snap Inc.
Distributed machine learning for improved privacy
EP3696704B1
( en )
*
2017-10-13
2022-07-13
Nippon Telegraph And Telephone Corporation
Synthetic data generation apparatus, method for the same, and program
US20220222752A1
( en )
*
2017-10-16
2022-07-14
Mitchell International, Inc.
Methods for analyzing insurance data and devices thereof
US10909266B2
( en )
*
2017-10-24
2021-02-02
Merck Sharp & Dohme Corp.
Adaptive model for database security and processing
US12231151B1
( en )
*
2017-10-30
2025-02-18
Atombeam Technologies Inc
Federated large codeword model deep learning architecture with homomorphic compression and encryption
EP3704583A4
( en )
*
2017-11-03
2021-08-11
Arizona Board of Regents on behalf of Arizona State University
SYSTEMS AND METHODS FOR PRIORITIZING SOFTWARE VULNERABILITIES FOR CORRECTIONAL PURPOSES
US11354590B2
( en )
*
2017-11-14
2022-06-07
Adobe Inc.
Rule determination for black-box machine-learning models
US11669769B2
( en )
*
2018-12-13
2023-06-06
Diveplane Corporation
Conditioned synthetic data generation in computer-based reasoning systems
US11727286B2
( en )
2018-12-13
2023-08-15
Diveplane Corporation
Identifier contribution allocation in synthetic data generation in computer-based reasoning systems
US11676069B2
( en )
2018-12-13
2023-06-13
Diveplane Corporation
Synthetic data generation using anonymity preservation in computer-based reasoning systems
US11640561B2
( en )
2018-12-13
2023-05-02
Diveplane Corporation
Dataset quality for synthetic data generation in computer-based reasoning systems
JP6649349B2
( en )
*
2017-11-21
2020-02-19
æ ªå¼ä¼ç¤¾ãã¯ããã¯ã»ã¹ãã¼ãã½ãªã¥ã¼ã·ã§ã³ãº
Measurement solution service provision system
US11886779B2
( en )
*
2017-11-27
2024-01-30
Siemens Industry Software Nv
Accelerated simulation setup process using prior knowledge extraction for problem matching
US10810320B2
( en )
*
2017-12-01
2020-10-20
At&T Intellectual Property I, L.P.
Rule based access to voluntarily provided data housed in a protected region of a data storage device
EP3499459A1
( en )
*
2017-12-18
2019-06-19
FEI Company
Method, device and system for remote deep learning for microscopic image reconstruction and segmentation
US10841331B2
( en )
*
2017-12-19
2020-11-17
International Business Machines Corporation
Network quarantine management system
EP3503117B1
( en )
*
2017-12-20
2024-12-18
Nokia Technologies Oy
Updating learned models
EP3503012A1
( en )
*
2017-12-20
2019-06-26
Accenture Global Solutions Limited
Analytics engine for multiple blockchain nodes
US11928716B2
( en )
*
2017-12-20
2024-03-12
Sap Se
Recommendation non-transitory computer-readable medium, method, and system for micro services
US11693989B2
( en )
*
2017-12-21
2023-07-04
Koninklijke Philips N.V.
Computer-implemented methods and nodes implementing performance estimation of algorithms during evaluation of data sets using multiparty computation based random forest
US12307350B2
( en )
2018-01-04
2025-05-20
Tesla, Inc.
Systems and methods for hardware-based pooling
CN111758108A
( en )
2018-01-17
2020-10-09
éå¦ä¹ 人工æºè½è¡ä»½æéå ¬å¸
System and method for modeling probability distributions
US11475350B2
( en )
*
2018-01-22
2022-10-18
Google Llc
Training user-level differentially private machine-learned models
US11656174B2
( en )
2018-01-26
2023-05-23
Viavi Solutions Inc.
Outlier detection for spectroscopic classification
US10810408B2
( en )
2018-01-26
2020-10-20
Viavi Solutions Inc.
Reduced false positive identification for spectroscopic classification
US11009452B2
( en )
*
2018-01-26
2021-05-18
Viavi Solutions Inc.
Reduced false positive identification for spectroscopic quantification
US11561791B2
( en )
2018-02-01
2023-01-24
Tesla, Inc.
Vector computational unit receiving data elements in parallel from a last row of a computational array
GB2571703A
( en )
*
2018-02-07
2019-09-11
Thoughtriver Ltd
A computer system
KR101880175B1
( en )
*
2018-02-13
2018-07-19
주ìíì¬ ë§í¬ë¡ì
Bio-information data providing method, bio-information data storing method and bio-information data transferring system based on multiple block-chain
EP3528179A1
( en )
*
2018-02-15
2019-08-21
Koninklijke Philips N.V.
Training a neural network
EP3528435B1
( en )
*
2018-02-16
2021-03-31
Juniper Networks, Inc.
Automated configuration and data collection during modeling of network devices
US10250381B1
( en )
*
2018-02-22
2019-04-02
Capital One Services, Llc
Content validation using blockchain
US11301951B2
( en )
*
2018-03-15
2022-04-12
The Calany Holding S. Ã R.L.
Game engine and artificial intelligence engine on a chip
US11940958B2
( en )
*
2018-03-15
2024-03-26
International Business Machines Corporation
Artificial intelligence software marketplace
WO2019182590A1
( en )
*
2018-03-21
2019-09-26
Visa International Service Association
Automated machine learning systems and methods
US10818288B2
( en )
2018-03-26
2020-10-27
Apple Inc.
Natural assistant interaction
US11245726B1
( en )
*
2018-04-04
2022-02-08
NortonLifeLock Inc.
Systems and methods for customizing security alert reports
US10707996B2
( en )
*
2018-04-06
2020-07-07
International Business Machines Corporation
Error correcting codes with bayes decoder and optimized codebook
US11262742B2
( en )
*
2018-04-09
2022-03-01
Diveplane Corporation
Anomalous data detection in computer based reasoning and artificial intelligence systems
CN112189206B
( en )
*
2018-04-09
2024-09-06
ç»´è¾¾æ°æ®æ¹æ¡å ¬å¸
Processing of personal data using machine learning algorithms and their applications
US11385633B2
( en )
*
2018-04-09
2022-07-12
Diveplane Corporation
Model reduction and training efficiency in computer-based reasoning and artificial intelligence systems
US10770171B2
( en )
*
2018-04-12
2020-09-08
International Business Machines Corporation
Augmenting datasets using de-identified data and selected authorized records
US11093640B2
( en )
<