ConceptioArchiveGoogle Patents
Google Patentsopen access

Generating a transformed dataset for use by a machine learning model in an … — Pure Storage, Inc. (US11768636B2)

Pure Storage, Inc. · Google Patents
Google Patents · Patents · License: Open Access
Open Source ↗
briangoldpurestorage
patent, google patents, intellectual property, US11768636B2, Pure Storage, Inc., Brian Gold, en, 2023

ABSTRACT

Abstract

Generating a transformed dataset for use by a machine learning model in an artificial intelligence infrastructure that includes one or more storage systems and one or more graphical processing unit (‘GPU’) servers, including: storing, within one or more storage systems, a transformed dataset generated by applying one or more transformations to a dataset that are identified based on one or more expected input formats of data received as input data by one or more machine learning models to be executed on one or more servers; and transmitting, from the one or more storage systems to the one or more servers without reapplying the one or more transformations on the dataset, the transformed dataset including data in the one or more expected formats of data to be received as input data by the one or more machine learning models.

Description

CROSS REFERENCE TO RELATED APPLICATIONS

This is a continuation application for patent entitled to a filing date and claiming the benefit of earlier-filed U.S. patent application Ser. No. 16/888,402, filed May 29, 2020, herein incorporated by reference in its entirety, which is a continuation of U.S. Pat. No. 10,671,435, issued Jun. 2, 2020, which claims priority from U.S. Provisional Patent Application No. 62/574,534, filed Oct. 19, 2017, U.S. Provisional Patent Application No. 62/576,523, filed Oct. 24, 2017, U.S. Provisional Patent Application No. 62/620,286, filed Jan. 22, 2018, U.S. Provisional Patent Application No. 62/648,368, filed Mar. 26, 2018, and U.S. Provisional Patent Application No. 62/650,736, filed Mar. 30, 2018.

BRIEF DESCRIPTION OF DRAWINGS

FIG. 1 A illustrates a first example system for data storage in accordance with some implementations.

FIG. 1 B illustrates a second example system for data storage in accordance with some implementations.

FIG. 1 C illustrates a third example system for data storage in accordance with some implementations.

FIG. 1 D illustrates a fourth example system for data storage in accordance with some implementations.

FIG. 2 A is a perspective view of a storage cluster with multiple storage nodes and internal storage coupled to each storage node to provide network attached storage, in accordance with some embodiments.

FIG. 2 B is a block diagram showing an interconnect switch coupling multiple storage nodes in accordance with some embodiments.

FIG. 2 C is a multiple level block diagram, showing contents of a storage node and contents of one of the non-volatile solid state storage units in accordance with some embodiments.

FIG. 2 D shows a storage server environment, which uses embodiments of the storage nodes and storage units of some previous figures in accordance with some embodiments.

FIG. 2 E is a blade hardware block diagram, showing a control plane, compute and storage planes, and authorities interacting with underlying physical resources, in accordance with some embodiments.

FIG. 2 F depicts elasticity software layers in blades of a storage cluster, in accordance with some embodiments.

FIG. 2 G depicts authorities and storage resources in blades of a storage cluster, in accordance with some embodiments.

FIG. 3 A sets forth a diagram of a storage system that is coupled for data communications with a cloud services provider in accordance with some embodiments of the present disclosure.

FIG. 3 B sets forth a diagram of a storage system in accordance with some embodiments of the present disclosure.

FIG. 4 sets forth a flow chart illustrating an example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 5 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 6 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 7 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 8 A sets forth a diagram illustrating an example computer architecture for implementing an artificial intelligence and machine learning infrastructure configured to fit within a single chassis according to some embodiments of the present disclosure.

FIG. 8 B sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 9 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 10 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 11 A sets forth a diagram illustrating an example artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 11 B sets forth a diagram illustrating an example computer architecture for implementing an artificial intelligence and machine learning infrastructure within a single chassis according to some embodiments of the present disclosure.

FIG. 11 C sets forth a diagram illustrating an example implementation of an artificial intelligence and machine learning infrastructure software stack according to some embodiments of the present disclosure.

FIG. 11 D sets forth a flow chart illustrating an example method for interconnecting a graphical processing unit layer and a storage layer of an artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 12 A sets forth a flow chart illustrating an example method of monitoring an artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 12 B sets forth a flow chart illustrating an example method of optimizing an artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 13 sets forth a flow chart illustrating an example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

FIG. 14 sets forth a flow chart illustrating an additional example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

FIG. 15 sets forth a flow chart illustrating an example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

FIG. 16 sets forth a flow chart illustrating an example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

DESCRIPTION OF EMBODIMENTS

Example methods, apparatuses, and products for data transformation caching in an artificial intelligence infrastructure in accordance with embodiments of the present disclosure are described with reference to the accompanying drawings, beginning with FIG. 1 A . FIG. 1 A illustrates an example system for data storage, in accordance with some implementations. System 100 (also referred to as “storage system” herein) includes numerous elements for purposes of illustration rather than limitation. It may be noted that system 100 may include the same, more, or fewer elements configured in the same or different manner in other implementations.

System 100 includes a number of computing devices 164 A-B. Computing devices (also referred to as “client devices” herein) may be embodied, for example, a server in a data center, a workstation, a personal computer, a notebook, or the like. Computing devices 164 A-B may be coupled for data communications to one or more storage arrays 102 A-B through a storage area network (‘SAN’) 158 or a local area network (‘LAN’) 160 .

The SAN 158 may be implemented with a variety of data communications fabrics, devices, and protocols. For example, the fabrics for SAN 158 may include Fibre Channel, Ethernet, Infiniband, Serial Attached Small Computer System Interface (‘SAS’), or the like. Data communications protocols for use with SAN 158 may include Advanced Technology Attachment (‘ATA’), Fibre Channel Protocol, Small Computer System Interface (‘SCSI’), Internet Small Computer System Interface (‘iSCSI’), HyperSCSI, Non-Volatile Memory Express (‘NVMe’) over Fabrics, or the like. It may be noted that SAN 158 is provided for illustration, rather than limitation. Other data communication couplings may be implemented between computing devices 164 A-B and storage arrays 102 A-B.

The LAN 160 may also be implemented with a variety of fabrics, devices, and protocols. For example, the fabrics for LAN 160 may include Ethernet (802.3), wireless (802.11), or the like. Data communication protocols for use in LAN 160 may include Transmission Control Protocol (‘TCP’), User Datagram Protocol (‘UDP’), Internet Protocol (‘IP’), HyperText Transfer Protocol (‘HTTP’), Wireless Access Protocol (‘WAP’), Handheld Device Transport Protocol (‘HDTP’), Session Initiation Protocol (‘SIP’), Real Time Protocol (‘RTP’), or the like.

Storage arrays 102 A-B may provide persistent data storage for the computing devices 164 A- B. Storage array 102 A may be contained in a chassis (not shown), and storage array 102 B may be contained in another chassis (not shown), in implementations. Storage array

102 A and 102 B may include one or more storage array controllers 110 A-D (also referred to as “controller” herein). A storage array controller 110 A-D may be embodied as a module of automated computing machinery comprising computer hardware, computer software, or a combination of computer hardware and software. In some implementations, the storage array controllers 110 A-D may be configured to carry out various storage tasks. Storage tasks may include writing data received from the computing devices 164 A-B to storage array 102 A-B, erasing data from storage array 102 A-B, retrieving data from <figure-callout id="102A" label="storage array" filenames="US11768636-20230926-D00000.png,US11768636-20230926-D00001.png" state

CROSS REFERENCE TO RELATED APPLICATIONS

This is a continuation application for patent entitled to a filing date and claiming the benefit of earlier-filed U.S. patent application Ser. No. 16/888,402, filed May 29, 2020, herein incorporated by reference in its entirety, which is a continuation of U.S. Pat. No. 10,671,435, issued Jun. 2, 2020, which claims priority from U.S. Provisional Patent Application No. 62/574,534, filed Oct. 19, 2017, U.S. Provisional Patent Application No. 62/576,523, filed Oct. 24, 2017, U.S. Provisional Patent Application No. 62/620,286, filed Jan. 22, 2018, U.S. Provisional Patent Application No. 62/648,368, filed Mar. 26, 2018, and U.S. Provisional Patent Application No. 62/650,736, filed Mar. 30, 2018.

BRIEF DESCRIPTION OF DRAWINGS

FIG. 1 A illustrates a first example system for data storage in accordance with some implementations.

FIG. 1 B illustrates a second example system for data storage in accordance with some implementations.

FIG. 1 C illustrates a third example system for data storage in accordance with some implementations.

FIG. 1 D illustrates a fourth example system for data storage in accordance with some implementations.

FIG. 2 A is a perspective view of a storage cluster with multiple storage nodes and internal storage coupled to each storage node to provide network attached storage, in accordance with some embodiments.

FIG. 2 B is a block diagram showing an interconnect switch coupling multiple storage nodes in accordance with some embodiments.

FIG. 2 C is a multiple level block diagram, showing contents of a storage node and contents of one of the non-volatile solid state storage units in accordance with some embodiments.

FIG. 2 D shows a storage server environment, which uses embodiments of the storage nodes and storage units of some previous figures in accordance with some embodiments.

FIG. 2 E is a blade hardware block diagram, showing a control plane, compute and storage planes, and authorities interacting with underlying physical resources, in accordance with some embodiments.

FIG. 2 F depicts elasticity software layers in blades of a storage cluster, in accordance with some embodiments.

FIG. 2 G depicts authorities and storage resources in blades of a storage cluster, in accordance with some embodiments.

FIG. 3 A sets forth a diagram of a storage system that is coupled for data communications with a cloud services provider in accordance with some embodiments of the present disclosure.

FIG. 3 B sets forth a diagram of a storage system in accordance with some embodiments of the present disclosure.

FIG. 4 sets forth a flow chart illustrating an example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 5 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 6 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 7 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 8 A sets forth a diagram illustrating an example computer architecture for implementing an artificial intelligence and machine learning infrastructure configured to fit within a single chassis according to some embodiments of the present disclosure.

FIG. 8 B sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 9 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 10 sets forth a flow chart illustrating an additional example method for executing a big data analytics pipeline in a storage system that includes compute resources and shared storage resources according to some embodiments of the present disclosure.

FIG. 11 A sets forth a diagram illustrating an example artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 11 B sets forth a diagram illustrating an example computer architecture for implementing an artificial intelligence and machine learning infrastructure within a single chassis according to some embodiments of the present disclosure.

FIG. 11 C sets forth a diagram illustrating an example implementation of an artificial intelligence and machine learning infrastructure software stack according to some embodiments of the present disclosure.

FIG. 11 D sets forth a flow chart illustrating an example method for interconnecting a graphical processing unit layer and a storage layer of an artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 12 A sets forth a flow chart illustrating an example method of monitoring an artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 12 B sets forth a flow chart illustrating an example method of optimizing an artificial intelligence and machine learning infrastructure according to some embodiments of the present disclosure.

FIG. 13 sets forth a flow chart illustrating an example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

FIG. 14 sets forth a flow chart illustrating an additional example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

FIG. 15 sets forth a flow chart illustrating an example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

FIG. 16 sets forth a flow chart illustrating an example method of data transformation caching in an artificial intelligence infrastructure that includes one or more storage systems and one or more GPU servers according to some embodiments of the present disclosure.

DESCRIPTION OF EMBODIMENTS

Example methods, apparatuses, and products for data transformation caching in an artificial intelligence infrastructure in accordance with embodiments of the present disclosure are described with reference to the accompanying drawings, beginning with FIG. 1 A . FIG. 1 A illustrates an example system for data storage, in accordance with some implementations. System 100 (also referred to as “storage system” herein) includes numerous elements for purposes of illustration rather than limitation. It may be noted that system 100 may include the same, more, or fewer elements configured in the same or different manner in other implementations.

System 100 includes a number of computing devices 164 A-B. Computing devices (also referred to as “client devices” herein) may be embodied, for example, a server in a data center, a workstation, a personal computer, a notebook, or the like. Computing devices 164 A-B may be coupled for data communications to one or more storage arrays 102 A-B through a storage area network (‘SAN’) 158 or a local area network (‘LAN’) 160 .

The SAN 158 may be implemented with a variety of data communications fabrics, devices, and protocols. For example, the fabrics for SAN 158 may include Fibre Channel, Ethernet, Infiniband, Serial Attached Small Computer System Interface (‘SAS’), or the like. Data communications protocols for use with SAN 158 may include Advanced Technology Attachment (‘ATA’), Fibre Channel Protocol, Small Computer System Interface (‘SCSI’), Internet Small Computer System Interface (‘iSCSI’), HyperSCSI, Non-Volatile Memory Express (‘NVMe’) over Fabrics, or the like. It may be noted that SAN 158 is provided for illustration, rather than limitation. Other data communication couplings may be implemented between computing devices 164 A-B and storage arrays 102 A-B.

The LAN 160 may also be implemented with a variety of fabrics, devices, and protocols. For example, the fabrics for LAN 160 may include Ethernet (802.3), wireless (802.11), or the like. Data communication protocols for use in LAN 160 may include Transmission Control Protocol (‘TCP’), User Datagram Protocol (‘UDP’), Internet Protocol (‘IP’), HyperText Transfer Protocol (‘HTTP’), Wireless Access Protocol (‘WAP’), Handheld Device Transport Protocol (‘HDTP’), Session Initiation Protocol (‘SIP’), Real Time Protocol (‘RTP’), or the like.

Storage arrays 102 A-B may provide persistent data storage for the computing devices 164 A- B. Storage array 102 A may be contained in a chassis (not shown), and storage array 102 B may be contained in another chassis (not shown), in implementations. Storage array

102 A and 102 B may include one or more storage array controllers 110 A-D (also referred to as “controller” herein). A storage array controller 110 A-D may be embodied as a module of automated computing machinery comprising computer hardware, computer software, or a combination of computer hardware and software. In some implementations, the storage array controllers 110 A-D may be configured to carry out various storage tasks. Storage tasks may include writing data received from the computing devices 164 A-B to storage array 102 A-B, erasing data from storage array 102 A-B, retrieving data from storage array 102 A-B and providing data to computing devices 164 A-B, monitoring and reporting of disk utilization and performance, performing redundancy operations, such as Redundant Array of Independent Drives (‘RAID’) or RAID-like data redundancy operations, compressing data, encrypting data, and so forth.

Storage array controller 110 A-D may be implemented in a variety of ways, including as a Field Programmable Gate Array (‘FPGA’), a Programmable Logic Chip (‘PLC’), an Application Specific Integrated Circuit (‘ASIC’), System-on-Chip (‘SOC’), or any computing device that includes discrete components such as a processing device, central processing unit, computer memory, or various adapters. Storage array controller 110 A-D may include, for example, a data communications adapter configured to support communications via the SAN 158 or LAN 160 . In some implementations, storage array controller 110 A-D may be independently coupled to the LAN 160 . In implementations, storage array controller 110 A-D may include an I/O controller or the like that couples the storage array controller 110 A-D for data communications, through a midplane (not shown), to a persistent storage resource 170 A-B (also referred to as a “storage resource” herein). The persistent storage resource 170 A-B main include any number of storage drives 171 A-F (also referred to as “storage devices” herein) and any number of non-volatile Random Access Memory (‘NVRAM’) devices (not shown).

In some implementations, the NVRAM devices of a persistent storage resource 170 A-B may be configured to receive, from the storage array controller 110 A-D, data to be stored in the storage drives 171 A-F. In some examples, the data may originate from computing devices 164 A-B. In some examples, writing data to the NVRAM device may be carried out more quickly than directly writing data to the storage drive 171 A-F. In implementations, the storage array controller 110 A-D may be configured to utilize the NVRAM devices as a quickly accessible buffer for data destined to be written to the storage drives 171 A-F. Latency for write requests using NVRAM devices as a buffer may be improved relative to a system in which a storage array controller 110 A-D writes data directly to the storage drives 171 A-F. In some implementations, the NVRAM devices may be implemented with computer memory in the form of high bandwidth, low latency RAM. The NVRAM device is referred to as “non-volatile” because the NVRAM device may receive or include a unique power source that maintains the state of the RAM after main power loss to the NVRAM device. Such a power source may be a battery, one or more capacitors, or the like. In response to a power loss, the NVRAM device may be configured to write the contents of the RAM to a persistent storage, such as the storage drives 171 A-F.

In implementations, storage drive 171 A-F may refer to any device configured to record data persistently, where “persistently” or “persistent” refers to a device&#39;s ability to maintain recorded data after loss of power. In some implementations, storage drive 171 A-F may correspond to non-disk storage media. For example, the storage drive 171 A-F may be one or more solid-state drives (‘SSDs’), flash memory based storage, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other implementations, storage drive 171 A-F may include mechanical or spinning hard disk, such as hard-disk drives (‘HDD’).

In some implementations, the storage array controllers 110 A-D may be configured for offloading device management responsibilities from storage drive 171 A-F in storage array 102 A-B. For example, storage array controllers 110 A-D may manage control information that may describe the state of one or more memory blocks in the storage drives 171 A-F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code for a storage array controller 110 A-D, the number of program-erase (‘TIE’) cycles that have been performed on a particular memory block, the age of data stored in a particular memory block, the type of data that is stored in a particular memory block, and so forth. In some implementations, the control information may be stored with an associated memory block as metadata. In other implementations, the control information for the storage drives 171 A-F may be stored in one or more particular memory blocks of the storage drives 171 A-F that are selected by the storage array controller 110 A-D. The selected memory blocks may be tagged with an identifier indicating that the selected memory block contains control information. The identifier may be utilized by the storage array controllers 110 A-D in conjunction with storage drives 171 A-F to quickly identify the memory blocks that contain control information. For example, the storage controllers 110 A-D may issue a command to locate memory blocks that contain control information. It may be noted that control information may be so large that parts of the control information may be stored in multiple locations, that the control information may be stored in multiple locations for purposes of redundancy, for example, or that the control information may otherwise be distributed across multiple memory blocks in the storage drive 171 A-F.

In implementations, storage array controllers 110 A-D may offload device management responsibilities from storage drives 171 A-F of storage array 102 A-B by retrieving, from the storage drives 171 A-F, control information describing the state of one or more memory blocks in the storage drives 171 A-F. Retrieving the control information from the storage drives 171 A-F may be carried out, for example, by the storage array controller 110 A-D querying the storage drives 171 A-F for the location of control information for a particular storage drive 171 A-F. The storage drives 171 A-F may be configured to execute instructions that enable the storage drive 171 A-F to identify the location of the control information. The instructions may be executed by a controller (not shown) associated with or otherwise located on the storage drive 171 A-F and may cause the storage drive 171 A-F to scan a portion of each memory block to identify the memory blocks that store control information for the storage drives 171 A-F. The storage drives 171 A-F may respond by sending a response message to the storage array controller 110 A-D that includes the location of control information for the storage drive 171 A-F. Responsive to receiving the response message, storage array controllers 110 A-D may issue a request to read data stored at the address associated with the location of control information for the storage drives 171 A-F.

In other implementations, the storage array controllers 110 A-D may further offload device management responsibilities from storage drives 171 A-F by performing, in response to receiving the control information, a storage drive management operation. A storage drive management operation may include, for example, an operation that is typically performed by the storage drive 171 A-F (e.g., the controller (not shown) associated with a particular storage drive 171 A-F). A storage drive management operation may include, for example, ensuring that data is not written to failed memory blocks within the storage drive 171 A-F, ensuring that data is written to memory blocks within the storage drive 171 A-F in such a way that adequate wear leveling is achieved, and so forth.

In implementations, storage array 102 A-B may implement two or more storage array controllers 110 A-D. For example, storage array 102 A may include storage array controllers 110 A and storage array controllers 110 B. At a given instance, a single storage array controller 110 A-D (e.g., storage array controller 110 A) of a storage system 100 may be designated with primary status (also referred to as “primary controller” herein), and other storage array controllers 110 A-D (e.g., storage array controller 110 B) may be designated with secondary status (also referred to as “secondary controller” herein). The primary controller may have particular rights, such as permission to alter data in persistent storage resource 170 A-B (e.g., writing data to persistent storage resource 170 A-B). At least some of the rights of the primary controller may supersede the rights of the secondary controller. For instance, the secondary controller may not have permission to alter data in persistent storage resource 170 A-B when the primary controller has the right. The status of storage array controllers 110 A-D may change. For example, storage array controller 110 A may be designated with secondary status, and storage array controller 110 B may be designated with primary status.

In some implementations, a primary controller, such as storage array controller 110 A, may serve as the primary controller for one or more storage arrays 102 A-B, and a second controller, such as storage array controller 110 B, may serve as the secondary controller for the one or more storage arrays 102 A-B. For example, storage array controller 110 A may be the primary controller for storage array 102 A and storage array 102 B, and storage array controller 110 B may be the secondary controller for storage array

102 A and 102 B. In some implementations, storage array controllers

110 C and 110 D (also referred to as “storage processing modules”) may neither have primary or secondary status. Storage array controllers

110 C and 110 D, implemented as storage processing modules, may act as a communication interface between the primary and secondary controllers (e.g., storage array controllers

110 A and 110 B, respectively) and storage array 102 B. For example, storage array controller 110 A of storage array 102 A may send a write request, via SAN 158 , to storage array 102 B. The write request may be received by both storage array controllers

110 C and 110 D of storage array 102 B. Storage array controllers

110 C and 110 D facilitate the communication, e.g., send the write request to the appropriate storage drive 171 A-F. It may be noted that in some implementations storage processing modules may be used to increase the number of storage drives controlled by the primary and secondary controllers.

In implementations, storage array controllers 110 A-D are communicatively coupled, via a midplane (not shown), to one or more storage drives 171 A-F and to one or more NVRAM devices (not shown) that are included as part of a storage array 102 A-B. The storage array controllers 110 A-D may be coupled to the midplane via one or more data communication links and the midplane may be coupled to the storage drives 171 A-F and the NVRAM devices via one or more data communications links. The data communications links described herein are collectively illustrated by data communications links 108 A-D and may include a Peripheral Component Interconnect Express (‘PCIe’) bus, for example.

FIG. 1 B illustrates an example system for data storage, in accordance with some implementations. Storage array controller 101 illustrated in FIG. 1 B may be similar to the storage array controllers 110 A-D described with respect to FIG. 1 A . In one example, storage array controller 101 may be similar to storage array controller 110 A or storage array controller 110 B. Storage array controller 101 includes numerous elements for purposes of illustration rather than limitation. It may be noted that storage array controller 101 may include the same, more, or fewer elements configured in the same or different manner in other implementations. It may be noted that elements of FIG. 1 A may be included below to help illustrate features of storage array controller 101 .

Storage array controller 101 may include one or more processing devices 104 and random access memory (‘RAM’) 111 . Processing device 104 (or controller 101 ) represents one or more general-purpose processing devices such as a microprocessor, central processing unit, or the like. More particularly, the processing device 104 (or controller 101 ) may be a complex instruction set computing (‘CISC’) microprocessor, reduced instruction set computing (‘RISC’) microprocessor, very long instruction word (‘VLIW’) microprocessor, or a processor implementing other instruction sets or processors implementing a combination of instruction sets. The processing device 104 (or controller 101 ) may also be one or more special-purpose processing devices such as an application specific integrated circuit (‘ASIC’), a field programmable gate array (‘FPGA’), a digital signal processor (‘DSP’), network processor, or the like.

The processing device 104 may be connected to the RAM 111 via a data communications link 106 , which may be embodied as a high speed memory bus such as a Double-Data Rate 4 (‘DDR4’) bus. Stored in RAM 111 is an operating system 112 . In some implementations, instructions 113 are stored in RAM 111 . Instructions 113 may include computer program instructions for performing operations in a direct-mapped flash storage system. In one embodiment, a direct-mapped flash storage system is one that addresses data blocks within flash drives directly and without an address translation performed by the storage controllers of the flash drives.

In implementations, storage array controller 101 includes one or more host bus adapters 103 A-C that are coupled to the processing device 104 via a data communications link 105 A-C. In implementations, host bus adapters 103 A-C may be computer hardware that connects a host system (e.g., the storage array controller) to other network and storage arrays. In some examples, host bus adapters 103 A-C may be a Fibre Channel adapter that enables the storage array controller 101 to connect to a SAN, an Ethernet adapter that enables the storage array controller 101 to connect to a LAN, or the like. Host bus adapters 103 A-C may be coupled to the processing device 104 via a data communications link 105 A-C such as, for example, a PCIe bus.

In implementations, storage array controller 101 may include a host bus adapter 114 that is coupled to an expander 115 . The expander 115 may be used to attach a host system to a larger number of storage drives. The expander 115 may, for example, be a SAS expander utilized to enable the host bus adapter 114 to attach to storage drives in an implementation where the host bus adapter 114 is embodied as a SAS controller.

In implementations, storage array controller 101 may include a switch 116 coupled to the processing device 104 via a data communications link 109 . The switch 116 may be a computer hardware device that can create multiple endpoints out of a single endpoint, thereby enabling multiple devices to share a single endpoint. The switch 116 may, for example, be a PCIe switch that is coupled to a PCIe bus (e.g., data communications link 109 ) and presents multiple PCIe connection points to the midplane.

In implementations, storage array controller 101 includes a data communications link 107 for coupling the storage array controller 101 to other storage array controllers. In some examples, data communications link 107 may be a QuickPath Interconnect (QPI) interconnect.

A traditional storage system that uses traditional flash drives may implement a process across the flash drives that are part of the traditional storage system. For example, a higher level process of the storage system may initiate and control a process across the flash drives. However, a flash drive of the traditional storage system may include its own storage controller that also performs the process. Thus, for the traditional storage system, a higher level process (e.g., initiated by the storage system) and a lower level process (e.g., initiated by a storage controller of the storage system) may both be performed.

To resolve various deficiencies of a traditional storage system, operations may be performed by higher level processes and not by the lower level processes. For example, the flash storage system may include flash drives that do not include storage controllers that provide the process. Thus, the operating system of the flash storage system itself may initiate and control the process. This may be accomplished by a direct-mapped flash storage system that addresses data blocks within the flash drives directly and without an address translation performed by the storage controllers of the flash drives.

The operating system of the flash storage system may identify and maintain a list of allocation units across multiple flash drives of the flash storage system. The allocation units may be entire erase blocks or multiple erase blocks. The operating system may maintain a map or address range that directly maps addresses to erase blocks of the flash drives of the flash storage system.

Direct mapping to the erase blocks of the flash drives may be used to rewrite data and erase data. For example, the operations may be performed on one or more allocation units that include a first data and a second data where the first data is to be retained and the second data is no longer being used by the flash storage system. The operating system may initiate the process to write the first data to new locations within other allocation units and erasing the second data and marking the allocation units as being available for use for subsequent data. Thus, the process may only be performed by the higher level operating system of the flash storage system without an additional lower level process being performed by controllers of the flash drives.

Advantages of the process being performed only by the operating system of the flash storage system include increased reliability of the flash drives of the flash storage system as unnecessary or redundant write operations are not being performed during the process. One possible point of novelty here is the concept of initiating and controlling the process at the operating system of the flash storage system. In addition, the process can be controlled by the operating system across multiple flash drives. This is in contrast to the process being performed by a storage controller of a flash drive.

A storage system can consist of two storage array controllers that share a set of drives for failover purposes, or it could consist of a single storage array controller that provides a storage service that utilizes multiple drives, or it could consist of a distributed network of storage array controllers each with some number of drives or some amount of Flash storage where the storage array controllers in the network collaborate to provide a complete storage service and collaborate on various aspects of a storage service including storage allocation and garbage collection.

FIG. 1 C illustrates a third example system 117 for data storage in accordance with some implementations. System 117 (also referred to as “storage system” herein) includes numerous elements for purposes of illustration rather than limitation. It may be noted that system 117 may include the same, more, or fewer elements configured in the same or different manner in other implementations.

In one embodiment, system 117 includes a dual Peripheral Component Interconnect (‘PCI’) flash storage device 118 with separately addressable fast write storage. System 117 may include a storage controller 119 . In one embodiment, storage controller 119 A-D may be a CPU, ASIC, FPGA, or any other circuitry that may implement control structures necessary according to the present disclosure. In one embodiment, system 117 includes flash memory devices (e.g., including flash memory devices 120 a - n ), operatively coupled to various channels of the storage device controller 119 . Flash memory devices 120 a - n , may be presented to the controller 119 A-D as an addressable collection of Flash pages, erase blocks, and/or control elements sufficient to allow the storage device controller 119 A-D to program and retrieve various aspects of the Flash. In one embodiment, storage device controller 119 A-D may

CLAIMS

Claims ( 20 )

What is claimed is:

1. A method comprising:

storing, within one or more storage systems, a transformed dataset generated by applying one or more transformations to a dataset that are identified based on one or more expected input formats of data received as input data by one or more machine learning models to be executed on one or more servers; and

transmitting, from the one or more storage systems to the one or more servers without reapplying the one or more transformations on the dataset, the transformed dataset including data in the one or more expected formats of data to be received as input data by the one or more machine learning models.

2. The method of claim 1 further comprising generating the transformed dataset by applying the transformations to an initial version of the dataset.

3. The method of claim 1 wherein transmitting the transformed dataset further comprises transmitting the transformed dataset from the one or more storage systems directly to application memory on at least one of the servers.

4. The method of claim 3 wherein transmitting the transformed dataset from the one or more storage systems directly to application memory on at least one of the servers further comprises transmitting the transformed data dataset from the one or more storage systems to the server via remote direct memory access (‘RDMA’).

5. The method of claim 1 further comprising executing, by one or more of the servers, one or more machine learning algorithms associated with the machine learning model using the transformed dataset as input.

6. The method of claim 1 further comprising:

scheduling, by a unified management plane, one or more transformations for one or more of the storage systems to apply to the dataset; and

scheduling, by the unified management plane, execution of one or more machine learning algorithms associated with the machine learning model by the one or more servers.

7. The method of claim 1 further comprising providing, by a unified management plane to the one or more servers, information describing the dataset, the one or more transformations applied to the dataset, and the transformed dataset.

8. An artificial intelligence infrastructure configured to carry out steps of:

storing, within one or more storage systems, a transformed dataset generated by applying one or more transformations to a dataset that are identified based on one or more expected input formats of data received as input data by one or more machine learning models to be executed on one or more servers; and

transmitting, from the one or more storage systems to the one or more servers without reapplying the one or more transformations on the dataset, the transformed dataset including data in the one or more expected formats of data to be received as input data by the one or more machine learning models.

9. The artificial intelligence infrastructure of claim 8 wherein the artificial intelligence infrastructure is further configured to carry out the step of generating, by at least one of the storage systems, the transformed dataset by applying the transformations to an initial version of the dataset.

10. The artificial intelligence infrastructure of claim 8 wherein transmitting the transformed dataset further comprises transmitting the transformed dataset from the one or more storage systems directly to application memory on at least one of the servers.

11. The artificial intelligence infrastructure of claim 10 wherein transmitting the transformed dataset from the one or more storage systems directly to application memory on the server further comprises transmitting the transformed data dataset from the one or more storage systems to the server via remote direct memory access (‘RDMA’).

12. The artificial intelligence infrastructure of claim 8 wherein the artificial intelligence infrastructure is further configured to carry out the step of executing, by one or more of the servers, one or more machine learning algorithms associated with the machine learning model using the transformed dataset as input.

13. The artificial intelligence infrastructure of claim 8 wherein the artificial intelligence infrastructure is further configured to carry out the steps of:

scheduling, by a unified management plane, one or more transformations for one or more of the storage systems to apply to the dataset; and

scheduling, by the unified management plane, execution of one or more machine learning algorithms associated with the machine learning model by the one or more servers.

14. The artificial intelligence infrastructure of claim 8 wherein the artificial intelligence infrastructure is further configured to carry out the step of providing, by a unified management plane to the one or more servers, information describing the dataset, the one or more transformations applied to the dataset, and the transformed dataset.

15. An apparatus comprising a computer processor, a computer memory operatively coupled to the computer processor, the computer memory having disposed within it computer program instructions that, when executed by the computer processor, cause the apparatus to carry out steps of:

storing, within one or more storage systems, a transformed dataset generated by applying one or more transformations to a dataset that are identified based on one or more expected input formats of data received as input data by one or more machine learning models to be executed on one or more servers; and

transmitting, from the one or more storage systems to the one or more servers without reapplying the one or more transformations on the dataset, the transformed dataset including data in the one or more expected formats of data to be received as input data by the one or more machine learning models.

16. The apparatus of claim 15 wherein the apparatus further comprises computer program instructions that, when executed by the computer processor, cause the apparatus to carry out step of generating the transformed dataset by applying the transformations to an initial version of the dataset.

17. The apparatus of claim 15 wherein transmitting the transformed dataset further comprises transmitting the transformed dataset from the one or more storage systems directly to application memory on at least one of the servers.

18. The apparatus of claim 15 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the steps of:

scheduling, by a unified management plane, one or more transformations for one or more of the storage systems to apply to the dataset; and

scheduling, by the unified management plane, execution of one or more machine learning algorithms associated with the machine learning model by the one or more servers.

19. The apparatus of claim 15 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the step of providing, by a unified management plane to the one or more servers, information describing the dataset, the one or more transformations applied to the dataset, and the transformed dataset.

20. The apparatus of claim 15 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the step of executing, by one or more of the servers, one or more machine learning algorithms associated with the machine learning model using the transformed dataset as input.

US18/146,807

2017-10-19

2022-12-27

Generating a transformed dataset for use by a machine learning model in an artificial intelligence infrastructure

Active

US11768636B2

( en )

Priority Applications (3)

Application Number

Priority Date

Filing Date

Title

US18/146,807

US11768636B2

( en )

2017-10-19

2022-12-27

Generating a transformed dataset for use by a machine learning model in an artificial intelligence infrastructure

US18/465,710

US12455705B2

( en )

2017-10-19

2023-09-12

Optimizing dataset transformations for use by machine learning models

US19/287,505

US20250362836A1

( en )

2017-10-19

2025-07-31

Reuse of transformed datasets in artificial intelligence pipelines via fingerprint-based selection

Applications Claiming Priority (8)

Application Number

Priority Date

Filing Date

Title

US201762574534P

2017-10-19

2017-10-19

US201762576523P

2017-10-24

2017-10-24

US201862620286P

2018-01-22

2018-01-22

US201862648368P

2018-03-26

2018-03-26

US201862650736P

2018-03-30

2018-03-30

US16/040,996

US10671435B1

( en )

2017-10-19

2018-07-20

Data transformation caching in an artificial intelligence infrastructure

US16/888,402

US11556280B2

( en )

2017-10-19

2020-05-29

Data transformation for a machine learning model

US18/146,807

US11768636B2

( en )

2017-10-19

2022-12-27

Generating a transformed dataset for use by a machine learning model in an artificial intelligence infrastructure

Related Parent Applications (1)

Application Number

Title

Priority Date

Filing Date

US16/888,402

Continuation

US11556280B2

( en )

2017-10-19

2020-05-29

Data transformation for a machine learning model

Related Child Applications (2)

Application Number

Title

Priority Date

Filing Date

US18/465,710

Continuation

US12455705B2

( en )

2017-10-19

2023-09-12

Optimizing dataset transformations for use by machine learning models

US19/287,505

Continuation

US20250362836A1

( en )

2017-10-19

2025-07-31

Reuse of transformed datasets in artificial intelligence pipelines via fingerprint-based selection

Publications (2)

Publication Number

Publication Date

US20230126789A1

US20230126789A1 ( en )

2023-04-27

US11768636B2

true

US11768636B2 ( en )

2023-09-26

Family

ID=66169307

Family Applications (13)

Application Number

Title

Priority Date

Filing Date

US16/040,846

Active

2038-12-27

US10671434B1

( en )

2017-10-19

2018-07-20

Storage based artificial intelligence infrastructure

US16/040,996

Active

2039-01-04

US10671435B1

( en )

2017-10-19

2018-07-20

Data transformation caching in an artificial intelligence infrastructure

US16/046,102

Active

US10275176B1

( en )

2017-10-19

2018-07-26

Data transformation offloading in an artificial intelligence infrastructure

US16/046,337

Active

US10275285B1

( en )

2017-10-19

2018-07-26

Data transformation caching in an artificial intelligence infrastructure

US16/047,649

Active

US10649988B1

( en )

2017-10-19

2018-07-27

Artificial intelligence and machine learning infrastructure

US16/888,402

Active

US11556280B2

( en )

2017-10-19

2020-05-29

Data transformation for a machine learning model

US16/888,135

Active

2038-08-11

US11210140B1

( en )

2017-10-19

2020-05-29

Data transformation delegation for a graphical processing unit (‘GPU’) server

US17/538,262

Active

US11803338B2

( en )

2017-10-19

2021-11-30

Executing a machine learning model in an artificial intelligence infrastructure

US18/146,807

Active

US11768636B2

( en )

2017-10-19

2022-12-27

Generating a transformed dataset for use by a machine learning model in an artificial intelligence infrastructure

US18/465,710

Active

US12455705B2

( en )

2017-10-19

2023-09-12

Optimizing dataset transformations for use by machine learning models

US18/497,214

Active

US12517685B2

( en )

2017-10-19

2023-10-30

Executing machine learning models using transformed datasets

US19/287,505

Pending

US20250362836A1

( en )

2017-10-19

2025-07-31

Reuse of transformed datasets in artificial intelligence pipelines via fingerprint-based selection

US19/423,390

Pending

US20260111154A1

( en )

2017-10-19

2025-12-17

Converting unstructured datasets into structured datasets

Family Applications Before (8)

Application Number

Title

Priority Date

Filing Date

US16/040,846

Active

2038-12-27

US10671434B1

( en )

2017-10-19

2018-07-20

Storage based artificial intelligence infrastructure

US16/040,996

Active

2039-01-04

US10671435B1

( en )

2017-10-19

2018-07-20

Data transformation caching in an artificial intelligence infrastructure

US16/046,102

Active

US10275176B1

( en )

2017-10-19

2018-07-26

Data transformation offloading in an artificial intelligence infrastructure

US16/046,337

Active

US10275285B1

( en )

2017-10-19

2018-07-26

Data transformation caching in an artificial intelligence infrastructure

US16/047,649

Active

US10649988B1

( en )

2017-10-19

2018-07-27

Artificial intelligence and machine learning infrastructure

US16/888,402

Active

US11556280B2

( en )

2017-10-19

2020-05-29

Data transformation for a machine learning model

US16/888,135

Active

2038-08-11

US11210140B1

( en )

2017-10-19

2020-05-29

Data transformation delegation for a graphical processing unit (‘GPU’) server

US17/538,262

Active

US11803338B2

( en )

2017-10-19

2021-11-30

Executing a machine learning model in an artificial intelligence infrastructure

Family Applications After (4)

Application Number

Title

Priority Date

Filing Date

US18/465,710

Active

US12455705B2

( en )

2017-10-19

2023-09-12

Optimizing dataset transformations for use by machine learning models

US18/497,214

Active

US12517685B2

( en )

2017-10-19

2023-10-30

Executing machine learning models using transformed datasets

US19/287,505

Pending

US20250362836A1

( en )

2017-10-19

2025-07-31

Reuse of transformed datasets in artificial intelligence pipelines via fingerprint-based selection

US19/423,390

Pending

US20260111154A1

( en )

2017-10-19

2025-12-17

Converting unstructured datasets into structured datasets

Country Status (1)

Country

Link

US

( 13 )

US10671434B1

( en )

Cited By (3)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20220334959A1

( en )

*

2021-04-14

2022-10-20

National Taiwan University

Method and apparatus for generating software test reports

US12061899B2

( en )

*

2021-10-28

2024-08-13

Red Hat, Inc.

Infrastructure as code (IaC) pre-deployment analysis via a machine-learning model

US12413417B2

( en )

2023-11-28

2025-09-09

Bank Of America Corporation

System and method for digitally marking artificial intelligence (AI) generated content

Families Citing this family (132)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20110010518A1

( en )

2005-12-19

2011-01-13

Srinivas Kavuri

Systems and Methods for Migrating Components in a Hierarchical Storage Network

US10552085B1

( en )

2014-09-09

2020-02-04

Radian Memory Systems, Inc.

Techniques for directed data migration

US10275320B2

( en )

2015-06-26

2019-04-30

Commvault Systems, Inc.

Incrementally accumulating in-process performance data and hierarchical reporting thereof for a data stream in a secondary copy operation

US11238164B2

( en )

*

2017-07-10

2022-02-01

Burstiq, Inc.

Secure adaptive data storage platform

US12184781B2

( en )

2017-07-10

2024-12-31

Burstiq, Inc.

Systems and methods for accessing digital assets in a blockchain using owner consent contracts

US11494692B1

( en )

2018-03-26

2022-11-08

Pure Storage, Inc.

Hyperscale artificial intelligence and machine learning infrastructure

US11861423B1

( en )

2017-10-19

2024-01-02

Pure Storage, Inc.

Accelerating artificial intelligence (‘AI’) workflows

US12067466B2

( en )

2017-10-19

2024-08-20

Pure Storage, Inc.

Artificial intelligence and machine learning hyperscale infrastructure

US10671434B1

( en )

2017-10-19

2020-06-02

Pure Storage, Inc.

Storage based artificial intelligence infrastructure

US10360214B2

( en )

2017-10-19

2019-07-23

Pure Storage, Inc.

Ensuring reproducibility in an artificial intelligence infrastructure

US11455168B1

( en )

2017-10-19

2022-09-27

Pure Storage, Inc.

Batch building for deep learning training workloads

US10951552B2

( en )

*

2017-10-30

2021-03-16

International Business Machines Corporation

Generation of a chatbot interface for an application programming interface

US10831591B2

( en )

2018-01-11

2020-11-10

Commvault Systems, Inc.

Remedial action based on maintaining process awareness in data storage management

EP3528435B1

( en )

*

2018-02-16

2021-03-31

Juniper Networks, Inc.

Automated configuration and data collection during modeling of network devices

US11137926B1

( en )

*

2018-03-30

2021-10-05

Veritas Technologies Llc

Systems and methods for automatic storage tiering

US11301776B2

( en )

*

2018-04-14

2022-04-12

International Business Machines Corporation

Memory-based data selection scheme for machine learning training on limited memory resources

EP3564873B1

( en )

2018-04-30

2022-11-30

Hewlett Packard Enterprise Development LP

System and method of decentralized machine learning using blockchain

EP3564883B1

( en )

2018-04-30

2023-09-06

Hewlett Packard Enterprise Development LP

System and method of decentralized management of device assets outside a computer network

EP3565218B1

( en )

2018-04-30

2023-09-27

Hewlett Packard Enterprise Development LP

System and method of decentralized management of multi-owner nodes using blockchain

US10771240B2

( en )

2018-06-13

2020-09-08

Dynamic Blockchains Inc

Dynamic blockchain system and method for providing efficient and secure distributed data access, data storage and data transport

US20200050443A1

( en )

2018-08-10

2020-02-13

Nvidia Corporation

Optimization and update system for deep learning models

US20200067851A1

( en )

*

2018-08-21

2020-02-27

Argela Yazilim ve Bilisim Teknolojileri San. ve Tic. A.S.

Smart software-defined network (sdn) switch

US10798592B2

( en )

*

2018-08-22

2020-10-06

At&amp;T Intellectual Property I, L.P.

Data parking within offline community system

US11403558B1

( en )

*

2018-09-18

2022-08-02

Iqvia Inc.

GxP artificial intelligence / machine learning (AI/ML) platform

US11062042B1

( en )

2018-09-26

2021-07-13

Splunk Inc.

Authenticating data associated with a data intake and query system using a distributed ledger system

US10983879B1

( en )

*

2018-10-31

2021-04-20

EMC IP Holding Company LLC

System and method for managing recovery of multi-controller NVMe drives

US11320995B2

( en )

2018-10-31

2022-05-03

Western Digital Technologies, Inc.

Transferring computational operations to controllers of data storage devices

CN111222903B

( en )

*

2018-11-27

2023-04-25

北京嘀嘀无限科技发展有限公司

System and method for processing data from an online on-demand service platform

US11379308B2

( en )

*

2018-12-10

2022-07-05

Zoox, Inc.

Data processing pipeline failure recovery

US20200192572A1

( en )

2018-12-14

2020-06-18

Commvault Systems, Inc.

Disk usage growth prediction system

EP4369229A3

( en )

*

2018-12-31

2024-09-25

INTEL Corporation

Securing systems employing artificial intelligence

US11799952B2

( en )

*

2019-01-07

2023-10-24

Intel Corporation

Computing resource discovery and allocation

US11537936B2

( en )

*

2019-01-17

2022-12-27

Servicenow, Inc.

Data set generation for testing of machine learning pipelines

US11966818B2

( en )

*

2019-02-21

2024-04-23

Hewlett Packard Enterprise Development Lp

System and method for self-healing in decentralized model building for machine learning using blockchain

US11436003B2

( en )

*

2019-03-26

2022-09-06

Flowfinity Wireless, Inc.

Non-stop internet-of-things (IoT) controllers

US11928559B2

( en )

*

2019-04-08

2024-03-12

Google Llc

Transformation for machine learning pre-processing

CN118113663A

( en )

2019-04-25

2024-05-31

伊姆西Ip控股有限责任公司

Method, apparatus and computer program product for managing a storage system

US11150978B2

( en )

2019-04-26

2021-10-19

Bank Of America Corporation

Automated system for intelligent error correction within an electronic blockchain ledger

US10817475B1

( en )

2019-05-03

2020-10-27

EMC IP Holding Company, LLC

System and method for encoding-based deduplication

US11138154B2

( en )

2019-05-03

2021-10-05

EMC IP Holding Company, LLC

System and method for offset-based deduplication

US10990565B2

( en )

2019-05-03

2021-04-27

EMC IP Holding Company, LLC

System and method for average entropy calculation

US10963437B2

( en )

2019-05-03

2021-03-30

EMC IP Holding Company, LLC

System and method for data deduplication

US10733158B1

( en )

*

2019-05-03

2020-08-04

EMC IP Holding Company LLC

System and method for hash-based entropy calculation

US12204781B2

( en )

*

2019-05-10

2025-01-21

Dell Products L.P.

System and method for performance based dynamic optimal block size data deduplication

CN111984364B

( en )

*

2019-05-21

2023-05-26

江苏艾蒂娜互联网科技有限公司

Artificial intelligence cloud platform towards 5G age

US11269859B1

( en )

*

2019-05-22

2022-03-08

Splunk Inc.

Correlating different types of data of a distributed ledger system

US11507562B1

( en )

2019-05-22

2022-11-22

Splunk Inc.

Associating data from different nodes of a distributed ledger system

CN110120957B

( en )

*

2019-06-03

2019-12-06

浙江鹏信信息科技股份有限公司

Safe disposal digital twin method and system based on intelligent scoring mechanism

JP7326903B2

( en )

*

2019-06-14

2023-08-16

富士フイルムビジネスイノベーション株式会社

Information processing device and program

CN112307113A

( en )

*

2019-07-29

2021-02-02

中兴通讯股份有限公司

Service request message sending method and distributed database architecture

CN112306925B

( en )

2019-08-02

2023-02-10

华为技术有限公司

Access request processing method, device, equipment and storage medium

CN110471766B

( en )

*

2019-08-06

2022-12-30

北京华恒盛世科技有限公司

GPU resource scheduling system and method based on CUDA

US11755884B2

( en )

2019-08-20

2023-09-12

Micron Technology, Inc.

Distributed machine learning with privacy protection

US11636334B2

( en )

*

2019-08-20

2023-04-25

Micron Technology, Inc.

Machine learning with feature obfuscation

US11392796B2

( en )

2019-08-20

2022-07-19

Micron Technology, Inc.

Feature dictionary for bandwidth enhancement

US11514372B2

( en )

*

2019-08-30

2022-11-29

Microsoft Technology Licensing, Llc

Automatically tuning parameters in a layered model framework

US11769070B2

( en )

2019-10-09

2023-09-26

Cornell University

Quantum computing based hybrid solution strategies for large-scale discrete-continuous optimization problems

US11455574B2

( en )

*

2019-11-21

2022-09-27

International Business Machines Corporation

Dynamically predict optimal parallel apply algorithms

US11294759B2

( en )

*

2019-12-05

2022-04-05

International Business Machines Corporation

Detection of failure conditions and restoration of deployed models in a computing environment

US11943293B1

( en )

*

2019-12-06

2024-03-26

Pure Storage, Inc.

Restoring a storage system from a replication target

US11237941B2

( en )

*

2019-12-12

2022-02-01

Cognizant Technology Solutions India Pvt. Ltd.

System and method for application transformation to cloud based on semi-automated workflow

US12014195B2

( en )

2019-12-12

2024-06-18

Cognizant Technology Solutions India Pvt. Ltd.

System for providing an adaptable plugin framework for application transformation to cloud

US11899570B2

( en )

2019-12-12

2024-02-13

Cognizant Technology Solutions India Pvt. Ltd.

System and method for optimizing assessment and implementation of microservices code for cloud platforms

US11775867B1

( en )

*

2019-12-18

2023-10-03

System Inc.

System and methods for evaluating machine learning models

US11218293B2

( en )

2020-01-27

2022-01-04

Hewlett Packard Enterprise Development Lp

Secure parameter merging using homomorphic encryption for swarm learning

US11748835B2

( en )

2020-01-27

2023-09-05

Hewlett Packard Enterprise Development Lp

Systems and methods for monetizing data in decentralized model building for machine learning using a blockchain

US12393883B2

( en )

2020-01-31

2025-08-19

Hewlett Packard Enterprise Development Lp

Adaptively synchronizing learning of multiple learning models

WO2021257128A2

( en )

*

2020-02-14

2021-12-23

Cornell University

Quantum computing based deep learning for detection, diagnosis and other applications

CN111461958A

( en )

*

2020-05-18

2020-07-28

江苏电力信息技术有限公司

System and method for controlling real-time detection and optimization processing of rapid multi-path data streams

US11693867B2

( en )

*

2020-05-18

2023-07-04

Google Llc

Time series forecasting

US11593393B1

( en )

2020-05-22

2023-02-28

Cigna Intellectual Property, Inc.

Systems and methods for providing automated integration and error resolution of records in complex data systems

EP3940607A1

( en )

*

2020-07-14

2022-01-19

Siemens Aktiengesellschaft

A neural network model, a method and modelling environment for configuring neural networks

US11816069B2

( en )

2020-07-27

2023-11-14

International Business Machines Corporation

Data deduplication in blockchain platforms

US11496521B2

( en )

2020-08-12

2022-11-08

International Business Machines Corporation

Feedback loop for security audit logs

US11651096B2

( en )

2020-08-24

2023-05-16

Burstiq, Inc.

Systems and methods for accessing digital assets in a blockchain using global consent contracts

JP2022057864A

( en )

*

2020-09-30

2022-04-11

富士フイルムビジネスイノベーション株式会社

Information processing equipment and information processing programs

US11829799B2

( en )

*

2020-10-13

2023-11-28

International Business Machines Corporation

Distributed resource-aware training of machine learning pipelines

KR102900828B1

( en )

*

2020-11-20

2025-12-16

삼성전자주식회사

Electronic apparatus and controlling method thereof

US12412123B2

( en )

*

2020-11-20

2025-09-09

Samsung Electronics Co., Ltd.

Electronic apparatus and controlling method thereof

KR20220087297A

( en )

2020-12-17

2022-06-24

삼성전자주식회사

Storage device executing processing code, and operating method thereof

US11645104B2

( en )

*

2020-12-22

2023-05-09

Reliance Jio Infocomm Usa, Inc.

Intelligent data plane acceleration by offloading to distributed smart network interfaces

US12602263B2

( en )

*

2020-12-30

2026-04-14

International Business Machines Corporation

Transparent data transformation and access for workloads in cloud environments

US20220245485A1

( en )

*

2021-02-04

2022-08-04

Netapp, Inc.

Multi-model block capacity forecasting for a distributed storage system

US11966340B2

( en )

2021-02-18

2024-04-23

International Business Machines Corporation

Automated time series forecasting pipeline generation

US11789779B2

( en )

*

2021-03-01

2023-10-17

Bank Of America Corporation

Electronic system for monitoring and automatically controlling batch processing

CN112906907B

( en )

*

2021-03-24

2024-02-23

成都工业学院

Method and system for layering management and distribution of machine learning pipeline model

Related documents

Record · ID 607043
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.