What technology area does this patent fall under?

Primary CPC classification H04L67/1097. Mapped technology areas include Electricity.

When was this patent published?

Publication date Tue Sep 12 2017 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.

What related patents are in patentsdb?

We list 3 related publications on this page (citations in our corpus or others sharing the same primary CPC).

Dynamic node group allocation

US9762672B2 · US · B2

Patent metadata
Field	Value
Publication number	US-9762672-B2
Application number	US-201514740050-A
Country	US
Kind code	B2
Filing date	Jun 15, 2015
Priority date	Jun 15, 2015
Publication date	Sep 12, 2017
Grant date	Sep 12, 2017

How to read this patent

A practical reading order for non-experts. Skip the full description unless you need deep technical detail.

Title
What the patent document calls the invention.
Abstract
A short plain-language summary of the technical disclosure.
Assignees and inventors
Who owns or filed the patent and who is credited as inventor.
Key dates
Filing, priority, publication, and grant dates set the timeline.
First independent claim
The legal scope of protection — read this for what is actually claimed.
CPC / IPC classifications
Technology tags used to group this patent with similar filings.
Citations and related patents
Prior art links and similar publications in this corpus.

Abstract

Official abstract text for this publication.

Provided are techniques for improving data locality for parallel applications running in a big data distributed file system with a dynamic node group. In response to a consumer job starting to read one or more files in a big data distributed file system having multiple nodes, node group information for the one or more files to be read is retrieved, wherein the node group information identifies nodes from the multiple nodes on which a producer job wrote the one or more files, and the consumer job is assigned to the nodes identified by the node group information to allow for local reading of the one or more files by the consumer job.

First claim

Opening claim text (preview).

What is claimed is: 1. A method, comprising: connecting a parallel application server to a data source structure, wherein the data source structure contains a big data distributed file system, wherein the big data distributed file system contains multiple nodes and data blocks; in response to the parallel application operating the data source structure within the multiple nodes of the big data distributed file system, the parallel application server and the data source structure performing read and write operations on the data blocks in a local mode setting; in response to the parallel application operating the data source structure outside of the multiple nodes of the big data distributed file system, the parallel application server and the data source structure performing read and write operations on the data blocks in a remote mode setting; in response to a consumer job starting to read one or more files in the big data distributed file system, retrieving node group information for the one or more files to be read, wherein the node group information identifies nodes from the multiple nodes on which a producer job wrote the one or more files; implementing a node grouping mechanism to read and write the data blocks within the local mode setting over the remote mode setting; assigning the consumer job to the nodes identified by the node group information to allow for reading of the one or more files by the consumer job within the local mode setting, wherein the local mode setting reads and writes the data blocks; in response to assigning the consumer job to the nodes identified by the node group information, generating a configuration file, wherein the configuration file comprises a dynamically generated configuration file and a non-dynamically generated configuration file; wherein the dynamically generated configuration file corresponds to the consumer job and the dynamically generated configuration file is dynamically assigned to the node group for the consumer job; in response to retrieving the node group information, requesting logical resources; executing the consumer job with the configuration file identifying the nodes on which the consumer job is to run; and in response to determining that logical resources cannot be allocated in the nodes identified by the node group information, attempting to allocate logical resources in nodes close to the nodes identified by the node group information. 2. The method of claim 1 , further comprising: storing a full path file name along with the node group information in a table. 3. The method of claim 1 , wherein software is provided as a service in a cloud environment. 4. A computer system, comprising: one or more processors, one or more computer-readable memories and one or more computer-readable, tangible storage devices; a parallel application server connected to a data source structure, wherein the data source structure contains a big data distributed file system, wherein the big data distributed file system contains multiple nodes and data blocks; and program instructions, stored on at least one of the one or more computer-readable, tangible storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, to perform: in response to the parallel application operating the data source structure within the multiple nodes of the big data distributed file system, the parallel application server and the data source structure performing read and write operations on the data blocks within a local mode setting; in response to the parallel application operating the data source structure outside the multiple nodes of the big data distributed file system, the parallel application server and the data source structure performing read and write operations on the data blocks within a remote mode setting; in response to a consumer job starting to read one or more files in the big data distributed file system, retrieving node group information for the one or more files to be read, wherein the node group information identifies nodes from the multiple nodes on which a producer job wrote the one or more files; implementing a node grouping mechanism to read and write the data blocks within the local mode setting over the remote mode setting; assigning the consumer job to the nodes identified by the node group information to allow for reading of the one or more files by the consumer job within the local mode setting, wherein the local mode setting reads and writes the data blocks; in response to assigning the consumer job to the nodes identified by the node group information, generating a configuration file, wherein the configuration file comprises a dynamically generated configuration file and a non-dynamically generated configuration file; wherein the dynamically generated configuration file corresponds to the consumer job and the dynamically generated configuration file is dynamically assigned to the node group for the consumer job; in response to retrieving the node group information, requesting logical resources; executing the consumer job with the configuration file identifying the nodes on which the consumer job is to run; and in response to determining that logical resources cannot be allocated in the nodes identified by the node group information, attempting to allocate logical resources in nodes close to the nodes identified by the node group information. 5. The computer system of claim 4 , wherein the operations further comprise: storing a full path file name along with the node group information in a table. 6. The computer system of claim 4 , wherein a Software as a Service (SaaS) is configured to perform the system operations. 7. A computer program product, the computer program product comprising a computer readable storage medium having program code embodied therewith, the program code executable by at least one processor to perform: connecting a parallel application server to a data source structure, wherein the data source structure contains a big data distributed file system, wherein the big data distributed file system contains multiple nodes and data blocks; in response to the parallel application operating the data source structure within the multiple nodes of the big data distributed file system, the parallel application server and the data source structure perform read and write operations on the data blocks within a local mode setting; in response to the parallel application operating the data source structure outside the multiple nodes of the big data distributed file system, the parallel application server and the data source structure perform read and write operations on the data blocks within a remote mode setting; in response to a consumer job starting to read one or more files in the big data distributed file system, retrieving node group information for the one or more files to be read, wherein the node group information identifies nodes from the multiple nodes on which a producer job wrote the one or more files; implementing a node grouping mechanism to read and write the data blocks within the local mode setting over the remote mode setting; assigning the consumer job to the nodes identified by the node group information to allow for reading of the one or more files by the consumer job within the local mode setting, wherein the local mode setting reads and writes the data blocks; in response to assigning the consumer job to the nodes identified by the node group information, generating a configuration file, wherein the configuration file comprises a dynamically generated configuration file and a non-dynamically generated configuration file; wherein the dynamically generated configuration file corresponds to the consumer job and th

Assignees

Inventors

Classifications

H04L67/1097Primary
for distributed storage of data in networks, e.g. transport arrangements for network file system [NFS], storage area networks [SAN] or network attached storage [NAS] · CPC title
H04L47/762
triggered by the network · CPC title
G06F17/30194
Physics · mapped topic
G06F16/182
Distributed file systems · CPC title

Patent family

Related publications grouped by family.

View patent family 57517416

External sources

Frequently asked questions

Answers are generated from the same data shown on this page.

What does patent US9762672B2 cover?: Provided are techniques for improving data locality for parallel applications running in a big data distributed file system with a dynamic node group. In response to a consumer job starting to read one or more files in a big data distributed file system having multiple nodes, node group information for the one or more files to be read is retrieved, wherein the node group information identifies …
Who is the assignee on this patent?: IBM
What technology area does this patent fall under?: Primary CPC classification H04L67/1097. Mapped technology areas include Electricity.
When was this patent published?: Publication date Tue Sep 12 2017 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.
What related patents are in patentsdb?: We list 3 related publications on this page (citations in our corpus or others sharing the same primary CPC).

How to read this patent

Abstract

First claim

Assignees

Inventors

Classifications

Patent family

External sources

Related patents

Knowledge-intensive data processing system

Gateway device, file server system, and file distribution method

Column based data transfer in extract, transform and load (ETL) systems

Frequently asked questions