Predicting protein structures over multiple iterations using recycling

US2023402133A1 · US · A1

Patent metadata
FieldValue
Publication numberUS-2023402133-A1
Application numberUS-202118034280-A
CountryUS
Kind codeA1
Filing dateNov 23, 2021
Priority dateNov 28, 2020
Publication dateDec 14, 2023
Grant date

How to read this patent

A practical reading order for non-experts. Skip the full description unless you need deep technical detail.

  1. Title

    What the patent document calls the invention.

  2. Abstract

    A short plain-language summary of the technical disclosure.

  3. Assignees and inventors

    Who owns or filed the patent and who is credited as inventor.

  4. Key dates

    Filing, priority, publication, and grant dates set the timeline.

  5. First independent claim

    The legal scope of protection — read this for what is actually claimed.

  6. CPC / IPC classifications

    Technology tags used to group this patent with similar filings.

  7. Citations and related patents

    Prior art links and similar publications in this corpus.

Abstract

Official abstract text for this publication.

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for predicting a structure of a protein comprising one or more chains. In one aspect, a method comprises, at each subsequent iteration after a first iteration in a sequence of iterations: obtaining a network input for the subsequent iteration that characterizes the protein; generating, from (i) structure parameters generated at a preceding iteration that precedes the subsequent iteration in the sequence, (ii) one or intermediate outputs generated by the protein structure prediction neural network while generating the structure parameters at the last iteration, or (iii) both, features for the subsequent iteration; and processing the features and the network input for the subsequent iteration using the protein structure prediction neural network to generate structure parameters for the subsequent iteration that define another predicted structure for the protein.

First claim

Opening claim text (preview).

What is claimed is: 1 . A method performed by one or more computers for predicting a structure of a protein comprising one or more chains, wherein each chain comprises a sequence of amino acids, the method comprising: at a first iteration of a sequence of iterations that comprises the first iteration followed by one or more subsequent iterations: obtaining a network input for the first iteration that characterizes the protein; processing the network input for the first iteration using a protein structure prediction neural network to generate structure parameters for the first iteration that define an initial predicted structure for the protein; at each subsequent iteration in the sequence of iterations: obtaining a network input for the subsequent iteration that characterizes the protein; generating, from (i) the structure parameters generated at a preceding iteration that precedes the subsequent iteration in the sequence, (ii) one or intermediate outputs generated by the protein structure prediction neural network while generating the structure parameters at the last iteration, or (iii) both, features for the subsequent iteration; and processing the features and the network input for the subsequent iteration using the protein structure prediction neural network to generate structure parameters for the subsequent iteration that define another predicted structure for the protein; and determining a final predicted structure for the protein from the structure parameters for the last iteration in the sequence. 2 . The method of claim 1 , wherein the network input for each iteration in the sequence of iterations is the same network input. 3 . The method of claim 1 , wherein processing the features and the network input for the subsequent iteration using the protein structure prediction neural network to generate structure parameters for the subsequent iteration that define a predicted structure for the protein comprises: generating a combined input from the features and the network input for the subsequent iteration; and processing the combined input using the protein structure prediction neural network to generate the structure parameters for the subsequent iteration. 4 . The method of claim 1 , wherein: the respective network input for each iteration comprises a respective initial pair embedding for each pair of amino acids in the protein, at each iteration, the structure prediction neural network is configured to repeatedly update the initial pair embeddings while generating the structure parameters for the iteration, generating the features comprises generating, from updated pair embeddings generated while generating the structure parameters at the preceding iteration, a transformed set of pair embeddings; and generating the combined input comprises combining the transformed set of pair embeddings and the initial pair embeddings for the subsequent iteration. 5 . The method of claim 1 , wherein: the respective network input for each iteration comprises an initial multiple sequence alignment (MSA) representation that represents a respective MSA corresponding to each chain in the protein, at each iteration, the structure prediction neural network is configured to generate one or more sets of single embeddings that each include a respective single embedding for each amino acid the protein while generating the structure parameters for the iteration, generating the features comprises generating, from one of the sets of single embeddings generated while generating the structure parameters at the preceding iteration, a transformed set of single embeddings; and generating the combined input comprises combining the transformed set of single embeddings and the initial MSA representation for the subsequent iteration. 6 . The method of claim 5 , wherein combining the transformed set of single embeddings and the initial MSA representation for the subsequent iteration comprises adding the transformed set of single embeddings to a first row of the initial MSA representation. 7 . The method of claim 1 , wherein: the respective network input for each iteration comprises a respective initial pair embedding for each pair of amino acids in the protein, at each iteration, the structure parameters specify, for each amino acid, a predicted 3-D spatial location of a specified atom in the amino acid in the structure of the protein; generating the features comprises: generating, from the predicted 3-D spatial locations for the amino acids specified by the structure parameters at the preceding iteration, a distance map that characterizes, for each pair of amino acids in the protein, a respective estimated distance between the pair of amino acids in the structure of the protein; and generating, from the distance map, a transformed distance map that has a same dimensionality as the initial pair embeddings; and generating the combined input comprises combining the transformed distance map and the initial pair embeddings for the subsequent iteration. 8 . The method of claim 1 , wherein: the respective network input for each iteration comprises a respective initial pair embedding for each pair of amino acids in the protein, at each iteration, the structure parameters specify a distance map that characterizes, for each pair of amino acids in the protein, a respective estimated distance between the pair of amino acids in the structure of the protein; generating the features comprises: generating, from the distance map specified by the structure parameters at the preceding iteration, a transformed distance map that has a same dimensionality as the initial pair embeddings; and generating the combined input comprises combining the transformed distance map and the initial pair embeddings for the subsequent iteration. 9 . The method of claim 1 , wherein: the respective network input for each iteration comprises a respective initial pair embedding for each pair of amino acids in the protein that is generated using a set of one or more template sequences and corresponding known structures for each of the template sequences; and generating the combined input comprises modifying the initial pair embeddings for the subsequent iteration by adding the protein and the structure prediction defined by the embeddings at the preceding iteration to the set of one or more template sequences and the corresponding known structures. 10 .- 25 . (canceled) 26 . A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for predicting a structure of a protein comprising one or more chains, wherein each chain comprises a sequence of amino acids, the operations comprising: at a first iteration of a sequence of iterations that comprises the first iteration followed by one or more subsequent iterations: obtaining a network input for the first iteration that characterizes the protein; processing the network input for the first iteration using a protein structure prediction neural network to generate structure parameters for the first iteration that define an initial predicted structure for the protein; at each subsequent iteration in the sequence of iterations: obtaining a network input for the subsequent iteration that characterizes the protein; generating, from (i) the structure parameters generated at a preceding iteration that precedes the subsequent iteration in the sequence, (ii) one or intermediate outputs generated by the protein structure prediction

Assignees

Inventors

Classifications

  • Feedforward networks · CPC title

  • Supervised learning · CPC title

  • G16B40/20Primary

    Supervised data analysis · CPC title

  • Protein or domain folding · CPC title

  • Drug targeting using structural data; Docking or binding prediction · CPC title

Patent family

Related publications grouped by family.

External sources

Frequently asked questions

Answers are generated from the same data shown on this page.

What does patent US2023402133A1 cover?
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for predicting a structure of a protein comprising one or more chains. In one aspect, a method comprises, at each subsequent iteration after a first iteration in a sequence of iterations: obtaining a network input for the subsequent iteration that characterizes the protein; generating, from (i) st…
Who is the assignee on this patent?
Deepmind Tech Ltd
What technology area does this patent fall under?
Primary CPC classification G16B40/20. Mapped technology areas include Physics.
When was this patent published?
Publication date Thu Dec 14 2023 00:00:00 GMT+0000 (Coordinated Universal Time) (A1). Legal status and post-grant events are not shown on this page.
What related patents are in patentsdb?
We list 8 related publications on this page (citations in our corpus or others sharing the same primary CPC).