“Watermarking Of Genomic Sequencing Data” in Patent Application Approval Process (USPTO 20230048167): Children’s Hospital Los Angeles
2023 MAR 08 (NewsRx) -- By a
This patent application is assigned to Children’s Hospital Los Angeles (
The following quote was obtained by the news editors from the background information supplied by the inventors: “The advent of next-generation sequencing (NGS) technologies led to the emergence of genomic medicine, which uses the genomic information to understand disease mechanisms and to guide patient care, such as for diagnostic, prognostic and therapeutic decision-making. As part of it, huge amount of genomic sequencing data have been generated for both research and clinical purposes with drastic more such data anticipated in the future. Genomics has been compared with other major sources of Big Data including astronomy, and may be considered the most demanding in terms of all four major aspects of Big Data, namely, data acquisition, storage, distribution, and analysis with the astronomical, or rather genomical, growth of DNA sequencing in terms of the overall sequencing capacity but also the number of human genomes sequenced each year and cumulatively.
“Biomedical research has benefited tremendously from the genomical growth of sequencing capacity. For example, cancer is considered a genetic disease. Using pediatric cancer as an example, pan-cancer analyses of pediatric tumors reveal a spectrum of nuclear somatic DNA alterations that vary by tumor type, and at least 8.5% of pediatric cancer patients have germline mutations in cancer predisposition genes. The patterns of these genomic alterations are distinctly different from one tumor type to another and one patient from another, which have been shown to be of diagnostic, prognostic and therapeutic importance and implications. For example, a comprehensive next-generation sequencing panel, OncoKids, was developed for pediatric cancers, which has demonstrated significant clinical utility in two years since its launch, with clinically significantly findings found in two thirds of 700 patients tested. Clinical exome sequencing tests, similarly, allowed for identification of pathogenic cancer predisposition variants in 8/106 (7.5%) patients tested. Such findings have all been enabled and empowered by the advent of massively parallel sequencing technologies, which led to 1 million fold decrease of the cost of sequencing a human genome since 2003, when the human genome project was completed. These genomic technologies have led to tremendously improved understanding of cancer etiology which, however, is only possible when the researchers and the patients are willing to share the genomic data. Again using research experience as an example the landscape of germline and somatic mitochondrial DNA mutations in pediatric cancers was able to be established from mining the matched tumor-normal whole genome sequencing data of 621 pediatric cancer patients, collected and shared by the St. Jude’s Children’s Hospital instead based on these patients informed consent.
“With the success of the 1000
“One such challenge relates to privacy concerns regarding access to and usage of the genomic data. The genomic sequencing data is deemed Personal Health Information (PHI) according to the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule, and also the General Data Protection Regulation recently established by the
“On the other hand, informed consent is now the essential component of any modern biomedical research involving human subjects. The notion of informed consent emerged after decades of atrocities, followed by tremendous efforts to address the problem that resulted in The Nuremberg Code, The Declaration of
“Alternative to any consent model is the ownership-based governance. The patients or participants of any genomic study ultimately own their data, and should have the governance of the data, which includes the right to control the data and also the right to assess the value of the data, with value being economic or intellectual. This provides the ultimate and most granular control of the participant’s data but requires a distributed model that 1) is participant-centric, 2) does not require any centralized management, and 3) provides the fine-grained control of the participant’s data. This model comes with significant technical challenges for the participants: a) to control what (portion of) data to share, with whom and for what duration, b) to track or trace data access, c) to prevent unauthorized access, d) to prevent or deter illicit duplication and usage of the data, and e) to potentially benefit financially from sharing the data.
“Either dynamic consent or ownership-based governance of accessing or sharing the genomic sequencing data, however, requires robust informatics tools to enable and to facilitate, in order to deal with the associated complexity while ensuring privacy preservation. Such algorithms or tools, however, are severely lacking. For example, there is currently a significant lack of computational or informatics tools to enable the implementation of any real ownership-based governance. Currently, the genomic sequencing data of an individual is shared or not shared, decrypted or encrypted, all as a whole. This makes executing dynamic consent or honoring data owner’s privacy concerns extremely challenging. As an example, the General Data Protection Regulation recently established by the
In addition to the background information obtained for this patent application, NewsRx journalists also obtained the inventors’ summary information for this patent application: “The disclosure addresses the above-described challenges by providing two sets of algorithms and methods that are usable (individually or in cooperation) to protect and secure patient data. In particular, the disclosed algorithms and methods may be used to provide a) a dynamic privacy preserving encryption scheme for data such as genomic sequencing data (e.g., to dynamically encrypt and decrypt genomic sequencing data of user-specified genomic regions, such as specific genes (e.g. APOE and BRCA1/2)) and b) a dynamic, robust, and data utility preserving algorithm for watermarking genomic sequencing data. Fine-grained control policies, such as the time period when the data could be decrypted and accessed and/or the entity to which the data is allowed to be distributed, are enabled using attribute-based encryption and/or watermarking. With such features, the methods and systems of the disclosure enhance the privacy of the genomic sequencing data by: (1) giving full control over genomic data to the data owner; (2) enabling flexible, efficient, and precise partial encryption and decryption of genomic data. Furthermore, these algorithms, as detailed below, empower individual data owners and let them control when and for what purpose to share a specific portion of their genomic sequencing data, all in a reliable and auditable manner. These algorithms may provide for i) a reduction in the cost of implementing and maintaining a dynamic consent platform because of the distributed nature of ownership-based governance, ii) a promotion and facilitation of genomic data sharing, iii) a support of “consent revocation”, and iv) a minimization of the “data holders” liability from improper handling of the participants’ data and the inability to honor the decisions of the participants thoroughly and in real-time. In this way, the disclosed features provide technical solutions to achieve principles of the above-described dynamic consent and ownership-based governance models, as well as other enhanced user controls regarding access to data, usage of data, and tracking/auditing of data.
“Example innovations described in the disclosure are the novel use of digital watermarking to enable the tracking and auditing of distributed data. The data is watermarked with selected watermarking elements (e.g., values of data, such as a selected alternate genomic base replacing a reference base determined by a sequence read) at selected locations in a file that are determined using a random seed that is based on a secret key.
“Further example innovations described in the disclosure are the novel use of a block cipher mode of operation, and the way the encryption is applied to genomic data. The data is encrypted once using the master encryption key; the data can be discarded after that. A novel index structure is designed to facilitate fine-grained data region control, to make sure that no additional data is exposed. The index allows the data owner to build keystreams (derived from the master encryption key) that decrypt specific genomic regions without sharing the master encryption key with other entities, and without the need to store the actual data. In some examples, the watermarking innovations may be combined, in full or in part, with the encryption/decryption innovations to provide further control over the genomic data.
“Achieving a trustworthy genomic data sharing is imperative if the benefits anticipated from large-scale data sharing are to be realized. The algorithms, methods, and systems described herein enable true ownership-based governance of genomic sequencing data and greatly simplify the attempts to implement dynamic patient consent for biomedical studies. Using the described mechanisms, the data owner will be able to specify and revoke authorizations for data access and use. Such owner-centered data management will improve the trust relationship between the data owner and the data users, removing the barriers for genomic data sharing. Furthermore, with greatly simplified ownership-based governance of the genomic sequencing data, the owner, instead of large diagnostic or healthcare companies that generate and hence control the genomic sequencing data, could potentially benefit financially from sharing his or her own data, as it should be. Furthermore, it is to be understood that genomic data is provided herein as an example, and the disclosed systems and methods may be applied to dynamically encrypt and/or decrypt any suitable data or file type.
“The foregoing and other objects, features, and advantages of the invention will become more apparent from the following detailed description, which proceeds with reference to the accompanying figures.”
The claims supplied by the inventors are:
“1. A method of dynamically applying a watermark to at least a portion of a file, the method comprising: generating, using information derived from a secret key, a first random seed; generating, using the first random seed, an ordered pseudorandom set of integers; generating, using dynamic attribute information, a second random seed; selecting, using the second random seed, a subset of the ordered pseudorandom set of integers, the subset corresponding to identifiers of data locations in the file; and modifying data at data locations in the file corresponding to at least a portion of the identifiers included in the subset to generate a watermarked file.
“2. The method of claim 1, wherein the dynamic attribute information includes entity information for an entity to which the file is being distributed, timing information corresponding to a validity time period for the file, a data usage policy for the file, and/or one or more other attributes of a policy for the data.
“3. The method of claim 1, wherein modifying the data comprises generating, using the first random seed, a pseudorandom integer and changing the data to a value that is based on the pseudorandom integer.
“4. The method of claim 1, further comprising determining which of the data locations corresponding to the identifiers of the subset meet selected criteria, and wherein the portion of the identifiers correspond to the identifiers of the subset that meet the selected criteria.
“5. The method of claim 1, further comprising assigning a selected quality score to modified data, the selected quality score being selected based on quality scores of data at each other data location in the file, wherein the selected quality score corresponds to a quality score below a threshold that is most frequently assigned to the data at each other data location in the file relative to other quality scores below the threshold.
“6. The method of claim 1, wherein an entity to which the file is being distributed is a first entity, and wherein the subset of the ordered pseudorandom set of integers is selected to only partially overlap with another subset or subsets of ordered pseudorandom sets of integers that is generated for watermarking the file for distribution to another, different entity or entities.
“7. The method of claim 1, wherein the file comprises a genomic data file that includes a sequencing data set, wherein the data locations in the file comprise reference bases in the sequencing data set, and wherein modifying the data comprises switching the reference bases in the data locations in the file corresponding to at least the portion of the identifiers included in the subset from the respective reference base to a selected alternative base.
“8. The method of claim 7, wherein the selected alternative base is selected based on a randomly generated number that is generated using the first random seed.
“9. The method of claim 7, further comprising determining which of the data locations corresponding to the identifiers of the subset meet selected criteria, wherein the portion of the identifiers correspond to the identifiers of the subset that meet the selected criteria, and wherein the selected criteria includes data locations that have a number of sequencing reads with the selected alternative base that is less than a threshold.
“10. The method of claim 1, wherein the watermarked file is a reference watermarked file, the method further comprising validating a targeted file by determining whether the watermark is present in the targeted file by generating a sequence of watermark elements based on information derived from the secret key and comparing the percentage of watermark elements discovered in the targeted file to the expected percentage of watermark elements that can be discovered by chance, estimated by a Monte Carlo simulation with random seeds.
“11. The method of claim 10, further comprising detecting collusion between two or more entities to attempt to modify or remove a watermark from the file by determining which watermark elements in the sequence of watermark elements generated during generation of the reference watermarked file are missing in the targeted file and which watermark elements in the sequence of watermark elements generated during generation of the reference watermarked file are present in the targeted file.
“12. The method of claim 1, further comprising dynamically encrypting the watermarked file, wherein the secret key is a watermarking secret key and the watermarked filed is formed of multiple blocks of ordered data to enable partial decryption of the watermarked file, and wherein dynamically encrypting the watermarked file comprises: generating, using an encryption secret key and one or more initialization vectors associated with the watermarked file, a keystream for the multiple blocks of ordered data of the watermarked file; encrypting the multiple blocks of ordered data of the watermarked file by performing a logical operation of the keystream with the multiple blocks of ordered data in a one-to-one correspondence; and building a file index of the watermarked file to identify location information of the multiple blocks of ordered data.
“13. A system for detecting and/or verifying a watermark in a file, the system comprising: a processor; and memory storing instructions executable by the processor to: generate, using information derived from a secret key associated with the watermark, a first random seed; generate, using the first random seed, an ordered pseudorandom set of integers; generate, using entity information for at least one entity to which the file was distributed and timing information corresponding to a validity time period for the file, a second random seed; select, using the second random seed, a subset of the ordered pseudorandom set of integers, the subset corresponding to identifiers of data locations in the file; generate a sequence of watermark elements, the watermark elements comprising expected values for associated locations in the file, the associated locations being selected based on the first random seed and the expected values being selected based on the second random seed; and compare the sequence of watermark elements to the file to determine whether the associated locations in the file are populated with the respective associated expected values.
“14. The system of claim 13, wherein the file is an encrypted file formed of multiple blocks of encrypted data, the method further comprising dynamically decrypting at least a portion of the file to generate a decrypted file, and wherein comparing the sequence of watermark elements to the file comprises comparing the sequence of watermark elements to the decrypted file.
“15. The system of claim 14, wherein dynamically decrypting at least a portion of the file comprises: receiving a request to decrypt at least one selected block of encrypted data of the file, responsive to validating the request, retrieving a portion of a keystream for the file, the portion of the keystream corresponding to the at least one selected block, and decrypting the at least one selected block by performing a logical operation of the portion of the keystream with the encrypted data of the at least one selected block to generate plaintext data corresponding only to the at least one selected block.
“16. The system of claim 15, wherein validating the request comprises comparing attributes of the request and a user making the request with one or more attributes associated with the user and/or policies bound with the encrypted data to determine if the user and the request are in compliance with the attributes and policies, respectively.
“17. The system of claim 15, wherein dynamically decrypting the file comprises decrypting selected portions of the file using the keystream while remaining portions of the file are not decryptable.
“18. The system of claim 15, wherein selected portions of the file are decryptable using the portion of the keystream while remaining portions of the file are not decryptable.
“19. The system of claim 15, wherein the encrypted data of the file is generated using an encryption secret key, the encryption secret key being used to generate the keystream, different portions of which are subsequently used for decrypting only respective portions of the file in respective decryption iterations without sharing the encryption secret key.
“20. The system of claim 15, wherein the at least one selected block of encrypted data comprises only a subset of the multiple blocks of encrypted data of the watermarked file.”
URL and more information on this patent application, see: Gai, Xiaowu; Ryutov, Alex; Ryutov, Tatyana. Watermarking Of Genomic Sequencing Data.
(Our reports deliver fact-based news of research and discoveries from around the world.)


Researchers Submit Patent Application, “System For Tracking Patient Referrals”, for Approval (USPTO 20230050346): Patent Application
ProAssurance Declares Quarterly Dividend
Advisor News
- Succession planning: Building the future of your practice
- From loss to security: Supporting widowed clients with life insurance
- Plan now for lower Social Security benefits later
- The conversation almost no advisor is having yet
- Why advisors should offer retirement-longevity planning
More Advisor NewsAnnuity News
- Empower Annuity Insurance Company of America Trademark Application for “EMPOWER WHAT’S NEXT” Filed: Empower Annuity Insurance Company of America
- Industry pushes back on linking ‘financial strength’ to annuity illustrations
- Sammons Enterprises & Sammons Financial Group Respond to Reports
- The Manhattan Life Insurance Company Acquires Union Security Life Insurance Company of New York
- Cayman Islands premier to meet with U.S. reinsurance regulators
More Annuity NewsHealth/Employee Benefits News
Life Insurance News
- Record IUL sales don’t diminish the need for continued customer engagement
- Benchmark International Successfully Facilitated the Transaction Between National Group Marketing Trust and New Era Life Insurance Companies
- Why the bond market is flexing its muscles, and why everyone needs to care
- An Application for the Trademark “LIVE TODAY, SECURE TOMORROW.” Has Been Filed by Security Mutual Life Insurance Company of New York: Security Mutual Life Insurance Company of New York
- Modern Woodmen board selects Shea Doyle as next president and CEO
More Life Insurance News