Whole Genome Sequencing

My understanding is that CRAM can be lossy by configuration but is structurally not so. Loss, if any, depends on what you tell the encoder to keep; i.e., which axes of “loss” matter and how much. The compression strategy is reference-based: instead of storing each base, CRAM stores the differences from a reference genome, plus enough metadata to reconstruct the read. To be lossless, you must have the reference file at decode time. (I have heard that you can embed the reference data in the file but this can present size problems.

If you choose loss, it can happen in three places:

Quality score compression. Q-scores can be kept verbatim or reduced by several levels. Since Q-scores are the single largest contributor to BAM/CRAM size, they are typically shrunk.

Read names. Can be preserved, replaced with auto-generated tokens (preserving pairing but losing the original instrument-level name), or dropped entirely. I’m not quite sure how this operates but it is claimed not to matter much for downstream analysis but might for forensics.

Auxiliary tags. A smaller contributor, you can select which BAM tags to retain. Common practice keeps the alignment-relevant tags (MD, NM, RG) and drops vendor-specific or large optional tags.

After listening to the Attia episode on genetic testing, I’ll have to wait to get my own. I don’t have a question where the answer couldn’t be obtained via biomarker testing, or a condition that genetics can only explain. And I definitely don’t need more data, just more insight.

He says:

Test with intention. Know what you’re looking for. Know what you do when you’ll find it out, and know what you do if you don’t.

3 Likes

I found this podcast interesting because he only really noted polygenic risk scores in passing. Compare this to Eric Topel’s Super Agers book where he extols the possibilities of polygenic risk scores. Attia dedicated almost all the show to a discussion of single/double genetic variants. As I seek to understand my own data, I do see the limitations of single/double genetic variations. Though I do note they seem helpful for pharmaceutical interactions. My guess is that only time will tell the true benefits of WGS and polygenic risk scores.

2 Likes

I’ve met Emily. This looks interesting…

2 Likes

Source: https://x.com/VamsiMootha/status/2059688512765452488?s=20

I wonder if it might be helpful to track the level of mtDNA mutations that a person has?

From Gemini:

Next-Generation Sequencing (NGS) architectures, including Whole Genome Sequencing (WGS) and targeted mitochondrial sequencing, can identify both the specific sequence variants in mitochondrial DNA (mtDNA) and quantify their heteroplasmy levels—the precise ratio of mutated mtDNA molecules to wild-type (normal) mtDNA molecules within a given sample.

Because cells contain hundreds to thousands of copies of the circular mitochondrial genome, quantifying this ratio is critical; the clinical or physiological impact of an mtDNA mutation is directly dependent on the percentage of shifted genomes.

Sequencing Methodologies and Detection Thresholds

The capacity to accurately quantify mtDNA mutation levels depends heavily on the specific sequencing approach and the depth of coverage achieved.

1. Standard Whole Genome Sequencing (WGS)

  • Mechanism: Standard WGS targets the entire cellular DNA extraction. Because mtDNA is highly abundant relative to nuclear DNA, a standard 30x nuclear genome sequencing run naturally yields an “off-target” mitochondrial coverage depth ranging from 100x to over 1,000x.
  • Sensitivity: This depth allows for the reliable detection and quantification of heteroplasmy levels down to approximately 1% to 5%. Any mutation existing below this frequency threshold generally falls into the baseline sequencing noise of standard WGS pipelines.

2. Targeted deep mtDNA Sequencing

  • Mechanism: This approach isolates or selectively amplifies the 16,569 base-pair mitochondrial genome using long-range Polymerase Chain Reaction (LR-PCR) prior to sequencing.
  • Sensitivity: By concentrating sequencing power exclusively on the mitochondrial genome, coverage depth frequently exceeds 10,000x to 100,000x. This extreme depth allows bioinformatic pipelines to confidently identify ultra-low frequency somatic mutations (micro-heteroplasmy) down to 0.1% or lower.

Technical Challenges and Confounding Factors

While technically feasible, accurate quantification of mtDNA mutations via sequencing must overcome two primary biological and methodological hurdles:

Nuclear Mitochondrial Segments (NUMTs)

Over evolutionary timescales, fragments of mtDNA have migrated and integrated into the nuclear genome, becoming pseudogenes known as NUMTs (Nuclear Mitochondrial DNA segments). Standard sequencing read-alignment tools can mistake these ancient, mutated nuclear fragments for true mitochondrial variants. Advanced bioinformatic filtering is mandatory to separate true mitochondrial reads from background NUMT sequences to prevent false-positive heteroplasmy readings.

Tissue Specificity and Mosaicism

The level of mtDNA mutations is not uniform throughout the human body. Somatic mtDNA mutations accumulate unevenly across different organs.

  • Blood (Liquid Biopsy): Easiest to sample, but rapidly dividing hematopoietic cells actively select against highly deleterious mtDNA mutations over time.
  • Post-Mitotic Tissues: High-energy, non-dividing tissues such as skeletal muscle, cardiac muscle, and cerebral cortex typically accumulate significantly higher levels of somatic mtDNA mutations with age.
  • Implication: A standard blood-derived WGS report may show 0% heteroplasmy for a specific mutation that sits at 40% heteroplasmy in the individual’s muscle tissue.

Relevance to Longevity and Geroscience

In the context of healthspan extension, mapping the accumulation of somatic mtDNA mutations provides a direct readout of mitochondrial decay.

Unlike nuclear DNA, mtDNA lacks protective histones and features less redundant repair mechanisms, leaving it highly susceptible to oxidative damage. While inherited mitochondrial diseases typically require a “biochemical threshold” of 60% to 90% heteroplasmy to manifest as clinical pathology, low-level age-related micro-heteroplasmy (sub-5% shifts spread across multiple loci) degrades electron transport chain efficiency, drives cellular senescence, and accelerates the energetic decline characteristic of biological aging.

While I agree with the general consensus, I do have a slightly different perspective. If an anti-aging intervention reduces Alzheimer’s risk at the expense of increasing cardiovascular mortality, the vast majority of people would rule it out immediately. However, for individuals who have undergone whole-genome sequencing and know they are predisposed to Alzheimer’s, this intervention becomes a compelling, albeit difficult, consideration.

1 Like

The problem is a question of which tissues to track it in. White Blood Cells are easy to access, but if you want to track it in any other type of tissue you need to start with a sample.

My personal view is to look at organ biomarkers and use those as a proxy for mtDNA mutations as they control protein production (and therefore function) via acetylation.

Hence for example you can assume if your kidney function is good that enough of the mtDNA in the kidney are in a good state.

Although the body shares mitochondria and therefore mtDNA this is a bit of a stochastic system of distributing the development clock.

2 Likes

That’s a good point. Also if people have very limited ability for prevention (limited time or resources) then knowing what to prioritize can be useful. This is in contrast to most of us that are super interested in longevity and try preventing everything.

I think all sorts of things have trade offs and hormones come to mind most obviously.

Estrogen is neuroprotective (and bone protective) but potentially a cancer risk factor. Say you have strong family history of breast cancer and AD - my sister.

Testosterone if you have a very low risk of CVD or higher risk of osteo or sarcopenia.

Tadalafil has good evidence for CVD protection but potentially harmful for LBD or maybe AD.

Some of these are all debatable but lots of things would seem to be harmful in one category and beneficial in others.

And never discount that some people are okay to die of CVD or cancer but don’t let them have AD. Or whatever preference for 1 disease over another that might not always be rational.

1 Like

$100 off…

1 Like

This is my new go recommendation for whole genome sequencing for two reasons:

First, they put a high value on privacy:

Second, their new .genome data format should greatly improve LLM token efficiency when analyzing a genome:

2 Likes

what are the differences between 3x, 30x, and 100x?

1 Like

These posts in this thread give a good idea:

2 Likes

Perhaps of interest to everyone here doing WGS: Ronjon Nag, Just Announced Superbio.ai - lets you work with large-scale data — genomics, proteomics, etc. 100GB+ — by asking questions in plain English

Anywhere to get whole genome sequencing in Australia?

If you use sequencing.com watch out for unexpected charges as a trial membership gets converted & billed as Premium (without notice). They don’t tell you that you can change your account setting to a free account.

4 Likes

Many thanks for this warning on sequencing.com. I was indeed charged $39 in the months of June and July. When I tried to opt out online, there was no option. When I called customer service, I was offered a number of scammy business offers and never connected to a customer service rep. I just wrote the email now asking for them to cancel the service and reimburse me. Very disappointed in this company.

1 Like

Happy to report that sequencing.com agreed to refund both charges immediately.

I am updating my view of the company from disappointing to somewhat wary.

2 Likes

Some other ideas for approaching this issue:

Source: Haseeb >|< on X: "Spent Saturday analyzing my genome using an LLM. The stack: * Genome sequencing: @nucleusgenomics (cheek swab kit via mail) - Takes about 3 weeks turnaround to get your results. They produce a health report but it's very high-level. You should feed that into your AI as https://t.co/ZKJU1tEp6Z" / X

1 Like