113 min read

Part V · Continuity  ·  Estimated reading time 100 minutes  ·  Prerequisites: Chapters 3, 27, 28

29. Genetics and Heredity

DNA, Genes, and the Molecular Basis of Inheritance

Part V · Continuity  ·  Estimated reading time 100 minutes  ·  Prerequisites: Chapters 3, 27, 28


Case File 29 — "Ninety-Second Percentile"

After her second cardiology visit, Amara Osei, 45, is referred to a genetic counselor. The question is simple to ask and hard to answer: does this run in her family, and if so, what exactly is "this"?

The counselor draws a pedigree.

Person Age Findings
Kwame Osei (Amara's father) Died at 58 Myocardial infarction. Smoker.
Adwoa Mensah (Amara's mother) 78 Hypertension since her 50s; osteoporosis; no coronary events
Amara Osei 45 Coronary artery disease, hypertension, type 2 diabetes, stage 3 CKD
Kofi Osei (Amara's brother) 53 Myocardial infarction at 51; hypertension
Nia Osei-Barrett (Amara's daughter) 24 Pregnant, 28 weeks. Normotensive. Competitive distance runner. No metabolic abnormality
Amara's son 19 Healthy

Amara's lipid panel: LDL cholesterol 168 mg/dL, HDL 38 mg/dL, triglycerides 210 mg/dL, total cholesterol 248 mg/dL. Lipoprotein(a) 42 nmol/L.

Because premature coronary disease in two siblings raises the question of an inherited lipid disorder, a diagnostic panel is sent. The result:

No pathogenic or likely pathogenic variant identified in LDLR, APOB, PCSK9, LDLRAP1, or the other genes on the familial hypercholesterolemia panel. One variant of uncertain significance in LDLR (a missense change not previously reported).

Polygenic risk score for coronary artery disease: 92nd percentile (European-ancestry reference panel; see report limitations).

Three questions to hold on to.

  1. If there is no single broken gene, what exactly is being inherited in this family? What is a polygenic risk score a score of?
  2. Why do Amara and Kofi have disease at 45 and 51 while Nia at 24 has none? Is that difference genetic — or is it time?
  3. Nia cannot change her genome. Given that, what can she actually change, and how much would it matter?

Learning Objectives

By the end of this chapter you should be able to:

  1. Define genome, chromosome, gene, allele, locus, genotype, and phenotype, and use each correctly.
  2. Interpret a karyotype and distinguish autosomes from sex chromosomes.
  3. Describe the structure of DNA — base pairing, antiparallel strands, the double helix — and explain why the structure is the copying mechanism.
  4. Explain semiconservative replication, name the principal enzymes, and describe proofreading and mismatch repair.
  5. Describe transcription and RNA processing, including splicing, and explain how alternative splicing lets one gene make many proteins.
  6. State the properties of the genetic code and describe translation at the ribosome.
  7. Classify mutations by type and predict the consequence of each.
  8. Explain the sickle cell mutation from single base change to clinical phenotype.
  9. Describe meiosis and quantify the three sources of gametic variation.
  10. Explain nondisjunction, name the major aneuploidies, and give the mechanism of the maternal age effect.
  11. Construct and interpret Punnett squares and calculate carrier risks.
  12. Distinguish the four Mendelian inheritance patterns from a pedigree and name a disease example of each.
  13. Explain penetrance, expressivity, pleiotropy, epistasis, anticipation, and imprinting.
  14. Explain mitochondrial inheritance and predict its pedigree pattern.
  15. Define a polygenic risk score correctly, and state three things it is not.
  16. Define heritability as a population statistic and explain why it says nothing about an individual.
  17. Explain gene–environment interaction with a quantified example.
  18. Describe DNA methylation and histone modification, and explain why gene expression is modifiable when sequence is not.
  19. Compare karyotyping, FISH, microarray, and sequencing, and explain the difference between a screening test and a diagnostic test.

29.1 The Vocabulary and the Material

Almost every confusion in genetics is a vocabulary confusion. Fix the words first.

Term Definition In this family
Genome The complete DNA sequence of an organism — 3.1 billion base pairs in ~2 metres of DNA per diploid cell Amara's genome differs from Nia's at millions of positions
Chromosome One continuous DNA molecule with its packaging proteins; 46 in a human somatic cell 22 autosome pairs + XX
Gene A DNA sequence that specifies a functional product (a protein, or a functional RNA); ~19,000–20,000 protein-coding genes, occupying about 1.5% of the genome LDLR is a gene
Locus The physical address of a gene on a chromosome LDLR is at 19p13.2
Allele One of the alternative sequences that can occupy a locus Amara has two LDLR alleles, one from each parent
Genotype The set of alleles an individual carries Amara: no pathogenic LDLR allele
Phenotype The observable characteristic Amara: LDL 168 mg/dL, coronary disease
Homozygous / heterozygous Two identical alleles / two different alleles at a locus
Diploid (2n) / haploid (n) 46 chromosomes / 23 chromosomes Somatic cells / gametes

Two of these carry most of the conceptual weight.

A gene is not a trait. A gene specifies a product. Whether that product produces a recognizable trait depends on the rest of the genome, on the cell it is expressed in, on when it is expressed, and on the environment. "The gene for heart disease" is not a category error because heart disease has no genes — it is a category error because it implies one gene, one disease, and a fixed relationship among them.

Genotype is not phenotype. The relationship between them is where all of clinical genetics lives, and §29.6 through §29.9 are essentially an extended tour of the ways the two can come apart.

The karyotype

A karyotype is an ordered display of an individual's chromosomes, arrested at metaphase when they are maximally condensed, stained, photographed, and arranged by size and centromere position. Human chromosomes are numbered 1 (largest, ~249 million base pairs) to 22 (smallest), plus the sex chromosomes.

  • Autosomes — pairs 1 through 22. Present in two copies in both sexes.
  • Sex chromosomesXX in a typical female, XY in a typical male. The Y is small (~57 Mb, fewer than 100 genes) and carries SRY, the testis-determining factor (§27.2). The X is large (~156 Mb, ~800 genes) and carries genes that have nothing whatever to do with sex — clotting factor VIII, dystrophin, the red and green opsins, glucose-6-phosphate dehydrogenase. This asymmetry is the entire reason X-linked disease behaves the way it does.

The two members of a pair are homologous chromosomes: same length, same centromere position, same gene loci in the same order — but not necessarily the same alleles. One came from each parent. That single fact generates Mendelian genetics.

A normal karyotype is written 46,XX or 46,XY. Abnormalities are written by exception: 47,XY,+21 (Down syndrome), 45,X (Turner syndrome), 47,XXY (Klinefelter syndrome).

Histology · Chromosomes, and Why You Can Only See Them for Twenty Minutes

Look down a microscope at a typical cell and you will not see chromosomes. You will see a nucleus with patchy chromatin: dark, dense heterochromatin at the nuclear periphery and around the nucleolus, and pale, dispersed euchromatin in between. That distinction is functional. Heterochromatin is compacted and transcriptionally silent; euchromatin is open and being read. A plasma cell churning out antibody has an enormously euchromatic, pale nucleus; a mature lymphocyte at rest has a small dense one. You can estimate a cell's transcriptional activity from the appearance of its nucleus, which is one of the more satisfying inferences in histology.

Chromosomes as discrete objects exist only during mitosis, and only because they must. DNA that is 2 metres long in a nucleus 6 micrometres across cannot be pulled to opposite poles as loose thread — it would tangle and break. So before division it is condensed roughly ten-thousand-fold by a hierarchy of packaging: DNA wound 1.65 times around a nucleosome core of eight histones (H2A, H2B, H3, H4 in duplicate), the resulting "beads on a string" coiled into a 30-nanometre fibre, that fibre looped onto a protein scaffold, and the loops coiled again.

Preparing a karyotype exploits this. Cells — classically peripheral blood lymphocytes — are stimulated to divide with phytohemagglutinin, cultured for 72 hours, then arrested in metaphase with colchicine, which poisons the mitotic spindle. A hypotonic solution swells the cells so the chromosomes spread apart, they are fixed and dropped onto a slide, and Giemsa staining produces the alternating light and dark G-bands — dark bands being AT-rich, gene-poor, late replicating; light bands GC-rich, gene-dense, early replicating. A trained cytogeneticist reads 400–550 bands per haploid set.

Two practical consequences follow from the method. First, karyotyping requires living, dividing cells, which is why it takes one to two weeks and why it fails on a non-viable sample. Second, its resolution is limited to about 5–10 million base pairs — a single band may contain a hundred genes. Anything smaller than a band is invisible, which is precisely the gap that FISH and microarray were invented to fill.


29.2 DNA Structure and Replication

The structure is the mechanism

DNA is a polymer of nucleotides, each consisting of a deoxyribose sugar, a phosphate, and one of four nitrogenous bases: adenine and guanine (double-ring purines) and thymine and cytosine (single-ring pyrimidines).

Three structural facts do all the work.

1 · The backbone is directional. The sugar–phosphate backbone runs from a free 5′ phosphate to a free 3′ hydroxyl. DNA polymerase can only add nucleotides to a 3′ hydroxyl, so synthesis runs 5′ → 3′, always, without exception. Essentially every complication in replication follows from this one restriction.

2 · The strands are antiparallel. The two strands run in opposite directions — one 5′→3′ left to right, the other 3′→5′. They must, because a purine must pair with a pyrimidine to keep the helix a constant 2 nm wide, and the geometry of that pairing only works head-to-tail.

3 · Base pairing is specific and complementary. A pairs with T via two hydrogen bonds; G pairs with C via three. Nothing else fits. Therefore each strand contains the complete information needed to rebuild the other, and the molecule's structure is simultaneously its storage format and its copying mechanism. Watson and Crick's understatement — "it has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material" — is the most consequential sentence in twentieth century biology.

Note also that G–C pairs, with three hydrogen bonds, are more stable than A–T pairs with two. GC-rich DNA melts at a higher temperature — which is why the AT-rich regions near replication origins are where the helix opens first, and why GC content determines the annealing temperature of every PCR primer ever designed.

Thread 1 · Structure Determines Function

Everywhere else in this book, structure predicts function. Here, structure is function.

A molecule built from two complementary antiparallel strands is not merely a good way to store information — it is simultaneously the only instruction needed to copy it. Separate the strands and each one specifies its partner completely. Nothing further had to be discovered: the moment the base-pairing geometry was solved, replication became obvious, and Watson and Crick said so in a single sentence at the end of a one-page paper.

The same principle runs all the way down. Replace one hydrophilic glutamate with one hydrophobic valine on the surface of β-globin, and you have created a sticky patch that docks into a pocket on the next molecule — which is the entire mechanism of sickle cell disease (§29.4). Delete three base pairs from CFTR and a protein misfolds; delete four and there is no protein at all. In molecular genetics, shape is not a clue to the job. Shape is the job.

   DNA: ANTIPARALLEL STRANDS AND COMPLEMENTARY BASE PAIRING

   5' end                                                     3' end
     │                                                          │
     P                                                          OH
     │                                                          │
   ┌─S─────A ═══════ T─────S─┐                    A = adenine   │
   │ │     (2 H-bonds)     │ │                    T = thymine   │
   │ P                     P │                    G = guanine   │
   │ │                     │ │                    C = cytosine  │
   ├─S─────G ══════ C──────S─┤                                  │
   │ │    (3 H-bonds)      │ │      A ═══ T   2 hydrogen bonds  │
   │ P                     P │      G ═══ C   3 hydrogen bonds  │
   │ │                     │ │      ↑ purine–pyrimidine keeps   │
   ├─S─────T ═══════ A─────S─┤        the helix a constant      │
   │ │                     │ │        2.0 nm wide               │
   │ P                     P │                                  │
   │ │                     │ │      CHARGAFF: %A = %T, %G = %C  │
   ├─S─────C ══════ G──────S─┤                                  │
   │ │                     │ │                                  │
     OH                    P
     │                     │
   3' end                5' end
     ◄──── strand 2 runs the OPPOSITE way ────►

   S = deoxyribose sugar   P = phosphate   ═ = hydrogen bonds

   ┌────────────────────────────────────────────────────────────────────┐
   │  WHY ANTIPARALLEL MATTERS: DNA polymerase adds nucleotides ONLY to │
   │  a free 3'-OH, so synthesis is ALWAYS 5'→3'. At a replication fork │
   │  moving in one direction, only ONE template can be copied          │
   │  continuously. The other must be copied backwards in pieces.       │
   └────────────────────────────────────────────────────────────────────┘

   REPLICATION IS SEMICONSERVATIVE
   ─────────────────────────────────────────────────────────────────────
        parent duplex          →       two daughter duplexes
        ▬▬▬▬▬▬▬▬▬▬ (old)              ▬▬▬▬▬▬▬▬▬▬ old
        ▬▬▬▬▬▬▬▬▬▬ (old)              ░░░░░░░░░░ new
                                       ░░░░░░░░░░ new
                                       ▬▬▬▬▬▬▬▬▬▬ old
   Each daughter keeps ONE parental strand. Meselson & Stahl, 1958.

   AT THE FORK
   ─────────────────────────────────────────────────────────────────────
                    ╭── LEADING STRAND: continuous, 5'→3', toward fork
       helicase     │   (one primer only)
        unwinds  ═══╪══════════════════════════════════►
       (topo-       │            3'───────────────────5' template
        isomerase   │
        relieves    │            5'───────────────────3' template
        supercoils) ╰── LAGGING STRAND: discontinuous, AWAY from fork
                        ◄──■■■■  ◄──■■■■  ◄──■■■■
                        OKAZAKI FRAGMENTS (100–200 nt in humans)
                        each needs its own RNA primer (primase),
                        primers removed, gaps filled (pol δ),
                        nicks sealed by LIGASE

   FIDELITY — three layers, multiplying
   ─────────────────────────────────────────────────────────────────────
   base-pairing specificity alone .......... ~1 error in 10⁵
   + polymerase 3'→5' PROOFREADING ......... ~1 error in 10⁷
   + MISMATCH REPAIR (MSH/MLH proteins) .... ~1 error in 10⁹–10¹⁰
   Net germline mutation rate: ~1.2 × 10⁻⁸ per base per generation
        = about 70 new mutations in every newborn genome

Figure 29.1 — DNA structure, antiparallel strands, complementary base pairing, and semiconservative replication.

Described: A ladder diagram shows two antiparallel DNA strands. The left strand runs from a five-prime phosphate end at the top to a three-prime hydroxyl at the bottom; the right strand runs in the opposite direction. Each rung is a base pair joining the two sugar–phosphate backbones: adenine pairs with thymine through two hydrogen bonds, and guanine pairs with cytosine through three. Pairing a double-ring purine with a single-ring pyrimidine keeps the helix a constant two nanometres wide, and Chargaff's rule follows — the percentage of adenine equals that of thymine and the percentage of guanine equals that of cytosine. A boxed note explains why antiparallel orientation matters: DNA polymerase can add nucleotides only to a free three-prime hydroxyl, so synthesis always proceeds five-prime to three-prime, and therefore at a replication fork travelling in one direction only one template can be copied continuously while the other must be copied backwards in pieces. A second panel shows semiconservative replication, in which each of the two daughter duplexes retains one parental strand and gains one new one, as demonstrated by Meselson and Stahl in 1958. A third panel diagrams the replication fork: helicase unwinds the duplex while topoisomerase relieves the resulting supercoils; the leading strand is synthesised continuously toward the fork from a single primer; the lagging strand is synthesised discontinuously away from the fork as Okazaki fragments of one hundred to two hundred nucleotides, each requiring its own RNA primer laid down by primase, with primers later removed, gaps filled by polymerase delta, and nicks sealed by ligase. A final panel gives the three multiplying layers of fidelity: base-pairing specificity alone gives about one error in ten to the fifth, polymerase three-prime to five-prime proofreading improves this to one in ten to the seventh, and mismatch repair by the MSH and MLH proteins improves it further to one in ten to the ninth or tenth, giving a net germline mutation rate of about 1.2 times ten to the minus eighth per base per generation, or roughly seventy new mutations in every newborn genome.

Replication, and what happens when repair fails

Human replication begins at 30,000–50,000 origins simultaneously — it must, because a single fork moving at ~50 nucleotides per second would take a month to copy one chromosome. Helicase unwinds the duplex, topoisomerase relieves the torsional strain ahead of it, single-strand binding proteins keep the strands apart, primase lays down short RNA primers, and DNA polymerases δ and ε extend them. On the leading strand this is continuous; on the lagging strand it proceeds in Okazaki fragments whose primers must be excised, whose gaps must be filled, and whose nicks must be sealed by DNA ligase.

The fidelity numbers in Figure 29.1 deserve attention because each layer is a distinct biological system, and each has a disease attached to it when it fails.

Repair system Fixes Failure causes
Proofreading (pol δ/ε exonuclease) Mispaired base just added POLE/POLD1 mutations → hypermutated colorectal and endometrial cancers
Mismatch repair (MLH1, MSH2, MSH6, PMS2) Mismatches and small loops that escaped proofreading Lynch syndrome — high lifetime risk of colorectal, endometrial, and other cancers; tumors show microsatellite instability
Nucleotide excision repair Bulky lesions, especially UV-induced pyrimidine dimers Xeroderma pigmentosum — ~1,000-fold increase in skin cancer; sunlight is disabling
Base excision repair Single damaged or deaminated bases, oxidative damage MUTYH-associated polyposis
Homologous recombination Double-strand breaks, using the sister chromatid BRCA1/BRCA2 — hereditary breast, ovarian, prostate, pancreatic cancer
Non-homologous end joining Double-strand breaks, without a template (error-prone) Immunodeficiency (it is also used deliberately in antibody gene rearrangement)

Notice the pattern: every one of these failures presents as cancer. That is not a coincidence. Cancer is, mechanistically, a disease of accumulated somatic mutation, and a person born with one arm of the repair machinery disabled accumulates mutations faster. The repair systems are the reason your genome is stable enough to be called yours.

Check Your Understanding 29.2

  1. Why must there be a lagging strand at all? Why can the cell not simply run both strands continuously?
  2. Base-pairing alone is accurate to about 1 error in 10⁵. Why is that catastrophically inadequate for a human genome?
Show answers
  1. Because DNA polymerase can add a nucleotide only to a free 3′ hydroxyl group, so it can synthesise only in the 5′→3′ direction, and the two template strands are antiparallel. A replication fork moves in one physical direction; relative to that direction, one template is oriented so its new strand grows toward the fork (leading, continuous) and the other so its new strand must grow away from it (lagging, discontinuous). The cell has no polymerase that works 3′→5′ — and there is a good reason it does not: proofreading requires removing an incorrect nucleotide and resuming, which is chemically straightforward at a 3′ end with the activating triphosphate on the incoming nucleotide, and would be impossible in the reverse direction. The lagging strand is the price of proofreading.
  2. Because the genome is 3.1 billion base pairs. At 1 error in 10⁵ you would introduce roughly 30,000 errors every time a cell divided, in every one of some 10¹³ divisions across a lifetime. Given ~20,000 genes occupying 1.5% of the genome, several hundred of those errors would land in coding sequence per division. No multicellular organism with a long lifespan could tolerate that. Proofreading and mismatch repair take the rate to about 1 in 10⁹–10¹⁰, which is roughly one to three errors per genome per division — low enough that most are in non-coding DNA and silent.

29.3 From Gene to Protein

The central dogma — DNA → RNA → protein — is a statement about information flow. Its most important modern amendment is that the arrow from DNA to protein is not one-to-one. Roughly 20,000 protein-coding genes give rise to well over 100,000 distinct proteins.

   FROM GENE TO PROTEIN — the numbered stages

  ╔══ NUCLEUS ═══════════════════════════════════════════════════════════╗
  ║                                                                      ║
  ║  ① TRANSCRIPTION — INITIATION                                        ║
  ║     transcription factors + RNA polymerase II assemble at the         ║
  ║     PROMOTER (TATA box ~25 bp upstream). ENHANCERS, which may be      ║
  ║     thousands of bp away, loop in to contact the complex.             ║
  ║     ──────────────────────────────────────────────────────────       ║
  ║     DNA  5'─ enhancer ····· PROMOTER │ EXON1 ─ intron ─ EXON2 ─ ...   ║
  ║                                                                      ║
  ║  ② ELONGATION — pol II reads the TEMPLATE strand 3'→5' and builds     ║
  ║     pre-mRNA 5'→3' at 20–50 nt/s. RNA uses URACIL in place of         ║
  ║     thymine and ribose in place of deoxyribose.                       ║
  ║                                                                      ║
  ║  ③ TERMINATION — polyadenylation signal AAUAAA; transcript cleaved.   ║
  ║                                                                      ║
  ║  ④ PROCESSING — three modifications, all essential                    ║
  ║     • 5' CAP (7-methylguanosine): ribosome binding + stability        ║
  ║     • 3' POLY-A TAIL (~200 A): stability + export                     ║
  ║     • SPLICING by the spliceosome (snRNPs): INTRONS excised at the    ║
  ║       GU...AG boundaries, EXONS joined                                ║
  ║                                                                      ║
  ║       pre-mRNA  [E1]──intron──[E2]──intron──[E3]──intron──[E4]        ║
  ║                       ╲          ╲            ╲                       ║
  ║       mature mRNA  cap-[E1][E2][E3][E4]-AAAAA...                      ║
  ║                                                                      ║
  ║       ALTERNATIVE SPLICING: the SAME pre-mRNA can be spliced          ║
  ║       differently in different cells —                                ║
  ║               cap-[E1][E2]    [E4]-AAAA   (exon 3 skipped)            ║
  ║               cap-[E1]    [E3][E4]-AAAA   (exon 2 skipped)            ║
  ║       ~95% of multi-exon human genes are alternatively spliced.       ║
  ║       ONE GENE → MANY PROTEINS. This is why 20,000 genes suffice.     ║
  ╚═══════════════════════════════╤══════════════════════════════════════╝
                                  │ ⑤ EXPORT through nuclear pore
  ╔══ CYTOPLASM ══════════════════╧══════════════════════════════════════╗
  ║  ⑥ TRANSLATION                                                       ║
  ║     INITIATION: small (40S) subunit + initiator tRNA scan to the      ║
  ║        first AUG → large (60S) subunit joins                          ║
  ║     ELONGATION: three sites on the ribosome —                         ║
  ║           A site (Arrival of charged tRNA)                            ║
  ║           P site (Peptide chain held here)                            ║
  ║           E site (Exit of the empty tRNA)                             ║
  ║        peptidyl transferase forms the bond — and it is rRNA, not      ║
  ║        protein: the ribosome is a RIBOZYME. ~2–6 aa/s.                ║
  ║     TERMINATION: a STOP codon (UAA, UAG, UGA) is read by a release    ║
  ║        factor — no tRNA exists for them.                              ║
  ║                                                                      ║
  ║     mRNA  5'-AUG GCC UUA AGC ... UAA-3'                               ║
  ║               Met Ala Leu Ser ... STOP                                ║
  ║                                                                      ║
  ║  ⑦ POST-TRANSLATIONAL MODIFICATION                                    ║
  ║     folding (chaperones) · cleavage (proinsulin → insulin + C-pep)    ║
  ║     glycosylation · phosphorylation · hydroxylation (collagen —       ║
  ║     needs vitamin C) · γ-carboxylation (clotting factors — needs      ║
  ║     vitamin K; the target of warfarin) · targeting · degradation      ║
  ║     by the ubiquitin–proteasome system                                ║
  ╚══════════════════════════════════════════════════════════════════════╝

Figure 29.2 — The flow from gene to protein, as seven numbered stages.

Described: A two-compartment diagram separates nuclear from cytoplasmic events. In the nucleus, stage one is transcription initiation, in which transcription factors and RNA polymerase II assemble at the promoter with its TATA box about twenty-five base pairs upstream, while enhancers thousands of base pairs away loop in to contact the complex. Stage two is elongation, in which polymerase reads the template strand three-prime to five-prime and builds pre-messenger RNA five-prime to three-prime at twenty to fifty nucleotides per second, using uracil in place of thymine and ribose in place of deoxyribose. Stage three is termination at the polyadenylation signal AAUAAA. Stage four is processing, comprising three essential modifications: a five-prime seven-methylguanosine cap for ribosome binding and stability, a three-prime poly-A tail of about two hundred adenines for stability and export, and splicing by the spliceosome, which excises introns at their GU and AG boundaries and joins exons. Alternative splicing is illustrated by showing the same pre-messenger RNA yielding different mature transcripts in different cells, one skipping exon three and another skipping exon two; about ninety-five percent of multi-exon human genes are alternatively spliced, so one gene makes many proteins, which is why twenty thousand genes suffice. Stage five is export through the nuclear pore. In the cytoplasm, stage six is translation: initiation, in which the forty-S subunit and initiator transfer RNA scan to the first AUG before the sixty-S subunit joins; elongation through the ribosome's three sites — the A site where charged transfer RNA arrives, the P site where the peptide chain is held, and the E site from which empty transfer RNA exits — with peptide bond formation catalysed by ribosomal RNA rather than protein, making the ribosome a ribozyme, at two to six amino acids per second; and termination when a stop codon, UAA, UAG, or UGA, is read by a release factor, since no transfer RNA exists for them. Stage seven is post-translational modification, including chaperone-assisted folding, proteolytic cleavage such as proinsulin to insulin and C-peptide, glycosylation, phosphorylation, hydroxylation of collagen requiring vitamin C, gamma-carboxylation of clotting factors requiring vitamin K and inhibited by warfarin, subcellular targeting, and degradation by the ubiquitin–proteasome system.

One gene, many proteins

The human genome contains about 20,000 protein-coding genes — roughly the same number as a nematode worm with 959 cells. The resolution of that apparent paradox is that a human gene is not a single instruction.

Alternative splicing is the main mechanism. About 95% of human multi-exon genes are alternatively spliced, producing on average several distinct mRNAs and therefore several distinct proteins. Cardiac and skeletal troponin T are alternative splice products of the same gene. So are the membrane-bound and secreted forms of immunoglobulin heavy chain — the same B cell switches from making a receptor to making an antibody by changing where it cleaves and polyadenylates one transcript.

Two further amplifications matter. Alternative promoters and polyadenylation sites produce transcripts with different regulatory ends. And post-translational modification multiplies again: the same polypeptide, differently phosphorylated or glycosylated, is functionally a different molecule. Proopiomelanocortin is cleaved into ACTH, β-endorphin, and MSH — three hormones with entirely different actions from one gene product, cut differently in different cells (Chapter 16).

The genetic code itself has five properties worth stating precisely:

  • Triplet. Three bases per codon; 4³ = 64 codons.
  • Degenerate (redundant). 61 codons specify 20 amino acids, so most amino acids have several codons — usually differing at the third position (wobble). This is why many third-position mutations are silent, and it is a built-in error buffer.
  • Unambiguous. Each codon specifies exactly one amino acid.
  • Non-overlapping and comma-less. Read three at a time from a fixed start, which is why a frameshift is so destructive.
  • Nearly universal. The same code runs in bacteria, plants, and humans — which is what makes recombinant human insulin producible in E. coli. The important exception is the mitochondrial genome, which uses a slightly different code.

29.4 Mutation

A mutation is a heritable change in DNA sequence. Most are harmless; a few are useful; the ones that cause disease are the ones we study, which badly distorts intuition about how common harm is.

Type What changes Consequence Example
Silent (synonymous) One base; codon still specifies the same amino acid Usually none — though it can alter splicing or translation speed Many third-position changes
Missense One base; one amino acid replaced Ranges from nothing to lethal, depending on where and what Sickle cell (Glu→Val); LDLR missense in familial hypercholesterolemia
Nonsense One base; codon becomes a STOP Truncated protein, usually non-functional; the transcript is often destroyed by nonsense-mediated decay Duchenne muscular dystrophy; β⁰-thalassemia
Frameshift Insertion or deletion of a number of bases not divisible by 3 Every codon downstream is misread; almost always a premature stop follows. Usually severe Tay-Sachs (HEXA); many BRCA1 variants
In-frame indel Insertion or deletion of a multiple of 3 One or a few residues lost or gained; reading frame preserved ΔF508 in cystic fibrosis — 3 bases deleted, one phenylalanine lost, protein misfolds
Splice-site Base change at an intron boundary Exon skipped or intron retained; often a frameshift downstream β-thalassemia; some LDLR variants
Copy number variant A whole segment duplicated or deleted Dosage change across many genes PMP22 duplication → Charcot-Marie-Tooth 1A; deletion → hereditary neuropathy with liability to pressure palsies
Repeat expansion A short repeat unit copied too many times Number of repeats determines severity and age of onset; expands further between generations Huntington (CAG), fragile X (CGG), myotonic dystrophy (CTG)
Regulatory Promoter, enhancer, or untranslated region Protein is normal but made in the wrong amount, place, or time Most variants contributing to polygenic disease

That last row is important and easy to skip past. The great majority of common disease- associated variants found by genome-wide association studies are not in coding sequence at all. They are regulatory. They change how much of a normal protein is made. That fact is what makes polygenic disease look so different from Mendelian disease, and it is the key to §29.8.

A worked example: sickle cell disease

The β-globin gene, HBB, sits at 11p15.4. In the sickle allele a single base changes in the sixth codon: GAG → GTG. The consequence chain is worth following link by link, because it is the cleanest demonstration in medicine of a genotype becoming a phenotype.

  1. The base. One A becomes a T. One of 3.1 billion.
  2. The codon. GAG (glutamate) becomes GTG (valine).
  3. The chemistry. Glutamate is negatively charged and hydrophilic and sits comfortably on the protein's surface. Valine is hydrophobic. The substitution creates a sticky hydrophobic patch on the outside of a soluble protein.
  4. The conformation. When hemoglobin releases oxygen it changes shape, and in the deoxy conformation a complementary hydrophobic pocket appears on the β chain of a neighboring molecule. The valine patch docks into it.
  5. The polymer. Molecules link into long 14-stranded helical fibers that stiffen and distort the cell into a crescent.
  6. The cell. Sickled cells are rigid and cannot deform to pass through capillaries 5 µm across; they adhere to endothelium; their membranes are damaged by repeated sickling–unsickling cycles.
  7. The disease. Vaso-occlusion → ischemic pain crises, stroke, acute chest syndrome, splenic infarction. Hemolysis → anemia, gallstones, and free hemoglobin that scavenges nitric oxide, producing pulmonary hypertension.
  8. The population. Heterozygotes (sickle trait) are largely asymptomatic and are protected against severe Plasmodium falciparum malaria, which is why the allele reaches frequencies above 10% where malaria is endemic. This is balanced polymorphism: the same allele is beneficial in one copy and harmful in two.

Every step is mechanical. Nothing is metaphorical. Chapter 17 develops the hematology.

Predict This

Two mutations occur in the same gene. Mutation A deletes 3 base pairs from the middle of exon 4. Mutation B deletes 4 base pairs from the same position.

Commit before reading on: which is likely to be worse, and by how much?

(Answer: B, and the difference is enormous rather than incremental. Deleting 3 base pairs removes exactly one codon — the protein loses one amino acid and everything downstream is normal. That is ΔF508 in cystic fibrosis, which produces a full-length protein that merely misfolds. Deleting 4 base pairs shifts the reading frame, so every codon downstream is misread, a premature stop codon almost always appears within 50 codons, and the transcript is usually destroyed by nonsense-mediated decay before it is ever translated. One base pair of difference separates "one amino acid missing" from "no protein at all." This is why the divisible-by-three rule is worth memorising.)


29.5 Meiosis and the Sources of Variation

Meiosis has two jobs: halve the chromosome number, so that fertilization restores it rather than doubling it, and generate variation, so that no two gametes are alike. One nuclear division would accomplish the first. It takes two divisions plus a recombination step to accomplish the second.

   MEIOSIS — two divisions, three sources of variation

   PARENT CELL  2n = 46, after S phase each chromosome has 2 sister
   chromatids                    ▌▌  ▌▌      (maternal ▌▌  paternal ▌▌)

   ╔═════════════ MEIOSIS I — REDUCTIONAL ═══════════════════════════════╗
   ║  PROPHASE I  (long: leptotene→zygotene→pachytene→diplotene→diakinesis)║
   ║    • homologues pair: SYNAPSIS, held by the synaptonemal complex     ║
   ║    • ★ SOURCE 1: CROSSING OVER at chiasmata                          ║
   ║        maternal  ▌▌▌▌▌▌▌▌▌▌      ▌▌▌▌░░░░░░                          ║
   ║                     ╳        →                                       ║
   ║        paternal  ░░░░░░░░░░      ░░░░▌▌▌▌▌▌                          ║
   ║        ~45–55 crossovers per human meiosis (more in oocytes).        ║
   ║        Every chromosome you pass on is a MOSAIC of your two parents. ║
   ║                                                                      ║
   ║  METAPHASE I  homologous PAIRS line up at the equator                ║
   ║    • ★ SOURCE 2: INDEPENDENT ASSORTMENT                              ║
   ║        each of 23 pairs orients maternal-left or maternal-right      ║
   ║        INDEPENDENTLY → 2²³ = 8,388,608 combinations per gamete       ║
   ║                                                                      ║
   ║  ANAPHASE I  HOMOLOGUES separate; sister chromatids stay together    ║
   ║              ← THIS is what makes it reductional: 2n → n             ║
   ║  TELOPHASE I → two cells, each n = 23, each chromosome still 2       ║
   ║              chromatids (and those chromatids are no longer          ║
   ║              identical, because of crossing over)                    ║
   ╚══════════════════════════════════════════════════════════════════════╝
   ╔═════════════ MEIOSIS II — EQUATIONAL (a "mitosis" of n cells) ═══════╗
   ║  no DNA replication. SISTER CHROMATIDS separate.                     ║
   ║  → FOUR haploid cells, all genetically different                     ║
   ╚══════════════════════════════════════════════════════════════════════╝

   ★ SOURCE 3: RANDOM FERTILIZATION — any of ~8.4 million maternal
     combinations × any of ~8.4 million paternal = 7 × 10¹³ zygotes,
     BEFORE crossing over is counted. With crossing over, effectively
     infinite. Only monozygotic twins escape this.

   ══ NONDISJUNCTION — when separation fails ══════════════════════════════
     NORMAL              NONDISJUNCTION in MI     NONDISJUNCTION in MII
       ▌▌ ░░                  ▌▌░░  --                ▌▌ ░░
       │   │                   │     │                │   │
      n  n  n  n            n+1 n+1  n-1 n-1        n  n  n+1  n-1
     all normal            ALL 4 gametes abnormal   2 normal, 2 abnormal

     Trisomy = n+1 gamete fertilised → 47 chromosomes
     Monosomy = n-1 gamete fertilised → 45 chromosomes (nearly all lethal;
                45,X is the ONLY viable human monosomy)

Figure 29.3 — Meiosis I and II, the three sources of variation, and nondisjunction.

Described: A parent cell with forty-six chromosomes, each duplicated into two sister chromatids after S phase, enters meiosis I, the reductional division. In prophase I — which passes through leptotene, zygotene, pachytene, diplotene, and diakinesis — homologous chromosomes pair in synapsis, held by the synaptonemal complex, and undergo crossing over at chiasmata, the first source of variation, with roughly forty-five to fifty-five crossovers per human meiosis and more in oocytes than in spermatocytes, so that every chromosome passed on is a mosaic of the parent's own two parents. In metaphase I, homologous pairs line up at the equator and each of the twenty-three pairs orients independently, the second source of variation, giving two to the twenty-third power, or 8,388,608, combinations per gamete. In anaphase I the homologues separate while sister chromatids stay together, which is what makes the division reductional, and telophase I yields two haploid cells whose chromosomes still consist of two now non-identical chromatids. Meiosis II is equational, with no DNA replication; sister chromatids separate, yielding four haploid cells, all genetically different. The third source of variation is random fertilization, multiplying roughly 8.4 million maternal combinations by 8.4 million paternal to give about seven times ten to the thirteenth possible zygotes before crossing over is even counted; only monozygotic twins escape this. A final panel contrasts normal segregation, which yields four normal haploid gametes, with nondisjunction in meiosis I, which yields four abnormal gametes — two with an extra chromosome and two missing one — and nondisjunction in meiosis II, which yields two normal gametes, one with an extra chromosome and one missing one. Fertilisation of an n-plus-one gamete produces trisomy with forty-seven chromosomes, and of an n-minus-one gamete produces monosomy with forty-five, which is nearly always lethal; 45,X is the only viable human monosomy.

Aneuploidy

Nondisjunction — failure of chromosomes or chromatids to separate — produces gametes with an extra or missing chromosome. About 20–30% of all human conceptions are aneuploid, and the great majority are lost before or shortly after implantation. Only a few aneuploidies are compatible with live birth.

Karyotype Name Incidence Features
47,+21 Down syndrome 1 in 700 live births Intellectual disability, characteristic facies, hypotonia; ~50% have congenital heart disease (especially AV canal defects); early Alzheimer pathology (APP is on chromosome 21)
47,+18 Edwards syndrome 1 in 5,000 Severe; most die in the first year
47,+13 Patau syndrome 1 in 16,000 Severe midline defects; most die in the first months
45,X Turner syndrome 1 in 2,500 female births Short stature, streak gonads, primary amenorrhea, coarctation, webbed neck. ~99% of 45,X conceptions miscarry. Usually from paternal loss
47,XXY Klinefelter syndrome 1 in 600 male births Tall, small firm testes, infertility, gynecomastia, low testosterone; often undiagnosed until infertility work-up
47,XYY 1 in 1,000 males Tall stature; usually otherwise unremarkable
47,XXX 1 in 1,000 females Usually unremarkable; mild learning differences

Note the pattern in the sex chromosome rows: they are all far milder than the autosomal trisomies. Two mechanisms explain it — X-inactivation silences all but one X in every cell (see the Development sidebar in §29.7), and the Y chromosome carries very few genes — so altering sex chromosome number changes gene dosage far less than adding an entire autosome.

Only about 95% of Down syndrome is free trisomy 21 from nondisjunction; roughly 3–4% arises from a Robertsonian translocation, in which chromosome 21 is fused to another acrocentric chromosome. This matters enormously for counseling: free trisomy has a low recurrence risk (~1%), whereas a parent carrying a balanced 14;21 translocation has a recurrence risk of 10–15% if the mother carries it, and a 21;21 translocation carrier has a recurrence risk of 100%. A karyotype therefore answers a question that a diagnosis alone cannot.

Aging · Telomeres, Somatic Mutation, and a Clone in the Blood

The genome ages, and it does so in three measurable ways that between them explain a good deal of Adwoa's biology at 78 and, increasingly, of Amara's at 45.

Telomere shortening. Chromosome ends carry tandem TTAGGG repeats — 10–15 kilobases at birth — capped by a protein complex (shelterin) that keeps them from being read as double-strand breaks. Because DNA polymerase cannot replicate the extreme 5′ end of the lagging strand (the end-replication problem), each division loses 50–200 base pairs. At a critical length the cap fails, the cell registers a DNA damage response, and it enters replicative senescence — the Hayflick limit. Senescent cells do not simply stop; they adopt a senescence-associated secretory phenotype, releasing IL-6, IL-8, and matrix metalloproteinases that drive chronic low-grade inflammation. Telomerase reverses the loss but is expressed only in germ cells, stem cells, activated lymphocytes — and in about 90% of cancers, which is how they become immortal. Short leukocyte telomeres are associated with cardiovascular disease and mortality, but the association is modest and the causal direction is contested; treat commercial telomere testing accordingly.

Somatic mutation accumulation. Every cell acquires mutations independently, at roughly 15–40 per year in a typical somatic tissue. By 70, a single skin or oesophageal cell may carry several thousand. Most are harmless passengers. The consequence is that an aged tissue is not a uniform population but a mosaic of competing clones, some carrying driver mutations in cancer genes — an entirely normal finding in histologically normal tissue.

Clonal hematopoiesis. The most clinically striking version. With age, a hematopoietic stem cell acquires a mutation — most often in DNMT3A, TET2, or ASXL1, all epigenetic regulators — that gives it a small competitive advantage. Its descendants expand until they constitute a measurable fraction of circulating blood cells. This is clonal hematopoiesis of indeterminate potential (CHIP): no cytopenia, no leukemia, just a clone. It is present in under 1% of people under 40, about 10% over 70, and more than 20% over 90.

CHIP raises the risk of hematological malignancy about tenfold — but the absolute risk remains low, around 0.5–1% per year. What is more surprising, and more relevant to Amara, is that CHIP carries a roughly doubled risk of coronary artery disease, independent of every conventional risk factor. The mechanism appears to be inflammatory: TET2-mutant macrophages in the atherosclerotic plaque secrete excess IL-1β and IL-6, accelerating plaque growth. Mouse experiments transplanting Tet2-deficient marrow reproduce the acceleration.

This is a genuinely new idea and worth pausing on. It means that some of a 60-year-old's cardiovascular risk is not inherited and not behavioral, but arises from mutations acquired in blood stem cells during that person's own lifetime — somatic genetics producing systemic disease, sitting in a category that neither the family history nor the polygenic score captures.

Clinical Connection · Down Syndrome — One Extra Chromosome, and Why Age Is the Risk Factor

Trisomy 21 is the most common autosomal aneuploidy compatible with life, at about 1 in 700 births. About 95% is free trisomy from nondisjunction; of those, 90–95% are maternal in origin, and about 75% arise in meiosis I.

The maternal-origin and meiosis-I concentration together point at one mechanism, and it is the one Chapter 28 introduced. A female's oocytes enter prophase I in fetal life and remain arrested there for one to five decades. During arrest, homologous chromosomes are held together by cohesin complexes loaded before birth and never replenished, protected at the centromere by shugoshin. Cohesin degrades over time. In an oocyte ovulated at 42, the bivalents have been held together by 42-year-old protein, and they destabilise: chiasmata slip toward the chromosome ends, kinetochores fail to bi-orient on the meiosis I spindle, and segregation goes wrong. The risk curve — roughly 1 in 1,500 at age 20, 1 in 350 at 35, 1 in 100 at 40, 1 in 30 at 45 — is a curve about protein degradation over decades, not about the uterus.

Two clinical points follow. Most babies with Down syndrome are born to women under 35, not because risk per pregnancy is higher in that group but because far more pregnancies occur in it — a base-rate effect that repeatedly confuses screening discussions. And because free trisomy recurrence risk is only about 1%, whereas translocation carriers face 10–100%, the parental karyotype changes counseling completely.

The clinical picture is a lesson in gene dosage: the phenotype is caused not by an abnormal gene but by 1.5-fold expression of a few hundred normal ones. Congenital heart disease in roughly half (particularly atrioventricular septal defects), duodenal atresia, hypothyroidism, atlantoaxial instability, leukemia risk raised 10–20 fold, and — because the amyloid precursor protein gene APP lies on chromosome 21 — Alzheimer neuropathology in essentially all individuals by age 40. Median survival has risen from about 25 years in 1983 to about 60 today, almost entirely through cardiac surgery and general medical care.

Check Your Understanding 29.5

  1. Nondisjunction in meiosis I produces four abnormal gametes; in meiosis II it produces two normal and two abnormal. Explain why.
  2. Why is 45,X the only viable human monosomy, when several trisomies are viable?
Show answers
  1. In meiosis I, the error occurs before the cell has divided at all, so both daughter cells inherit the mistake — one gets both homologues (n+1), the other gets neither (n−1) — and each then divides in meiosis II to produce two gametes like itself. All four are abnormal. In meiosis II, the first division was correct, so both daughter cells are normally haploid; the error occurs in only one of them, and only that cell's two products are abnormal (one n+1, one n−1). The other daughter cell produces two normal gametes. The distinction is diagnostically useful, since parental origin and division of error can be determined from polymorphic markers.
  2. Because of dosage. A monosomy leaves a single copy of every gene on that chromosome, and for hundreds of genes, half the normal product is insufficient — autosomal monosomies are uniformly lethal very early. The X is the exception because of X-inactivation: in a typical female, one X is already silenced in every cell, so a 45,X individual is not far from the functional dosage of a 46,XX individual for most X-linked genes. What she lacks is the product of the pseudoautosomal genes that escape inactivation — notably SHOX, which is why short stature is the most consistent feature of Turner syndrome — and the second X required for normal ovarian maintenance, which is why the gonads regress to fibrous streaks. Even so, about 99% of 45,X conceptions are lost, so "viable" is relative.

29.6 Mendelian Inheritance

Gregor Mendel's experiments succeeded because he chose traits controlled by single genes with two clean alleles and complete dominance. Most human traits are not like that — which is why §29.7 and §29.8 exist. But single-gene inheritance is the necessary foundation, and roughly 7,000 human conditions follow it.

Dominant means the phenotype appears when only one copy of the allele is present. Recessive means the phenotype appears only when both copies are present. Note that these are properties of the phenotype, not of the allele: sickle cell anemia is recessive at the level of clinical disease, codominant at the level of hemoglobin electrophoresis (both HbA and HbS are detectable in a heterozygote), and dominant with respect to malaria resistance. The same allele can be dominant, recessive, or codominant depending on which phenotype you measure. This is the single most useful thing to know about the word "dominant."

Punnett squares and carrier arithmetic

   A PUNNETT SQUARE FOR A CARRIER COUPLE
   Cystic fibrosis: autosomal recessive.  F = normal allele, f = CFTR variant

                            FATHER  Ff  (carrier, unaffected)
                                 ┌──────────┬──────────┐
                                 │    F     │    f     │
                    ┌────────────┼──────────┼──────────┤
                    │            │   FF     │   Ff     │
   MOTHER  Ff       │     F      │  normal  │ CARRIER  │
   (carrier,        │            │  1/4     │  1/4     │
    unaffected)     ├────────────┼──────────┼──────────┤
                    │            │   Ff     │   ff     │
                    │     f      │ CARRIER  │ AFFECTED │
                    │            │  1/4     │  1/4     │
                    └────────────┴──────────┴──────────┘

   PER PREGNANCY:  1/4 affected · 1/2 carrier · 1/4 homozygous normal
                   = the classic 1 : 2 : 1 genotypic, 3 : 1 phenotypic ratio

   ── THE THREE ARITHMETIC TRAPS ────────────────────────────────────────
   ① Each pregnancy is INDEPENDENT. Three unaffected children do not make
      the fourth safer. The coin has no memory.
   ② An UNAFFECTED child of two carriers is 2/3 likely to be a carrier,
      not 1/2 — because the "affected" box has been excluded by
      observation. (Ff + Ff + FF, minus ff = 2 carriers out of 3.)
   ③ POPULATION risk requires the carrier frequency. In Northern European
      ancestry, CF carrier frequency ≈ 1/25.

   ── WORKED: a healthy man whose SISTER has CF, with a partner of
      Northern European ancestry and no family history ───────────────────
      P(man is a carrier)      = 2/3        (trap ② — his parents are
                                             obligate carriers)
      P(partner is a carrier)  = 1/25       (population frequency)
      P(both carriers AND affected child) = 2/3 × 1/25 × 1/4 = 2/300
                                          = 1 in 150 per pregnancy
      Now the partner is CARRIER-SCREENED and is NEGATIVE. A panel that
      detects 90% of variants leaves a residual risk:
      P(carrier | negative test) = (1/25 × 0.10) ÷ [(1/25 × 0.10)
                                    + (24/25)]  ≈ 1/240
      revised risk = 2/3 × 1/240 × 1/4 ≈ 1 in 1,440 per pregnancy
      A negative screen does not give zero. It gives a smaller number.

Figure 29.4 — A Punnett square for a carrier couple, with the three arithmetic traps and a worked risk calculation.

Described: A two-by-two Punnett square crosses two heterozygous carriers of a cystic fibrosis variant, each with genotype capital-F lowercase-f. The four boxes give one quarter FF, homozygous normal; two quarters Ff, unaffected carriers; and one quarter ff, affected — the classic one-to-two-to-one genotypic and three-to-one phenotypic ratio, applying to each pregnancy independently. Three arithmetic traps are listed. First, each pregnancy is independent, so three unaffected children do not reduce the risk to the fourth. Second, an unaffected child of two carriers has a two-thirds probability of being a carrier rather than one half, because observing that the child is unaffected excludes the affected box, leaving two carrier outcomes out of three possible ones. Third, population risk calculations require the carrier frequency, which for cystic fibrosis in people of Northern European ancestry is about one in twenty-five. A worked example follows for a healthy man whose sister has cystic fibrosis and whose partner has Northern European ancestry and no family history: the man's carrier probability is two thirds, the partner's is one in twenty-five, and the probability of an affected child in any pregnancy is two thirds times one twenty-fifth times one quarter, or about one in one hundred fifty. If the partner then has a negative carrier screen on a panel detecting ninety percent of variants, her residual carrier probability falls to about one in two hundred forty by Bayesian revision, and the couple's per-pregnancy risk falls to about one in one thousand four hundred forty. The example closes by emphasising that a negative screen yields a smaller number, never zero.

The four patterns and how to tell them apart

   THE FOUR MENDELIAN PEDIGREE PATTERNS
   □ unaffected male   ○ unaffected female   ■ affected male   ● affected
   female   ⊡ carrier male   ⊙ carrier female   │ mating   ┬ offspring

  ══ AUTOSOMAL DOMINANT ══════════ ══ AUTOSOMAL RECESSIVE ═══════════════
   I    ■────○                      I    ⊡────⊙        both parents
        │                                │            unaffected carriers
   II ┌─┼──┬──┐                     II ┌─┼──┬──┬──┐
      ■  □  ●  □                       ⊙  ⊡  ●  □
      │                                      │
   III┌┴─┬──┐                        III ────┴────  (none affected unless
      ●  □  ■                                       partner is a carrier)

   RULES                            RULES
   • EVERY affected has an          • Affected individuals often have
     affected PARENT (unless          UNAFFECTED parents — the pattern
     new mutation)                    "SKIPS" generations
   • NO SKIPPING                    • ~25% of sibs of an affected are
   • M = F affected equally           affected; M = F equally
   • MALE-to-MALE transmission      • CONSANGUINITY raises the odds
     OCCURS (excludes X-linked)       sharply
   • 50% of offspring affected      • Often ENZYME deficiencies (50% of
   • Often STRUCTURAL proteins        normal enzyme is usually enough,
     or receptors (one bad copy       so carriers are well)
     disrupts a complex)
   EXAMPLES Huntington · Marfan ·   EXAMPLES cystic fibrosis · sickle
   familial hypercholesterolemia ·  cell · Tay-Sachs · phenylketonuria ·
   achondroplasia · NF1 · ADPKD     hemochromatosis · Wilson disease

  ══ X-LINKED RECESSIVE ══════════ ══ X-LINKED DOMINANT ═════════════════
   I    □────⊙  carrier mother      I    ■────○  affected FATHER
        │                                │
   II ┌─┼──┬──┬──┐                  II ┌─┼──┬──┬──┐
      ■  □  ⊙  ○                       ●  ●  ●  □   ALL daughters
      │        │                       │            affected,
   III└──►     └──► carrier            │            NO sons
      NO male-to-male                III┌┴─┬──┐
                                        ●  □  ■  (affected mother →
   RULES                                          50% of ALL children)
   • MALES predominate heavily      RULES
   • NEVER male-to-male (a father   • NO male-to-male (as X-linked
     gives his son a Y)               recessive)
   • ALL daughters of an affected   • An affected FATHER transmits to
     male are obligate carriers       ALL daughters and NO sons ← the
   • Carrier mother → 50% of sons     single most diagnostic rule
     affected, 50% of daughters     • An affected MOTHER transmits to
     carriers                         50% of children of either sex
   • Affected females need TWO      • FEMALES affected roughly 2× as
     hits — rare                      often as males, usually more mildly
   EXAMPLES hemophilia A & B ·        (X-inactivation mosaicism)
   Duchenne muscular dystrophy ·    • Male lethality in some conditions
   red-green colour blindness ·     EXAMPLES X-linked hypophosphatemic
   G6PD deficiency · Fabry          rickets · Rett syndrome (male lethal)
                                    · incontinentia pigmenti

   ═════ THE THREE QUESTIONS THAT SORT ANY PEDIGREE ══════════════════════
   1. Does it SKIP generations?     no → dominant | yes → recessive
   2. Is there MALE-to-MALE
      transmission?                 yes → AUTOSOMAL (excludes X-linked)
   3. Are the sexes affected
      EQUALLY?                      no, males ≫ → X-linked recessive
                                    no, females ~2× → X-linked dominant

Figure 29.5 — The four Mendelian inheritance patterns, side by side, with their identifying rules.

Described: Four pedigrees are shown side by side with standard symbols — squares for males, circles for females, filled for affected, dotted for carriers. In autosomal dominant inheritance, every affected individual has an affected parent unless the variant is new, the trait does not skip generations, males and females are affected equally, male-to-male transmission occurs and therefore excludes X-linkage, and half of an affected person's offspring are affected; the genes involved are often structural proteins or receptors, and examples include Huntington disease, Marfan syndrome, familial hypercholesterolemia, achondroplasia, neurofibromatosis type 1, and autosomal dominant polycystic kidney disease. In autosomal recessive inheritance, affected individuals commonly have unaffected carrier parents so the trait appears to skip generations, about a quarter of the siblings of an affected person are affected, males and females are affected equally, and consanguinity sharply raises the odds; the genes involved are typically enzymes, because half the normal enzyme level usually suffices, and examples include cystic fibrosis, sickle cell disease, Tay-Sachs disease, phenylketonuria, hemochromatosis, and Wilson disease. In X-linked recessive inheritance, males predominate heavily, male-to-male transmission never occurs because a father gives his son a Y chromosome, all daughters of an affected male are obligate carriers, a carrier mother transmits to half her sons as affected and half her daughters as carriers, and affected females are rare because they require two hits; examples include hemophilia A and B, Duchenne muscular dystrophy, red-green colour blindness, glucose-6-phosphate dehydrogenase deficiency, and Fabry disease. In X-linked dominant inheritance there is likewise no male-to-male transmission, an affected father transmits to all his daughters and none of his sons — the single most diagnostic rule — an affected mother transmits to half her children of either sex, females are affected about twice as often as males and usually more mildly because of X-inactivation mosaicism, and some conditions are lethal in males; examples include X-linked hypophosphatemic rickets, Rett syndrome, and incontinentia pigmenti. A closing panel gives three sorting questions: whether the trait skips generations, which distinguishes dominant from recessive; whether male-to-male transmission occurs, which establishes autosomal inheritance; and whether the sexes are affected equally, with a heavy male excess indicating X-linked recessive and a roughly twofold female excess indicating X-linked dominant.

Carrier states, penetrance, and expressivity

Three concepts break the tidy genotype-to-phenotype mapping even within single-gene disease.

Carrier. A heterozygote for a recessive allele, clinically unaffected. Carriers are common: about 1 in 25 people of Northern European ancestry carries a CF variant, and on a large expanded panel roughly one person in three or four carries something. Being a carrier is normal.

Penetrance is the proportion of people with a genotype who show any phenotype. It is a population statistic and it is all-or-none per person. Huntington disease has essentially 100% penetrance given normal lifespan; a pathogenic BRCA1 variant has 55–72% lifetime penetrance for breast cancer, meaning roughly a third of carriers never develop it.

Expressivity is how severely the phenotype appears among those who show it. Neurofibromatosis type 1 has near-complete penetrance and famously variable expressivity — one family member may have a few café-au-lait macules and another optic gliomas and skeletal dysplasia, with the same mutation.

Keep the pair straight: penetrance is whether; expressivity is how much.

Clinical Connection · Cystic Fibrosis — Autosomal Recessive, With the Arithmetic Done

CFTR on chromosome 7 encodes a cAMP-regulated chloride channel in the apical membrane of secretory epithelia. Loss of function means chloride — and osmotically, water — is not secreted onto the epithelial surface, so airway surface liquid is depleted, mucus becomes thick and poorly cleared, and ducts obstruct. Note that in sweat glands the channel works in the opposite direction, reabsorbing chloride from the duct, which is why CF patients lose excess salt in sweat and why the sweat chloride test (>60 mmol/L) remains the diagnostic standard. One gene, two directions of transport, depending on which epithelium.

The consequences follow the ducts: chronic pulmonary infection and bronchiectasis, pancreatic exocrine insufficiency with malabsorption (and eventually endocrine failure — CF-related diabetes), meconium ileus in newborns, biliary cirrhosis, and congenital bilateral absence of the vas deferens causing obstructive azoospermia in about 98% of men with CF.

Over 2,000 CFTR variants are known. ΔF508, present on about 70% of CF alleles in Northern European populations, deletes three base pairs and therefore one phenylalanine at position 508: an in-frame deletion producing a full-length protein that misfolds and is degraded before it reaches the membrane. That mechanistic detail is now therapeutic — "corrector" drugs help the protein fold and traffic, and "potentiator" drugs open channels that do reach the surface, and the combination has changed the disease's trajectory. Median predicted survival has risen from under 10 years in the 1960s to beyond 50 today.

For the carrier arithmetic — including why an unaffected sibling of an affected person is 2/3 rather than 1/2 likely to be a carrier, and how a negative screen changes the number without zeroing it — work through the calculation in Figure 29.4.

Clinical Connection · Huntington Disease and Anticipation

HTT on chromosome 4 contains a CAG trinucleotide repeat near its 5′ end, translated into a polyglutamine tract. Repeat number determines everything:

CAG repeats Outcome
≤ 26 Normal, stable
27–35 Normal phenotype, but the allele is unstable and may expand in transmission
36–39 Reduced penetrance — may or may not develop disease
≥ 40 Full penetrance — disease is certain given normal lifespan
≥ 60 Juvenile onset

Expanded huntingtin misfolds, aggregates, and is toxic, killing medium spiny neurons of the striatum preferentially. The clinical triad is chorea, progressive cognitive decline, and psychiatric disturbance, with onset typically between 35 and 50 and death 15–20 years later.

Anticipation — earlier onset and greater severity in successive generations — is the striking feature, and it has a molecular cause. The repeat is unstable during DNA replication, particularly during spermatogenesis, where slippage during the many mitotic divisions of spermatogonial stem cells expands it further. A father with 42 repeats may pass on 50; a mother with 42 usually passes on 42 or 43. Juvenile-onset Huntington disease is therefore almost always paternally inherited, and this is one of the few places where the sex of the transmitting parent predicts the phenotype without imprinting being involved.

Huntington is also the ethical case study of predictive testing. It is autosomal dominant, fully penetrant above 40 repeats, testable decades before onset, and — as yet — untreatable. Roughly only 5–20% of at-risk individuals choose to be tested, which is a datum worth sitting with: the majority of people offered certain knowledge about their own future decline decline it. Formal protocols require pre-test counseling, a waiting period, and support, and testing of asymptomatic minors is not offered.

Clinical Connection · Hemophilia — Reading an X-Linked Pedigree

Hemophilia A (factor VIII deficiency, 1 in 5,000 male births) and hemophilia B (factor IX, 1 in 30,000) are X-linked recessive, and their pedigrees demonstrate every rule in Figure 29.5.

A male has only one X. If it carries the variant he has no second copy to compensate, so he is hemizygous and affected. A female with one variant copy is a carrier — usually with factor levels around 50% of normal, enough for hemostasis, though skewed X-inactivation can push a carrier's level low enough to cause bleeding, which is why the old assumption that carriers are always asymptomatic has been abandoned.

Three predictions follow, and each is diagnostic:

  • An affected father transmits to none of his sons (they receive his Y) and all of his daughters become carriers (they receive his only X). Male-to-male transmission of hemophilia does not happen.
  • A carrier mother transmits to half her sons (affected) and half her daughters (carriers).
  • An affected female requires an affected father and a carrier mother — rare but possible, and increasingly seen as treated men reach reproductive age.

The bleeding phenotype tracks residual factor activity: below 1% is severe, with spontaneous haemarthroses; 1–5% moderate; 5–40% mild, presenting only after surgery or trauma. Note what fails — the intrinsic pathway of secondary hemostasis — which is why the platelet count and bleeding time are normal, the PTT is prolonged, and the PT is normal. Chapter 17 develops the cascade. The historically famous pedigree is Queen Victoria's, in which a hemophilia B variant — almost certainly a new mutation in Victoria herself, since there were no affected ancestors — spread through the Spanish, German, and Russian royal families across three generations exactly as the rules predict.


29.7 Beyond Mendel

Most human variation is not Mendelian. Here are the ways it departs, in rough order of how far they take you from Mendel.

Incomplete dominance. The heterozygote is intermediate. Familial hypercholesterolemia is the clinical example: a heterozygote has roughly half the normal number of LDL receptors and an LDL of 200–400 mg/dL; a homozygote has almost none and an LDL of 600–1,000 mg/dL, with coronary disease in childhood. Dose of functional protein maps onto phenotype almost linearly.

Codominance. Both alleles are fully expressed. The ABO system is the standard example: an I^A I^B individual makes both the A and the B transferase and displays both antigens — blood type AB. Neither masks the other (Chapter 17).

Multiple alleles. More than two alleles exist in the population, though any individual has only two. ABO has three common alleles (I^A, I^B, i). The HLA loci have thousands, which is why tissue matching is difficult and why HLA typing is nearly individual-specific.

Pleiotropy. One gene, many unrelated effects. Marfan syndrome arises from FBN1, encoding fibrillin-1, a microfibril protein — so the phenotype appears wherever microfibrils matter: aortic root dilation and dissection, lens dislocation, tall stature with long limbs and fingers, pectus deformity, scoliosis, and dural ectasia. Nothing about that list is arbitrary once you know where the protein is. Sickle cell disease is equally pleiotropic for the same kind of reason.

Epistasis. One gene masks another. The Bombay phenotype is the classic: the H antigen is the substrate onto which A and B transferases build their antigens, and a person homozygous for a non-functional FUT1 gene (genotype hh) makes no H substrate. Such a person may carry I^A I^B and still type as O, because there is nothing for the transferases to act on. They can receive blood only from another Bombay individual — and would be catastrophically mis-typed by a routine ABO test.

Polygenic (multifactorial) inheritance. Many loci of small effect, plus environment, producing a continuous distribution. Height, blood pressure, LDL cholesterol, BMI, bone density, and susceptibility to nearly every common disease. This is §29.8, and it is the answer to Amara's case.

Sex-linked vs sex-influenced vs sex-limited. Distinguish these carefully:

  • Sex-linked — the gene is on a sex chromosome (hemophilia).
  • Sex-influenced — the gene is autosomal, but the phenotype differs by sex because the hormonal environment differs. Male-pattern baldness is the standard example: expressed as dominant in men and recessive in women, because expression requires androgen.
  • Sex-limited — the gene is autosomal but the phenotype can only appear in one sex (an autosomal gene affecting uterine anatomy, or male precocious puberty).

Mitochondrial inheritance. The mitochondrial genome is a 16,569-base-pair circle encoding 37 genes: 13 proteins (all subunits of the oxidative phosphorylation complexes), 22 tRNAs, and 2 rRNAs. Three properties follow, and they produce a pedigree unlike any nuclear pattern:

  1. Strictly maternal transmission. Paternal mitochondria entering at fertilization are ubiquitinated and destroyed (§28.1). An affected mother transmits to all her children; an affected father transmits to none. That single asymmetry identifies the pattern instantly.
  2. Heteroplasmy. A cell contains hundreds to thousands of mitochondrial genomes, and mutant and normal genomes coexist in varying proportions. There is no "heterozygote."
  3. Threshold effect. Symptoms appear only when the mutant fraction exceeds a tissue-specific threshold, typically 60–90%. Because tissues with the highest ATP demand — brain, retina, cardiac and skeletal muscle, cochlea, pancreatic beta cell — cross the threshold first, the clinical syndromes cluster there: LHON (sudden bilateral optic neuropathy in young adults), MELAS (encephalopathy with stroke-like episodes and lactic acidosis), MERRF, and maternally inherited diabetes with deafness.

Heteroplasmy also explains why severity varies so widely within one family, and why prenatal prediction is unreliable: the fraction of mutant genomes a given oocyte receives passes through a random bottleneck during oogenesis.

Genomic imprinting. For roughly 100 human genes, expression depends on the parent of origin — one copy is silenced by methylation laid down in the germ line, so the individual is functionally hemizygous. This breaks the most basic Mendelian assumption, which is that it does not matter which parent an allele came from. See the Development sidebar below.

Development · Imprinting, X-Inactivation, and the Calico Cat

Two developmental phenomena silence one copy of a gene while leaving the sequence untouched. Both are epigenetic, both are established early in development, and both produce inheritance patterns that look impossible on Mendelian assumptions.

Genomic imprinting. During gametogenesis, certain loci are marked by DNA methylation according to which sex is producing the gamete. The mark is erased and re-established each generation. At an imprinted locus, only one parental copy is expressed, so a deletion or a mutation on the active copy produces disease while the same lesion on the silent copy produces nothing.

The paradigm is the 15q11-13 region:

  • Prader-Willi syndrome results from loss of the paternal contribution — deletion (~70%), maternal uniparental disomy (~25%, in which both copies came from the mother), or an imprinting defect. The genes there are maternally imprinted (silenced), so losing the paternal copy leaves no expression. Phenotype: neonatal hypotonia and poor feeding, then hyperphagia and obesity, short stature, hypogonadism, and intellectual disability.
  • Angelman syndrome results from loss of the maternal contribution at the UBE3A gene in the same interval — which is paternally imprinted in neurons. Phenotype: severe intellectual disability, absent speech, ataxia, seizures, and a characteristically happy demeanour.

Same chromosomal region, opposite parent of origin, two entirely different diseases. Nothing in Mendel predicts this. Note also its relevance to Chapter 28: imprinted genes are heavily concentrated in genes controlling placental and fetal growthIGF2 is paternally expressed and growth-promoting, H19 maternally expressed and growth-limiting — which is the empirical basis of the "parental conflict" hypothesis and the reason imprinting disorders cluster around growth abnormalities and are over-represented after assisted reproduction.

X-inactivation (Lyonization). Females have two X chromosomes; males have one. Without compensation, females would produce double the dose of ~800 genes. Around days 10–14 of development, each cell randomly and permanently silences one X, condensing it into a Barr body visible at the nuclear periphery. The choice is random per cell but is then inherited by all that cell's descendants.

The consequence is that every female is a mosaic — roughly half her cells expressing her mother's X and half her father's, in patches whose size reflects how early inactivation occurred. The visible demonstration is the calico cat: the coat-colour gene for orange versus black is X-linked, so a heterozygous female is patched orange and black, with white from a separate autosomal gene. A calico cat is essentially always female; a male calico is almost always XXY.

The medical consequences are substantial. Carriers of X-linked recessive disease are mosaics, not uniformly unaffected — a hemophilia carrier with skewed inactivation can bleed, and a carrier of Duchenne muscular dystrophy can have a raised creatine kinase and mild weakness. Females affected by X-linked dominant conditions are typically milder than males because only half their cells express the variant. And about 15% of X-linked genes escape inactivation, which is why 45,X is not phenotypically neutral.

Check Your Understanding 29.7

  1. A pedigree shows an affected father, all four of his daughters affected, and none of his three sons affected. Name the inheritance pattern and justify it in one sentence.
  2. A woman has MELAS. She has four children; all four have some degree of symptoms, but the severity ranges from mild hearing loss to severe encephalopathy. Explain both observations.
Show answers
  1. X-linked dominant. A father gives his single X to every daughter and his Y to every son, so if the responsible allele is a dominant one on the X, transmission to all daughters and no sons is obligatory. Autosomal dominant would affect about half of each sex; X-linked recessive would leave the daughters as unaffected carriers; mitochondrial inheritance would affect none of a father's children at all.
  2. All four are affected because mitochondrial DNA is transmitted exclusively through the oocyte's cytoplasm, so every child of an affected mother inherits her mutant genomes. The severity varies because of heteroplasmy and the threshold effect: each oocyte receives a different proportion of mutant to normal mitochondrial genomes through a random bottleneck in oogenesis, and a tissue shows dysfunction only once the mutant fraction exceeds its own threshold — lowest in the most ATP-dependent tissues. A child receiving a 40% mutant load may cross the threshold only in the cochlea, while one receiving 85% crosses it in brain and muscle as well.

29.8 Polygenic and Multifactorial Disease

This is the section Amara's case exists for.

Her panel found no pathogenic variant. That result is not a failure of the test, and it is not reassurance. It is a statement that her disease is not built the way Mendelian disease is built.

What a polygenic risk score is

A genome-wide association study genotypes hundreds of thousands to millions of common variants — mostly single nucleotide polymorphisms, most of them in regulatory rather than coding DNA — in very large numbers of cases and controls, and asks which variants are more common in cases. For coronary artery disease, such studies have now identified several hundred genome-wide significant loci, and modelling suggests the true number of contributing variants runs into the millions, most with effects far too small to detect individually.

The effect sizes are the point. A typical CAD-associated variant changes risk by 1–5% relative — an odds ratio of 1.02 to 1.05. Individually, meaningless. Collectively, not.

A polygenic risk score is the weighted sum:

PRS = Σ (number of risk alleles at each locus × the effect size estimated for that locus)

computed across up to several million positions, and then expressed as a percentile against a reference population. Amara's score is at the 92nd percentile, meaning 92% of people in that reference population have a lower burden of CAD-associated variants than she does.

Concretely, in large cohorts, individuals in the top 5% of a CAD polygenic score carry roughly a 3- to 4-fold increased risk of coronary events compared with the middle of the distribution — comparable in magnitude to carrying a pathogenic LDLR variant. And because the top 5% of a distribution contains far more people than the ~1 in 250 who carry familial hypercholesterolemia, polygenic risk accounts for many times more coronary disease in the population than all monogenic causes combined.

   TWO ARCHITECTURES OF INHERITED RISK

   ══ MONOGENIC (e.g. FAMILIAL HYPERCHOLESTEROLEMIA) ══════════════════
   One gene. Large effect. Bimodal — you carry it or you do not.

     number of                    ████
     people      ██████████████   ████
                ████████████████  ████
               ██████████████████ ████
        ───────────────────────────────────────────► untreated LDL
                  ~100 mg/dL       ~250 mg/dL
                  NON-CARRIERS     CARRIERS (~1 in 250 people)
                                   3–13× lifetime CHD risk
                                   Segregates ~50:50 in a family
                                   ONE test → a categorical answer

   ══ POLYGENIC (e.g. AMARA'S CAD RISK) ═══════════════════════════════
   Millions of variants, each worth 1–5%. Continuous. Everyone has a score.

     number of                    ╭──╮
     people                    ╭──╯  ╰──╮
                            ╭──╯        ╰──╮
                         ╭──╯              ╰──╮
                     ╭───╯                    ╰───╮
                 ╭───╯                            ╰───╮        ▼ AMARA
             ╭───╯                                    ╰───╮   92nd %ile
        ─────┴──────────────────────────────────────────┴─┴──────────►
             1st        25th       50th       75th    92nd  99th
                                                       ▲
             ← lower risk ────────────────── higher risk →

     top 5%  ≈ 3–4× risk of coronary events vs the middle of the
             distribution — comparable in magnitude to FH, but far
             more COMMON, so it explains far more disease overall

   ══ THE FOUR DIFFERENCES THAT MATTER CLINICALLY ═════════════════════
                        MONOGENIC              POLYGENIC
   Inheritance          clean Mendelian        clusters, no clean pattern
   Family pedigree      segregates 50:50       "runs in the family"
   Effect of one variant huge                  1–5%
   Test result          categorical            a percentile
   Modifiability        risk is high whatever  large ABSOLUTE risk
                        you do (but treatment  reduction from lifestyle
                        works very well)       and from LDL lowering

Figure 29.6 — Monogenic and polygenic risk architectures contrasted.

Described: Two distributions are contrasted. The monogenic architecture, illustrated by familial hypercholesterolemia, is bimodal: a large peak of non-carriers with untreated LDL cholesterol around one hundred milligrams per decilitre and a small separate peak of carriers around two hundred fifty, affecting about one person in two hundred fifty, carrying a three- to thirteen-fold lifetime coronary risk, segregating fifty-fifty within a family, and yielding a categorical answer from a single test. The polygenic architecture, illustrated by Amara's coronary risk, is a continuous bell curve in which everyone has a score built from millions of variants each worth only one to five percent; Amara sits at the ninety-second percentile. People in the top five percent of the distribution carry roughly three- to four-fold the coronary event risk of those in the middle — comparable in magnitude to familial hypercholesterolemia but far more common, so polygenic risk explains far more disease overall. A summary table lists four clinically important differences: monogenic conditions show clean Mendelian inheritance while polygenic risk merely clusters in families with no clean pattern; monogenic pedigrees segregate fifty-fifty while polygenic ones only appear to run in the family; a single monogenic variant has a huge effect while a single polygenic variant has an effect of one to five percent; a monogenic test result is categorical while a polygenic result is a percentile; and monogenic risk remains high regardless of behaviour though it responds well to treatment, whereas polygenic risk admits a large absolute reduction from lifestyle and from LDL lowering.

What a polygenic risk score is not

Four things, and each has caused real harm when forgotten.

1 · It is not diagnostic. It reports a position in a distribution, not a disease. Amara's 92nd percentile does not say she has coronary disease; her angiogram says that. A 20th percentile does not say a person is safe — most coronary events occur in people with average scores, because that is where nearly everyone is.

2 · It is not deterministic. It is a probability shift. Plenty of people at the 95th percentile never have an event, and plenty at the 10th do. Genetic risk changes the odds; it does not fix the outcome.

3 · It is ancestry-dependent, and this is a serious limitation. The large majority of genome-wide association data has been collected in cohorts of European ancestry. Scores derived from those cohorts lose accuracy when applied to people of other ancestries — typically two- to fivefold reduced predictive power in individuals of African ancestry — because linkage disequilibrium patterns (which variants travel together) differ between populations, so the tagged variant may no longer sit near the causal one, and because allele frequencies and effect sizes differ. The Osei family is of West African descent. Amara's report says "European-ancestry reference panel" and carries a limitations note for exactly this reason. Her 92nd percentile is a real signal, but it is measured with a ruler calibrated on a different population, and honest counseling says so.

4 · It is not causal. An associated variant is a signpost, not a mechanism. Most sit in regulatory regions and their target gene is often unknown. A polygenic score is a prediction instrument, not an explanation.

Heritability, explained correctly

This is the most misused number in genetics, and using it correctly is genuinely worth the effort.

Heritability (h²) is the proportion of the variance of a trait in a particular population, in a particular environment, that is attributable to genetic variance.

Read that definition three times. Then note what it does not say.

  • It says nothing about an individual. "Height is 80% heritable" does not mean 80% of your height comes from your genes and 20% from your diet. That statement is not even wrong; it is meaningless. Heritability is a statement about differences between people, not about the composition of any person.
  • It is not fixed. It depends on how much environmental variation the population contains. If you gave everyone in a population an identical environment, heritability would rise toward 1.0 — not because genes became more important, but because the environmental variance you divided by disappeared. Conversely, in a population with wildly unequal nutrition, heritability of height falls. Heritability partly measures how equal an environment is.
  • It says nothing whatever about modifiability. This is the error that matters clinically.

Two examples settle the last point permanently.

Height is roughly 80% heritable, and average adult height in the Netherlands rose about 20 cm over 150 years. Genes explain most of the differences between Dutch people at any one time; nutrition explains the change in the mean across time. High heritability and a large environmental effect coexisted without contradiction.

Phenylketonuria is essentially 100% heritable — the phenotype is determined entirely by genotype in ordinary environments — and it is completely preventable by a low-phenylalanine diet started in the first weeks of life. Maximal heritability, maximal modifiability. Anyone who tells you a highly heritable condition cannot be changed has misunderstood the statistic.

Gene–environment interaction, quantified

The useful question is never "genes or environment?" It is "how much does the environment move the outcome at a given level of genetic risk?"

For coronary disease that question has been answered directly. In a study of more than 55,000 people across four cohorts, participants were stratified by polygenic risk score and by a composite lifestyle score (non-smoking, no obesity, regular physical activity, healthy diet):

Unfavourable lifestyle Favourable lifestyle
High genetic risk (top quintile) 10.7% 10-year event rate 5.1%
Low genetic risk (bottom quintile) 5.8% 3.1%

Two readings, both important. First, within the high-genetic-risk group, favourable lifestyle roughly halved the event rate — an absolute reduction of 5.6 percentage points over ten years. Second, notice that a person at high genetic risk with a favourable lifestyle (5.1%) does better than a person at low genetic risk with an unfavourable one (5.8%). Genetic risk and behavioural risk are on comparable scales, and they are largely independent — meaning they add.

The same logic appears in lipid biology, and it is the key to Question 2. Mendelian randomization studies compare people who inherited naturally lower LDL through common variants with people who did not. Lifelong exposure to an LDL that is 1 mmol/L (~39 mg/dL) lower is associated with roughly an 80–88% lower risk of coronary disease. Starting a statin at 60 and lowering LDL by the same 1 mmol/L reduces risk by roughly 20–25%.

Same molecule, same magnitude of reduction, four-fold difference in benefit. The difference is duration. Atherosclerosis is driven by the cumulative exposure of the arterial wall to apoB-containing lipoproteins — the area under the LDL-versus-time curve — and a variant that lowers LDL from conception has been working for six decades before the drug is even prescribed.

Reading a pedigree for a common complex disease

Mendelian pedigrees are read by pattern. Polygenic pedigrees have no pattern, and must be read by load instead. Four features raise the estimate:

  1. Number of affected first-degree relatives. Each roughly doubles risk.
  2. Age at onset. Early onset implies a heavier genetic contribution. The standard clinical definition of a positive family history for premature cardiovascular disease is a first-degree relative affected before 55 in men, 65 in women.
  3. Severity, and whether the relative had the usual risk factors. An MI at 51 in a non-smoking, normal-weight relative is more informative than one at 70 in a lifelong smoker.
  4. Both sides of the family, and clustering of related traits — hypertension, diabetes, dyslipidemia, and stroke are correlated phenotypes drawing on overlapping variants.
   THE OSEI FAMILY PEDIGREE
   □ unaffected male    ○ unaffected female    ■ / ● affected
   ⧄ deceased (diagonal)   ↗ proband arrow    numbers = age or age at death

   I        ⊠────────────────○  ADWOA MENSAI, 78
            KWAME OSEI          HYPERTENSION (from ~50s)
            d. 58, MI           osteoporosis · no coronary events
            smoker              ┌── LDL never measured before 60
            │                   │
            └─────────┬─────────┘
                      │
   II   ┌─────────────┴──────────────┐
        ■  KOFI OSEI, 53             ● AMARA OSEI, 45   ↗ PROBAND
        MI at 51                       CAD · hypertension · type 2
        hypertension                   diabetes · CKD stage 3
        │                              LDL 168 · HDL 38 · TG 210
        │                              LDLR/APOB/PCSK9: NO pathogenic
        │                              variant.  PRS(CAD) = 92nd %ile
        │                              │
        │                        ┌─────┴─────┐
   III  ○                        ○           □
        (Kofi's daughter,        NIA, 24     brother, 19
         28, unaffected)         pregnant 28 wk · normotensive
                                 competitive runner · no metabolic
                                 abnormality · PRS not measured

   ═══ HOW TO READ THIS ══════════════════════════════════════════════════
   NO — it does NOT show a Mendelian pattern.
       – not autosomal dominant: Adwoa (78) has hypertension but NO
         coronary disease, and both of her children have coronary events;
         a dominant coronary allele should be traceable through one
         affected parent
       – not recessive: two generations affected, no consanguinity
       – not X-linked: father-to-daughter AND father-to-son involvement
         patterns are both present in generation II
   YES — it DOES show high FAMILIAL LOAD:
       – 3 of 4 individuals in generations I–II have cardiovascular
         disease
       – TWO premature events: Kwame at 58, Kofi at 51 (male < 55 = the
         clinical threshold for "premature")
       – correlated traits clustering: hypertension in 3, diabetes in 1,
         dyslipidaemia in 1
   ⇒ CONCLUSION: polygenic burden plus shared environment, not a single
       broken gene. The genetic test result and the pedigree AGREE.

Figure 29.7 — The Osei family pedigree, drawn with standard symbols and annotated.

Described: A three-generation pedigree of the Osei family. Generation one shows Kwame Osei, who died at fifty-eight of a myocardial infarction and was a smoker, married to Adwoa Mensah, now seventy-eight, who has had hypertension since her fifties and has osteoporosis but no coronary events. Generation two shows their two children: Kofi Osei, fifty-three, who had a myocardial infarction at fifty-one and has hypertension; and Amara Osei, forty-five, marked as the proband, with coronary artery disease, hypertension, type 2 diabetes, and stage 3 chronic kidney disease, an LDL of 168, HDL of 38, triglycerides of 210, no pathogenic variant in LDLR, APOB, or PCSK9, and a coronary artery disease polygenic risk score at the ninety-second percentile. Generation three shows Kofi's unaffected twenty-eight-year-old daughter, and Amara's two children: Nia, twenty-four, pregnant at twenty-eight weeks, normotensive, a competitive runner with no metabolic abnormality and no polygenic score measured; and a healthy nineteen-year-old son. An interpretation panel explains that the pedigree does not show a Mendelian pattern: it is not autosomal dominant, because Adwoa at seventy-eight has hypertension but no coronary disease while both of her children have had coronary events; it is not recessive, since two generations are affected without consanguinity; and it is not X-linked, because both father-to-daughter and father-to-son involvement appear in generation two. It does show high familial load: three of the four individuals in generations one and two have cardiovascular disease, two events were premature — Kwame at fifty-eight and Kofi at fifty-one, below the male threshold of fifty-five — and correlated traits cluster, with hypertension in three members, diabetes in one, and dyslipidaemia in one. The conclusion is that the family carries a polygenic burden together with a shared environment rather than a single broken gene, and that the genetic test result and the pedigree agree with each other.

Clinical Connection · Familial Hypercholesterolemia — the Contrast Case

Amara's panel was sent to look for familial hypercholesterolemia, and understanding why it was a reasonable test — and what its negative result means — requires understanding what FH actually is.

FH is autosomal dominant, prevalence about 1 in 250, caused most often by loss-of-function variants in LDLR (the LDL receptor), less often in APOB (the ligand the receptor binds), and occasionally by gain-of-function variants in PCSK9 (a protease that degrades the receptor — more PCSK9 activity means fewer receptors). All three converge on one mechanism: fewer functional LDL receptors on hepatocytes, so LDL is cleared from plasma more slowly, so plasma LDL rises and stays high from birth.

Heterozygous FH Homozygous FH
Frequency 1 in 250 1 in 250,000–1,000,000
Receptor function ~50% ~0–15%
Untreated LDL 200–400 mg/dL 600–1,000 mg/dL
Untreated coronary event 4th–5th decade in men 1st–2nd decade
Physical signs Tendon xanthomas (Achilles, extensor tendons), corneal arcus before 45, xanthelasma Cutaneous xanthomas in infancy

Note that this is incomplete dominance at the biochemical level — receptor number and LDL concentration track gene dosage almost linearly — even though it is called dominant clinically.

Why Amara's negative result does not mean "no genetic risk." FH is one architecture of inherited high LDL: a single large-effect variant. Her LDL of 168 mg/dL is elevated but sits well below the usual FH range, she has no tendon xanthomas, and her lipid pattern — moderately high LDL, low HDL at 38, triglycerides at 210 — is the pattern of metabolic syndrome and insulin resistance, not of an isolated receptor defect. Her genetic contribution is spread across millions of small-effect variants affecting lipoprotein metabolism, blood pressure, endothelial function, coagulation, and inflammation.

The practical difference at the bedside is smaller than students expect. Both architectures respond to LDL lowering; both benefit from starting early; and in both, the earlier the exposure is reduced, the greater the benefit. What changes is who else needs testing: FH demands cascade screening of every first-degree relative, because each has a 50% chance of carrying the same variant and of needing treatment from childhood. Polygenic risk does not segregate that cleanly — but it still clusters, which is why the pedigree, not the panel, is what identifies Nia and her brother as people worth watching.

Exercise & Sport · The Genetics of Athletic Performance, and Where the Claims Outrun the Data

Nia runs a 3:04 marathon. Her mother does not run. It is natural to ask how much of that is genetic, and the honest answer is a useful case study in reading genetic claims critically.

What is well established. Twin and family studies put the heritability of maximal oxygen uptake (V̇O₂max) at roughly 50%, and the heritability of the response to training — the change in V̇O₂max after a standardized program — at around 47%. The HERITAGE Family Study trained 481 sedentary people identically for 20 weeks and found V̇O₂max gains ranging from essentially zero to over 1,000 mL/min, with responses clustering strongly within families. Trainability itself is heritable, and non-response is real. That is a genuinely important finding, because it means a program that works for one athlete may do very little for another, and neither is a matter of effort.

The two famous polymorphisms, and what they actually show.

ACTN3 encodes α-actinin-3, a structural protein of the Z-disc found only in fast-twitch (type II) fibres. The common R577X variant introduces a premature stop codon; about 18% of people worldwide are XX and produce none of the protein at all. The findings are consistent in direction and modest in size: the XX genotype is markedly under-represented among elite sprint and power athletes (in some cohorts of elite sprinters, essentially absent), slightly over-represented among elite endurance athletes, and associated with small differences in muscle fibre composition and sprint performance in the general population. The mechanism is plausible — α-actinin-2 substitutes for the missing protein, and the substitution appears to shift fibres toward a more oxidative profile. But the variant explains only about 1–3% of the variance in sprint performance, and XX individuals have won Olympic sprint medals.

ACE insertion/deletion polymorphism sits in a non-coding region and correlates with serum ACE activity. The I allele (lower activity) has been associated with endurance performance and the D allele with power. Here the literature is considerably weaker: findings replicate inconsistently across cohorts and sports, and meta-analyses show effects near zero for several outcomes.

Where commercial claims outrun the evidence. Direct-to-consumer "sports genetics" panels typically test a handful of variants — often ACTN3 and ACE — and return advice about whether a person is a "power" or "endurance" athlete, or which training program to follow. Three problems:

  1. Effect sizes are tiny. Even the best-supported variants explain a few percent of variance in a trait that is influenced by hundreds to thousands of loci plus training history, biomechanics, psychology, and opportunity.
  2. Prediction is not the same as association. A variant can be reliably over-represented in elite athletes and still be almost useless for predicting whether this child will become one, because the base rate of becoming an elite athlete is tiny.
  3. No trial has shown that genotype-guided training beats well-designed conventional training. Position statements from sports science bodies are unambiguous that genetic testing for talent identification in children is not supported.

The defensible summary is the one that respects both halves of the evidence: genetics constrains the ceiling and shapes the response to training, and no current test can tell an individual where their ceiling is. Measured trainability — actually training someone and recording the response — remains far more informative than any panel.

Thread 3 · The Body Is Integrated

A polygenic risk score is integration made numerical.

The several hundred known coronary artery disease loci do not sit in a "heart disease" pathway, because there is no such pathway. They sit in LDL receptor regulation and lipoprotein metabolism (Chapter 24), in blood pressure control through the renin–angiotensin system and renal sodium handling (Chapters 16 and 26), in endothelial nitric oxide signalling (Chapter 19), in coagulation (Chapter 17), in vascular smooth muscle and extracellular matrix biology (Chapter 4), and in inflammation (Chapter 21). Some act on the vessel wall directly; others act only through the kidney; others only through the liver.

They converge because the arterial wall integrates them. Every one of those systems delivers its output to the same few centimetres of coronary endothelium, over decades, and the plaque is the running total.

That is why a single number can summarize hundreds of unrelated mechanisms, and also why the number is not an explanation. It is an integral, and the systems being integrated are the previous twenty-four chapters of this book.


29.9 Epigenetics

Nia cannot change her genome. She can change which parts of it are being read, and in some tissues she already has.

Epigenetics refers to heritable — through cell division, and sometimes through generations — changes in gene expression that do not involve any change in DNA sequence. The distinction is worth stating carefully because it is routinely overstated in popular writing: epigenetic marks are real, mechanistically well characterized, and mitotically heritable, and the evidence that lifestyle durably rewrites them in ways that change disease outcomes is much thinner than the enthusiasm suggests.

The two principal mechanisms

   EPIGENETIC MODIFICATION OF CHROMATIN

   ══ SILENT / CLOSED ══════════════   ══ ACTIVE / OPEN ═══════════════
   HETEROCHROMATIN                     EUCHROMATIN

   histone tails: DEACETYLATED         histone tails: ACETYLATED
   (HDACs removed the acetyl groups)   (HATs added acetyl groups)
        │                                   │
        │ tails keep their POSITIVE         │ acetyl NEUTRALISES the
        │ charge → grip the NEGATIVE        │ tail's positive charge →
        │ DNA backbone tightly              │ grip on DNA loosens
        ▼                                   ▼
     ●━●━●━●━●━●━●   nucleosomes          ●───●────●───●   nucleosomes
     tightly packed, ~30 nm fibre         spaced out, DNA accessible
     ▲                                    ▲
     │ CpG islands METHYLATED             │ CpG islands UNMETHYLATED
     │   ─C─G─  →  ─C(CH₃)─G─             │
     │   recruits methyl-CpG binding      │ transcription factors and
     │   proteins → recruits HDACs        │ RNA pol II can bind the
     │   → a self-reinforcing loop        │ promoter
     ▼                                    ▼
     TRANSCRIPTION BLOCKED                TRANSCRIPTION PROCEEDS

   ─────────────────────────────────────────────────────────────────────
   HISTONE MARKS worth knowing
     H3K4me3    trimethylated lysine 4    →  ACTIVE promoter
     H3K27ac    acetylated lysine 27      →  ACTIVE enhancer
     H3K27me3   trimethylated lysine 27   →  REPRESSED (Polycomb)
     H3K9me3    trimethylated lysine 9    →  constitutive heterochromatin
   Methylation of histones can activate OR repress depending on which
   residue — unlike DNA methylation, which is essentially always
   repressive at a promoter.

   ─────────────────────────────────────────────────────────────────────
   WHAT WRITES AND ERASES
     DNA methyltransferases (DNMT1 maintains, DNMT3A/3B establish)
     TET enzymes (oxidise 5-mC toward demethylation)
     HATs / HDACs · histone methyltransferases / demethylases
     Methyl groups all come from S-ADENOSYLMETHIONINE — which is
     supplied by ONE-CARBON METABOLISM, i.e. by FOLATE and B₁₂.
     ← this is the molecular bridge to §28.4 and to the Dutch Hunger
       Winter cohort

   ─────────────────────────────────────────────────────────────────────
   WHEN THE PATTERN IS SET AND RESET
     fertilisation ──► genome-wide DEMETHYLATION ──► re-methylation
        (imprinted loci ESCAPE this erasure — §29.7)
     primordial germ cells ──► a SECOND erasure, then sex-specific
        re-establishment of imprints during gametogenesis
     ⇒ two reprogramming windows per generation, and they are why
       transgenerational inheritance of acquired marks is HARD, and
       why claims of it require strong evidence

Figure 29.8 — Epigenetic modification of chromatin: DNA methylation and histone modification.

Described: Two chromatin states are contrasted. On the left, silent or closed heterochromatin: histone tails have been deacetylated by histone deacetylases, so the tails keep their positive charge and grip the negatively charged DNA backbone tightly, nucleosomes are tightly packed into a thirty-nanometre fibre, and CpG islands are methylated — a cytosine followed by a guanine carries a methyl group, which recruits methyl-CpG binding proteins, which in turn recruit histone deacetylases in a self-reinforcing loop — so transcription is blocked. On the right, active or open euchromatin: histone tails have been acetylated by histone acetyltransferases, which neutralises the tails' positive charge and loosens their grip on DNA, nucleosomes are spaced apart, DNA is accessible, CpG islands are unmethylated, and transcription factors and RNA polymerase II can bind the promoter, so transcription proceeds. A list of histone marks follows: trimethylation of histone H3 lysine 4 marks active promoters, acetylation of lysine 27 marks active enhancers, trimethylation of lysine 27 marks Polycomb-repressed regions, and trimethylation of lysine 9 marks constitutive heterochromatin — histone methylation can therefore activate or repress depending on the residue, unlike DNA methylation, which is essentially always repressive at a promoter. The writers and erasers are listed: DNA methyltransferases, with DNMT1 maintaining existing patterns and DNMT3A and DNMT3B establishing new ones; TET enzymes, which oxidise 5-methylcytosine toward demethylation; histone acetyltransferases and deacetylases; and histone methyltransferases and demethylases. All methyl groups derive from S-adenosylmethionine, supplied by one-carbon metabolism and therefore by folate and vitamin B12, which is the molecular bridge to neural tube defects and to the Dutch Hunger Winter cohort. A final panel describes when patterns are set and reset: at fertilisation the genome is demethylated and then re-methylated, with imprinted loci escaping the erasure, and in primordial germ cells a second erasure is followed by sex-specific re-establishment of imprints during gametogenesis — two reprogramming windows per generation, which is why transgenerational inheritance of acquired marks is difficult and why claims of it require strong evidence.

DNA methylation adds a methyl group to the 5-carbon of cytosine, almost always where a cytosine is followed by a guanine (a CpG dinucleotide). The genome contains about 28 million CpG sites, and roughly 70–80% are methylated. Clusters of CpGs called CpG islands sit at about 60% of gene promoters, and these are usually unmethylated in active genes. Methylation of a promoter CpG island silences the gene — directly, by blocking transcription factor binding, and indirectly, by recruiting methyl-CpG-binding proteins that bring histone deacetylases with them.

Histone modification alters the accessibility of DNA by changing how tightly nucleosomes grip it. Acetylation of lysine residues on histone tails neutralizes their positive charge, loosening their electrostatic grip on the phosphate backbone and opening the chromatin. Deacetylation closes it. Methylation of histones can do either, depending on which residue is marked.

Real examples, stated at the right confidence

Exercise remodels the skeletal muscle methylome, acutely and durably. A single bout of intense exercise produces measurable hypomethylation of the promoters of PGC-1α, PDK4, PPARδ, and TFAM in human vastus lateralis within 20 minutes to 3 hours, followed by increased transcription of those genes. PGC-1α is the master regulator of mitochondrial biogenesis — so this is the molecular first step of aerobic adaptation, and it is an epigenetic step. Longer training programs alter methylation at thousands of sites across the muscle genome. Notably, some of the hypomethylation induced by resistance training persists through weeks of detraining, and muscle retrained afterward grows faster — a plausible molecular substrate for the long-observed phenomenon of "muscle memory."

This is Nia's most direct answer, and it is worth stating precisely: her training is not changing her DNA sequence, and it is not changing her polygenic score. It is changing which genes her muscle, liver, adipose tissue, and vascular endothelium are transcribing.

Prenatal nutrition leaves durable marks. The Dutch Hunger Winter cohort (§28.9) showed reduced methylation at the IGF2 differentially methylated region in individuals exposed to famine periconceptionally — still detectable in blood six decades later — and not in their unexposed same-sex siblings. This is the strongest human evidence linking a defined prenatal exposure to a durable epigenetic difference. The effect size is small (a few percentage points of methylation), and whether it is causal for the associated adult disease remains unproven.

Other well-replicated examples. Smoking produces robust hypomethylation at the AHRR locus, so reliable that it is used to verify smoking status; the mark reverses slowly over years after quitting. Epigenetic clocks — weighted combinations of methylation at a few hundred CpG sites — predict chronological age within a few years, and the residual (biological age minus chronological age) predicts mortality. Exercise, diet, and sleep are associated with modestly "younger" clock readings.

What not to claim

Three restraints, stated plainly because this is the part of genetics most prone to overreach.

  • Correlation is not causation here either. Most human epigenetic findings are associations measured in blood, which may not reflect the tissue that matters.
  • Effect sizes are small. Methylation differences attributable to lifestyle are typically a few percentage points at individual sites.
  • Transgenerational inheritance in humans is not established. Two genome-wide reprogramming events per generation — at fertilization and in primordial germ cells — erase nearly all marks. A grandmother's famine exposure affecting a grandchild can be explained without invoking transmitted epigenetic marks at all, because a pregnant woman contains her daughter's developing germ cells, so a "three-generation" exposure is really a direct one. Genuine transgenerational epigenetic inheritance is well demonstrated in plants and some animals; in humans it remains an open question.

The honest, useful claim: DNA sequence is fixed; gene expression is not; expression responds to behavior, nutrition, sleep, and environment through mechanisms that are chemically identifiable and in several cases measurable. That is enough to be worth acting on, and it does not require overstating.

Thread 2 · Homeostasis Is the Master Concept

The genome is not a blueprint that is executed once. It is the parts list and the control system, and gene expression is itself homeostatic.

Every negative feedback loop in this book eventually runs through transcription. Low iron stabilizes the mRNA for the transferrin receptor and represses ferritin translation, so the cell imports more iron and stores less. Hypoxia stabilizes HIF-1α, which drives transcription of erythropoietin, VEGF, and glycolytic enzymes — a receptor, a control center, and a set of effectors, all inside one cell. Cortisol, thyroid hormone, and the sex steroids are all transcription factors in disguise: their receptors bind DNA directly (Chapter 16).

Fit these into the four boxes from Chapter 1 and the picture is complete. The variable is the concentration of a gene product. The receptor is a molecule that binds a metabolite or a hormone. The control center is the promoter and its transcription factors. The effector is RNA polymerase. And the epigenetic marks of this section are the mechanism by which the set point of that loop can be adjusted for the long term — which is exactly what development does, and exactly what §28.9's thrifty phenotype proposes goes wrong.


29.10 Advanced Topic · Genetic Testing and Counseling

Choosing a test by resolution

Test Resolution Detects Misses Turnaround
Karyotype 5–10 Mb Aneuploidy, large deletions/duplications, balanced translocations, mosaicism Anything smaller than a band 1–2 weeks; needs dividing cells
FISH ~100 kb–1 Mb, targeted A specific, pre-specified locus; works on interphase nuclei Anything not asked about Hours to days
Chromosomal microarray 50–100 kb genome-wide Copy number variants anywhere; loss of heterozygosity Balanced rearrangements (no dosage change); point mutations 1–2 weeks
Sanger sequencing Single base, one gene Point mutations in a specified gene Everything else Days
Gene panel / exome Single base, many genes; exome ≈ 1–2% of the genome but ~85% of known pathogenic variants Point mutations, small indels Deep intronic and most regulatory variants; poor for repeat expansions and CNVs Weeks
Genome sequencing Single base, whole genome Nearly everything, including structural and non-coding Interpretation, mostly Weeks

The clinical skill is matching resolution to question. Suspected Down syndrome: karyotype (and it must be a karyotype rather than FISH, because it distinguishes free trisomy from a translocation and therefore changes recurrence risk). Suspected 22q11.2 deletion: FISH or microarray. Unexplained developmental delay: microarray first. Suspected familial hypercholesterolemia: a gene panel — which is what Amara had.

Results are reported on the ACMG five-tier scale: pathogenic, likely pathogenic, variant of uncertain significance (VUS), likely benign, benign. Amara's report contains a VUS, and the correct clinical handling of a VUS is to treat it as a non-result: it does not change management, it should not trigger cascade testing, and it may be reclassified in either direction as databases grow. VUS results are substantially more common in people of non-European ancestry for the same reason polygenic scores transfer poorly — the reference databases are less complete.

Screening across the lifespan

Carrier screening. Historically ancestry-based (Tay-Sachs in Ashkenazi Jewish populations, sickle cell in African-ancestry populations, thalassemia in Mediterranean and Southeast Asian populations). Increasingly expanded and pan-ethnic, testing hundreds of conditions at once — partly because ancestry is a poor proxy in admixed populations, and partly because sequencing costs collapsed. On a large panel, roughly one person in three or four is a carrier of something.

Prenatal screening. Cell-free DNA screening (non-invasive prenatal testing) analyses fragments of DNA in maternal plasma from about 10 weeks; 5–15% of that DNA is fetal in origin — strictly, placental, from apoptotic trophoblast, which is the source of its characteristic false positives from confined placental mosaicism. Sensitivity for trisomy 21 exceeds 99% and specificity is about 99.9%. But positive predictive value depends on prevalence, and this is where counseling most often fails: at age 40 a positive cfDNA result for trisomy 21 has a PPV around 90%, while at age 25 it is closer to 50%, and for rare microdeletions it can fall below 10%. A screening test with excellent sensitivity can still be wrong more often than right when the condition is rare. Every positive requires a diagnostic test: chorionic villus sampling at 10–13 weeks or amniocentesis from 15 weeks, each with a procedure-related loss risk of roughly 0.1–0.3%.

Newborn screening. A heel-prick blood spot in the first days of life, testing 30–60+ conditions depending on jurisdiction. The paradigm is phenylketonuria: an autosomal recessive deficiency of phenylalanine hydroxylase affecting about 1 in 10,000–15,000 births, which causes severe irreversible intellectual disability if untreated and essentially none if a phenylalanine-restricted diet begins in the first weeks. The criteria that justify population screening are all met — the condition is serious, detectable presymptomatically, and treatable, and the test is cheap and accurate. Also screened: congenital hypothyroidism, cystic fibrosis, sickle cell disease, severe combined immunodeficiency, and medium-chain acyl-CoA dehydrogenase deficiency.

Pharmacogenomics

Genotype predicts drug response, and two examples come directly from Amara's medication list.

  • CYP2C19 and clopidogrel. Clopidogrel is a prodrug requiring CYP2C19 for activation. Loss-of-function alleles (2, 3) produce reduced platelet inhibition and a measurably higher rate of stent thrombosis and recurrent events after percutaneous coronary intervention. Roughly 30% of people of European ancestry and up to 50% of East Asian ancestry carry at least one loss-of-function allele. Genotype-guided selection of an alternative agent is now recommended in several guidelines.
  • SLCO1B1 and statin myopathy. SLCO1B1 encodes the hepatic transporter that takes statins out of the circulation. Reduced-function variants raise systemic simvastatin exposure several- fold and increase the risk of myopathy substantially.

Others worth knowing: CYP2C9/VKORC1 and warfarin dosing; TPMT/NUDT15 and thiopurine toxicity; HLA-B*57:01 and abacavir hypersensitivity, where pre-testing has essentially eliminated a once-common and potentially fatal reaction; and HLA-B*15:02 and carbamazepine-induced Stevens–Johnson syndrome in individuals of Southeast Asian ancestry.

The ethical dimensions, briefly and factually

  • Predictive information about an untreatable condition. Huntington disease is the case study, and the fact that only 5–20% of at-risk individuals choose testing is data, not squeamishness. The right not to know is a recognized principle.
  • Testing minors. Standard practice is to defer predictive testing for adult-onset conditions until the individual can consent, unless childhood intervention would change outcomes.
  • Duty to warn relatives. A pathogenic BRCA1 or Lynch syndrome variant is medically relevant to siblings and children who have not consented to anything. Practice strongly favours encouraging the patient to inform relatives rather than clinician disclosure.
  • Secondary findings. Sequencing done for one reason finds things about others. Professional guidelines specify a list of actionable genes to be reported regardless of indication, with an opt-out.
  • Discrimination. In the United States, the Genetic Information Nondiscrimination Act (2008) prohibits genetic discrimination in health insurance and employment — and explicitly does not cover life, disability, or long-term care insurance. Patients considering testing should be told this.
  • Direct-to-consumer testing. Most consumer products are genotyping arrays, not sequencing: they check pre-specified positions. A consumer "BRCA test" may check three Ashkenazi founder variants out of more than a thousand known pathogenic ones, so a negative result is close to uninformative for most people. Raw-data reinterpretation by third-party tools produces high false-positive rates and should always be confirmed in a clinical laboratory.
  • Privacy and genetic relatives. Your genome is 50% shared with each parent, child, and sibling. Consenting to share it is partly consenting on their behalf — a point made vivid by forensic genetic genealogy, which identifies individuals through the database entries of third cousins who never consented to anything.

Imaging · Karyotype and FISH — the Photography of Genetics

Every other Imaging sidebar in this book has described a way of seeing structure inside a living body. Cytogenetics is the same enterprise at a scale a thousand times smaller, and the same principle governs it: the physical method determines what the picture can show.

The karyotype is a photograph. Metaphase chromosomes are stained with Giemsa, imaged, and digitally arranged in order. Its strength is that it is unbiased and whole-genome — you are looking at all 46 chromosomes at once and you may find something nobody suspected, including a balanced translocation, in which two chromosomes have exchanged segments with no net gain or loss. That is invisible to every dosage-based method, and it matters, because a balanced carrier is healthy but produces unbalanced gametes and recurrent miscarriage. Its weakness is resolution: 5–10 million base pairs, hundreds of genes per band.

FISH is a targeted spotlight. A fluorescently labelled DNA probe complementary to a chosen sequence is hybridized to the sample; the locus lights up, and you count the signals. Two signals is normal, one is a deletion, three is a duplication or trisomy. It resolves down to about 100 kilobases, works on non-dividing interphase nuclei, and returns an answer in hours — which is why it is used for rapid aneuploidy detection on amniotic fluid and for specific microdeletion syndromes such as 22q11.2. Its weakness is the mirror image of the karyotype's strength: it only shows you what you asked about. A FISH probe for chromosome 21 will not notice a translocation involving chromosome 5.

Chromosomal microarray is a whole-genome dosage map. Sample DNA is hybridized to hundreds of thousands of probes tiling the genome, and the signal intensity at each reports copy number. It finds deletions and duplications ten to a hundred times smaller than a karyotype can, anywhere in the genome, without needing dividing cells. And it is blind to balanced rearrangements, because nothing has changed in dosage.

Cell-free DNA screening images a pregnancy from a maternal blood tube. No probe touches the fetus at all; the fetal-placental fraction of plasma DNA is sequenced and counted, and an excess of chromosome-21 fragments is inferred statistically.

The parallel with radiology is exact. The karyotype is the plain radiograph — cheap, whole-field, low resolution, occasionally the only thing that shows the finding. FISH is the targeted ultrasound. Microarray is the CT. Sequencing is the MRI. And in genetics as in radiology, the commonest clinical error is ordering the highest-resolution study when a lower-resolution one answers the question, or a targeted study when you do not yet know where to point it.


Chapter Summary

§29.1 A genome is the full DNA sequence; a gene is a sequence specifying a product; an allele is one version of it at a locus; genotype is what you carry and phenotype is what shows. Humans have 22 autosome pairs plus XX or XY. Homologous chromosomes carry the same loci in the same order but not the same alleles — which is the whole basis of Mendelian genetics.

§29.2 DNA's structure is its copying mechanism: antiparallel strands, complementary base pairing (A–T with two hydrogen bonds, G–C with three), and 5′→3′ synthesis only. Replication is semiconservative, continuous on the leading strand and discontinuous in Okazaki fragments on the lagging strand. Fidelity is built in three multiplying layers — base-pairing, proofreading, and mismatch repair — taking the error rate to about 1 in 10⁹. Each repair system, when inherited broken, presents as a cancer predisposition syndrome.

§29.3 Transcription produces pre-mRNA, which is capped, polyadenylated, and spliced. Because about 95% of multi-exon genes are alternatively spliced, ~20,000 genes yield well over 100,000 proteins. The genetic code is triplet, degenerate, unambiguous, non-overlapping, and nearly universal. Translation runs at the ribosome — itself a ribozyme — through A, P, and E sites, and post-translational modification multiplies the products again.

§29.4 Mutations are silent, missense, nonsense, frameshift, in-frame indel, splice-site, copy number, repeat expansion, or regulatory. The divisible-by-three rule separates a lost amino acid from a destroyed protein. Sickle cell disease traces a single A→T through a Glu→Val substitution, a hydrophobic patch, polymerization of deoxyhemoglobin, cell rigidity, vaso-occlusion and hemolysis, and finally a population-level balanced polymorphism against malaria.

§29.5 Meiosis halves chromosome number and generates variation through crossing over (45–55 events per meiosis), independent assortment (2²³ combinations), and random fertilization. Nondisjunction produces aneuploidy: trisomies 21, 18, and 13, and the sex chromosome aneuploidies, which are milder because of X-inactivation and the Y's small gene content. The maternal age effect arises from cohesin degrading in oocytes arrested since fetal life.

§29.6 Four Mendelian patterns are distinguished by three questions: does it skip generations, is there male-to-male transmission, and are the sexes affected equally. Punnett squares give per-pregnancy probabilities; the unaffected sibling of an affected person is 2/3 rather than 1/2 likely to be a carrier; and a negative screen reduces risk without abolishing it. Penetrance is whether; expressivity is how much.

§29.7 Beyond Mendel: incomplete dominance, codominance, multiple alleles, pleiotropy, epistasis, polygenic inheritance, sex-influenced and sex-limited traits, mitochondrial inheritance (strictly maternal, heteroplasmic, with a threshold effect), and genomic imprinting, in which the parent of origin determines expression.

§29.8 A polygenic risk score sums millions of small-effect variants into a percentile. The top 5% for coronary disease carries 3–4× the risk of the middle — comparable to monogenic FH, but far more common. It is not diagnostic, not deterministic, ancestry-dependent, and not causal. Heritability is a population variance statistic that says nothing about an individual and nothing about modifiability — height and PKU prove both points. Gene–environment interaction is large and quantifiable: at high genetic risk, favourable lifestyle roughly halves ten-year coronary event rates.

§29.9 Epigenetic marks — DNA methylation at CpG islands and histone modification — control which genes are read without altering sequence. Exercise demethylates the promoters of mitochondrial biogenesis genes within hours; prenatal famine leaves detectable marks six decades later. Effect sizes are modest, most human data are associative, and two reprogramming windows per generation make transgenerational inheritance in humans unproven. Sequence is fixed; expression is not.

§29.10 Tests are chosen by resolution: karyotype for whole-genome low-resolution and balanced rearrangements, FISH for a targeted question, microarray for submicroscopic copy number, sequencing for point mutations. A variant of uncertain significance is a non-result. Screening tests are not diagnostic tests, and a highly sensitive screen can still be wrong more often than right when the condition is rare. Pharmacogenomics already changes prescribing — CYP2C19 and clopidogrel most relevantly for this family.

The Three Threads in Chapter 29

Structure → Function. The double helix is the clearest case in all of biology: the molecule's shape — two complementary antiparallel strands — is simultaneously how information is stored and how it is copied. Nothing else about DNA had to be discovered for replication to become obvious. The same principle runs down to a single residue: replacing one hydrophilic glutamate with one hydrophobic valine creates a sticky patch on the surface of hemoglobin, and that patch is the entire mechanism of sickle cell disease.

Homeostasis. Gene expression is a control system, not a script. Transcription factors are comparators, promoters are control centers, and RNA polymerase is the effector — and epigenetic marks are how the set points of those loops are adjusted for the long term.

Integration. A polygenic risk score is a single number summarizing several hundred loci acting through lipoprotein metabolism, blood pressure control, endothelial function, coagulation, and inflammation. They converge because the arterial wall integrates them over decades. The score is an integral of the previous twenty-four chapters.


Case File 29 · Resolution

Question 1 — If there is no single broken gene, what exactly is being inherited?

What Amara inherited is not a gene. It is a distribution position.

Coronary artery disease is polygenic. Several hundred loci reach genome-wide significance and the true number of contributing variants runs into the millions, each shifting risk by 1–5%. Each of Amara's parents passed her one allele at every one of those positions, and by chance she received a combination that sits at the 92nd percentile of the reference distribution. Her brother Kofi, drawing independently from the same two parents, evidently drew similarly. Nia received half of Amara's alleles and half of her father's, so her expected score is intermediate between her parents — but the variance around that expectation is wide, and she could plausibly sit anywhere from well below average to above her mother.

What those variants actually do is the part worth understanding, because it explains why the disease looks the way it does in this family. They are not variants in a "heart disease gene." Most are regulatory, and they are spread across systems: LDL receptor expression and lipoprotein handling, renal sodium transport and renin–angiotensin signalling, endothelial nitric oxide production, coagulation and platelet reactivity, vascular smooth muscle and matrix biology, and inflammatory signalling. Amara's lipid pattern — LDL 168, HDL 38, triglycerides 210 — is the signature of insulin resistance rather than of a receptor defect, which is exactly what the negative LDLR panel predicts and exactly what a polygenic architecture produces.

So the honest sentence for the family is: they inherited a slightly unfavourable version of about a dozen physiological systems, none of them broken, all of them tilted in the same direction. The pedigree agrees with the score. Three of the four adults in generations I and II have cardiovascular disease, two of the events were premature, and correlated traits cluster — but the disease does not segregate 50:50 from an affected parent, it does not skip generations in a recessive pattern, and Adwoa at 78 has hypertension without coronary disease. That is familial load, not Mendelian inheritance, and no test would have found a single gene because there is not one.

One caveat belongs in this answer. The score was computed against a European-ancestry reference panel, and the Osei family is of West African descent. Polygenic scores lose two- to fivefold predictive accuracy across that ancestry boundary. Her 92nd percentile is a real signal measured with a mis-calibrated instrument, and it should be weighted accordingly — alongside a family history that is entirely unambiguous.

Question 2 — Why do Amara and Kofi have disease at 45 and 51 while Nia at 24 has none? Is that genetics or is it time?

It is overwhelmingly time, and the reason is worth stating precisely, because it is the most useful single idea in preventive cardiology.

Nia shares roughly half of Amara's genome. Whatever polygenic burden Amara carries, Nia carries a substantial and randomly selected fraction of it, and she has carried it since conception. Her genotype today is exactly what it will be at 50. Nothing about her genome will change between now and her mother's current age.

What will change is cumulative exposure. Atherosclerosis is not an event; it is an integral. The arterial wall accumulates apoB-containing lipoprotein particles that are retained by the subendothelial matrix, oxidised, taken up by macrophages, and built into plaque, and the rate of that accumulation is proportional to the concentration of those particles multiplied by the time they are present. The quantity that predicts disease is the area under the LDL-versus-age curve.

The Mendelian randomization data make this quantitative. Inheriting common variants that lower LDL by 1 mmol/L (about 39 mg/dL) from conception is associated with an 80–88% lower lifetime risk of coronary disease. Taking a statin from age 60 and achieving the identical 1 mmol/L reduction lowers risk by about 20–25%. Same molecule, same magnitude, four-fold difference — and the only variable that differs is how many years the reduction was in force.

Apply that to this family. Amara has had approximately twenty-five additional years of an LDL of 168 mg/dL, of blood pressure that has been elevated for at least three, of insulin resistance progressing to diabetes, and of twenty years of rotating night shift with its attendant disturbance of circadian metabolic control. Kofi has had a similar accumulation. Nia has had a normal blood pressure, no metabolic abnormality, and a training history that has been actively protective. She is not a different genetic person from her mother. She is her mother at twenty-four, before the integral had accumulated.

Two qualifications keep this honest. Some of the age difference is genetic: Nia may simply have drawn a more favourable half of the alleles, and there is no way to know without measuring. And her risk is not zero now — subclinical atherosclerosis begins in the second and third decades, and the plaque she will have at 50 is already starting. The point is not that she is safe. It is that the difference between her and her mother is measured in exposure-years, and exposure-years are the one variable in this entire chapter that is still available to be changed.

Question 3 — Nia cannot change her genome. What can she actually change, and how much would it matter?

Four things, in descending order of how well the evidence supports them.

1 · The magnitude and duration of her lipoprotein exposure — the largest lever. LDL is the causal driver, and benefit scales with concentration multiplied by years. Nia should know her LDL number now rather than at 40, because the intervention with the greatest lifetime effect is the one started earliest. Diet quality, weight maintenance, and — if her LDL proves elevated despite that — pharmacological lowering begun in her thirties rather than her fifties, are all acting on the same integral.

2 · Blood pressure and glucose over decades. Both are polygenically influenced and both are strongly modifiable. Her mother's hypertension and diabetes were not present at 24 either. This is where sustained aerobic training, sodium moderation, sleep, and weight stability do their work — and note that Nia is currently in one of the few periods of life when blood pressure is under direct clinical surveillance, because she is pregnant.

3 · Gene expression, which is genuinely modifiable even though sequence is not. Her running is not metaphorically changing her biology. A single bout of intense exercise produces measurable hypomethylation of the PGC-1α, PDK4, and PPARδ promoters in skeletal muscle within hours, followed by increased transcription and mitochondrial biogenesis. Training remodels methylation at thousands of sites, and some of it persists through detraining. Her genotype specifies which proteins she can make; her behaviour substantially determines which ones she does make, in which tissues, at what level.

How much does it matter? This has been measured directly, in people stratified by polygenic score. Among those in the top quintile of genetic risk for coronary disease, a favourable lifestyle — not smoking, not obese, physically active, eating well — was associated with a ten-year coronary event rate of 5.1% versus 10.7% with an unfavourable one. Roughly a halving, and an absolute reduction of 5.6 percentage points. More striking still: a person at high genetic risk with a favourable lifestyle did better than a person at low genetic risk with an unfavourable one.

4 · Two things that are not lifestyle and matter anyway. Because Amara's disease is polygenic, cascade genetic testing of relatives is not indicated — but clinical cascade screening is: Nia and her brother should have a lipid panel and blood pressure documented now, because family history is itself a validated risk factor independent of any score. And Nia's pregnancy is currently an unrepeatable natural stress test of her cardiovascular and metabolic reserve (§28.9): if she were to develop preeclampsia or gestational diabetes, that would identify her decades early as someone whose vascular and beta-cell reserve is limited. She has developed neither.

What should not be claimed. Epigenetics does not let Nia rewrite her risk; it lets her change expression, with modest effect sizes at individual loci. Exercise will not move her polygenic score, which is fixed. And a favourable lifestyle halved the relative rate in that study — it did not abolish risk, and 5.1% is not zero.

The accurate summary is the one worth giving a patient: her genome sets the slope; her exposure-years set how far along it she travels; and she is twenty-four, which means almost all of those years are still unspent.


Systems Integration Case File · Entry 29

Entry 29 — The pedigree, and what it does and does not explain

New findings.

Amara, 45 LDL 168 mg/dL · HDL 38 · TG 210 · Lp(a) 42 nmol/L
Genetic panel No pathogenic variant in LDLR, APOB, PCSK9, LDLRAP1. One LDLR VUS
Polygenic risk score, CAD 92nd percentile (European-ancestry reference; see limitations)
Pharmacogenomics CYP2C19 *1/*2 — intermediate metabolizer
Family Father d. 58 MI · brother Kofi MI at 51 · mother Adwoa 78, hypertension only
Nia, 24 Normotensive, no metabolic abnormality, PRS not measured

Your entry:

1 · ADD. State in two or three sentences what the genetic findings add to your model of Amara that the clinical findings did not already contain — and be careful to say what they do not add.

2 · CONNECT. Link the genetic findings to at least two systems already in your file, stating the direction of causation. One of your links must involve the CYP2C19 result and a decision made in an earlier chapter.

3 · PREDICT. Chapter 30 is about aging. Predict one way in which Adwoa at 78, Amara at 45, and Nia at 24 will differ in a variable you have already recorded — and say whether you expect the difference to be genetic, cumulative, or both.

Model responses — read only after writing your own

1 · ADD. The genetic findings convert "runs in the family" into a mechanism and a magnitude: Amara's risk is polygenic, distributed across millions of small-effect variants acting through lipoprotein handling, blood pressure control, endothelial function, and inflammation, and her burden sits near the top of the distribution. Critically, they add something negative as well — the absence of a pathogenic LDLR variant means there is no single-gene lesion to cascade-test relatives for, and the LDLR VUS must be treated as a non-result. What they do not add is prognosis for any individual, a causal explanation, or a well-calibrated number, since the score was computed against a European-ancestry reference panel and the family is of West African descent.

2 · CONNECT. Genetics → hepatic lipoprotein handling → arterial wall: polygenic variants that modestly reduce LDL receptor expression and impair triglyceride clearance cause a sustained LDL of 168 and TG of 210, which over twenty-five years causes subendothelial apoB retention and plaque, which caused the coronary lesions found in Chapter 18. Genetics → renal and vascular control → cardiac load: variants affecting renal sodium handling and the renin–angiotensin system cause hypertension, which causes increased afterload and left ventricular remodeling and accelerates the glomerular injury behind her stage 3 CKD — closing the cardiorenal loop from Chapter 26 with an upstream genetic term. Pharmacogenomics → hematology → cardiology: Amara is a CYP2C19 intermediate metabolizer, and clopidogrel is a prodrug requiring CYP2C19 for activation — so the antiplatelet regimen chosen after her NSTEMI in Chapter 17 may be delivering less platelet inhibition than intended, raising the risk of stent thrombosis. This is a genotype changing a decision that was already made about a different organ system.

3 · PREDICT. Blood pressure is the cleanest candidate: Nia 102/58, Amara 168/98 before treatment, Adwoa hypertensive since her fifties. The prediction is that Chapter 30 will show arterial stiffening — progressive fragmentation of elastin and accumulation of cross-linked collagen in the aortic wall — raising systolic pressure and pulse pressure with age in everyone, so that isolated systolic hypertension is Adwoa's pattern while Amara's is a combined elevation. The difference is both: the polygenic burden sets each woman's starting slope, and the arterial wall accumulates the consequences over decades. A second good answer is bone density, where Adwoa's osteoporosis reflects both a heritable peak bone mass and forty years of post-menopausal loss, and where Nia's running is currently building the peak she will spend the rest of her life drawing down.


Review

Level 1 · Recall

25.1 Which base pairs with adenine in DNA?

a) guanine    b) cytosine    c) thymine    d) uracil

Answer

c — thymine, through two hydrogen bonds. Uracil (d) replaces thymine in RNA, so adenine pairs with uracil in an RNA duplex or in an RNA–DNA hybrid, but not in DNA. Guanine pairs with cytosine through three hydrogen bonds, which is why GC-rich DNA melts at a higher temperature.

25.2 A mutation that deletes two base pairs from the middle of a coding exon most likely produces:

a) a silent change    b) a frameshift with a premature stop    c) the loss of one amino acid    d) a repeat expansion

Answer

b. Two is not divisible by three, so every codon downstream is misread and a premature stop codon almost always appears within a few dozen codons; the transcript is usually degraded by nonsense-mediated decay. (c) would require a deletion of exactly three base pairs — the ΔF508 situation in cystic fibrosis. This one-base-pair distinction separates a mildly abnormal protein from no protein at all.

25.3 Crossing over occurs during:

a) mitosis    b) prophase I of meiosis    c) metaphase II of meiosis    d) fertilization

Answer

b — prophase I, at chiasmata within the synaptonemal complex, with roughly 45–55 crossovers per human meiosis. It is one of three sources of gametic variation, alongside independent assortment at metaphase I and random fertilization.

25.4 A pedigree shows affected individuals in every generation, both sexes affected about equally, and several instances of an affected father with an affected son. The pattern is:

a) autosomal dominant    b) autosomal recessive    c) X-linked recessive    d) mitochondrial

Answer

a — autosomal dominant. No skipping means dominant; male-to-male transmission excludes any X-linked pattern, because a father gives his son a Y; and equal involvement of the sexes excludes X-linkage as well. Mitochondrial inheritance (d) is excluded outright, since an affected father transmits to none of his children.

25.5 Which is transmitted exclusively from mother to all of her children?

a) an X-linked recessive allele    b) a mitochondrial DNA variant    c) an imprinted paternal allele    d) a Y-linked gene

Answer

b. Paternal mitochondria entering the oocyte at fertilization are ubiquitinated and destroyed, so mitochondrial DNA is transmitted only through the oocyte's cytoplasm, and every child of an affected mother inherits it. Severity varies because of heteroplasmy and the threshold effect. An X-linked recessive allele (a) goes to half of a carrier mother's children; Y-linked genes (d) pass strictly father to son.

25.6 Heritability of 0.8 for a trait means:

a) 80% of an individual's value for the trait is caused by their genes b) 80% of the variance of the trait in that population is attributable to genetic variance c) the trait cannot be changed by the environment d) 80% of people with the risk genotype will develop the trait

Answer

b. Heritability is a statistic about variance in a population, in a particular environment. (a) is the classic error: heritability says nothing about the composition of any individual. (c) is the clinically dangerous error — phenylketonuria is essentially 100% heritable and completely preventable by diet. (d) describes penetrance, a different concept entirely.

25.7 DNA methylation of a promoter CpG island typically:

a) activates transcription    b) silences transcription    c) causes a frameshift    d) changes the DNA sequence

Answer

b — silences it, both by physically obstructing transcription factor binding and by recruiting methyl-CpG-binding proteins that bring histone deacetylases, closing the chromatin in a self-reinforcing loop. Note (d): methylation adds a chemical group to a cytosine but does not change which base it is, which is exactly why epigenetic marks are reversible while mutations are not.

25.8 A cell-free DNA screen at 12 weeks is positive for trisomy 21 in a 25-year-old. The correct next step is:

a) treat the result as diagnostic    b) offer a diagnostic test such as CVS or amniocentesis    c) repeat the cell-free DNA test    d) no further action

Answer

b. Cell-free DNA screening has excellent sensitivity (>99%) and specificity (~99.9%), but positive predictive value depends on prevalence, and at 25 the prior probability of trisomy 21 is low enough that a positive result is right only about half the time. The DNA analysed is also placental rather than fetal, so confined placental mosaicism is a real source of false positives. A screen is never diagnostic; confirmation requires CVS or amniocentesis.

Level 2 · Comprehension

25.9 Explain how approximately 20,000 protein-coding genes produce well over 100,000 proteins.

Model answer

Three mechanisms multiply, and the first does most of the work.

Alternative splicing. After transcription, introns are removed and exons joined by the spliceosome — but which exons are retained can differ between cells and conditions. About 95% of human multi-exon genes are alternatively spliced, so one gene routinely yields several distinct mRNAs and therefore several proteins. Cardiac and skeletal troponin T come from one gene; so do the membrane-bound and secreted forms of immunoglobulin heavy chain.

Alternative promoters and polyadenylation sites produce transcripts with different regulatory ends and sometimes different first exons, changing localization or stability.

Post-translational modification multiplies again: cleavage (proopiomelanocortin is cut into ACTH, β-endorphin, and MSH in different cells), glycosylation, phosphorylation, and other modifications produce functionally distinct molecules from an identical polypeptide.

The conceptual point is that "one gene, one protein" was a useful early approximation that is simply false in humans, and its failure is why gene count correlates so poorly with organismal complexity — a nematode has about as many genes as we do.

25.10 A couple are both carriers of a recessive condition. Their first three children are unaffected. What is the risk to the fourth pregnancy, and what is the probability that their eldest child is a carrier? Explain both.

Model answer

The risk to the fourth pregnancy is 1 in 4, exactly as it was for the first. Each fertilization is an independent event: which allele each parent transmits is determined afresh at each meiosis, and the previous outcomes have no influence. This is the most common intuitive error in genetic counseling, and it is the gambler's fallacy in medical clothing.

The eldest child is 2 in 3 likely to be a carrier, not 1 in 2. The Punnett square gives four equally likely genotypes: FF, Ff, fF, ff. We have observed that the child is unaffected, which eliminates ff. Of the three remaining possibilities, two are carriers. This is conditional probability — the observation changed the denominator, not the underlying odds — and it is exactly the calculation that drives real carrier-risk arithmetic in a family with an affected relative.

25.11 Amara's polygenic risk score is at the 92nd percentile. Explain what that number does and does not tell her, and why the reference population matters.

Model answer

What it tells her: that the weighted sum of her coronary-disease-associated alleles is higher than that of 92% of people in the reference population. Individuals in the top few percent of such a distribution have roughly three- to four-fold the coronary event rate of those in the middle — comparable in magnitude to carrying a pathogenic familial hypercholesterolemia variant, though arising from millions of small effects rather than one large one.

What it does not tell her: whether she has coronary disease (that is a clinical question, and in her case already answered), whether she will have an event (it is a probability shift, not a determination), why she is at risk (most associated variants are regulatory and their causal targets are often unknown), or what she should do (management follows measured LDL, blood pressure, glucose, and clinical findings, none of which the score replaces).

Why the reference population matters: polygenic scores are built from genome-wide association studies conducted overwhelmingly in people of European ancestry. The variants genotyped are usually not the causal ones but markers that happen to sit nearby, and which markers travel with which causal variants — the linkage disequilibrium structure — differs between populations, as do allele frequencies and effect sizes. Applied across an ancestry boundary, predictive accuracy typically falls two- to fivefold. For a family of West African descent, a score calibrated on European-ancestry data is a real signal measured with the wrong ruler, and the counseling should say so explicitly rather than reporting a percentile as though it were a laboratory value.

Level 3 · Clinical Application

25.12 A healthy 30-year-old man's sister has cystic fibrosis. His partner has no family history and is of Northern European ancestry (population carrier frequency 1 in 25). Calculate their per-pregnancy risk of an affected child. Then recalculate after the partner has a negative carrier screen on a panel that detects 90% of variants.

Model answer

Before screening.

His parents must both be carriers, since they had an affected child. He is unaffected, so his genotype is FF, Ff, or fF with equal probability: P(carrier) = 2/3.

His partner has no family history, so her carrier probability is the population frequency: 1/25.

If both are carriers, each pregnancy carries a 1/4 risk.

P(affected child) = 2/3 × 1/25 × 1/4 = 2/300 = 1 in 150.

After a negative screen with 90% detection.

The screen only misses a carrier 10% of the time, so use Bayes' theorem. Prior odds of carrier : non-carrier = 1/25 : 24/25. The probability of testing negative is 0.10 if she is a carrier and 1.0 if she is not.

P(carrier | negative) = (1/25 × 0.10) ÷ [(1/25 × 0.10) + (24/25 × 1.0)] = 0.004 ÷ (0.004 + 0.96) ≈ 1 in 240.

Revised per-pregnancy risk = 2/3 × 1/240 × 1/4 ≈ 1 in 1,440.

The clinical point. A negative screen reduced the risk roughly tenfold — from 1 in 150 to 1 in 1,440 — but it did not eliminate it, because no panel detects every variant. Counseling that reports a negative screen as "you're clear" is wrong. If greater certainty is wanted, the correct next step is to sequence the affected sister to identify the family's specific variants, then test him directly for those; that converts his 2/3 into a definite yes or no and is far more informative than any population panel.

25.13 A 6-month-old boy has hypotonia, poor feeding, and failure to thrive; his mother had noted decreased fetal movement. Genetic testing shows a deletion in the paternally inherited chromosome 15q11-13. Name the condition, explain the mechanism, and explain why an identically sized deletion on the maternal chromosome 15 would produce a different disease.

Model answer

This is Prader-Willi syndrome. The classic course is exactly as described — decreased fetal movement, neonatal hypotonia and feeding difficulty — followed after age 2 by hyperphagia and obesity, short stature, hypogonadism, and intellectual disability.

The mechanism is genomic imprinting. Several genes in the 15q11-13 interval are expressed only from the paternal chromosome, because the maternal copies are silenced by methylation marks established during oogenesis. An individual therefore has only one functional copy of those genes to begin with. Deleting the paternal copy leaves no expression at all, and the phenotype follows. The same result arises from maternal uniparental disomy — inheriting both chromosome 15s from the mother, so that even though the chromosome count is normal, both copies are imprinted off.

An identically sized deletion on the maternal chromosome 15 removes UBE3A, which is imprinted in the opposite direction — silenced on the paternal chromosome in neurons and expressed only from the maternal one. Losing the maternal copy therefore abolishes neuronal UBE3A expression and produces Angelman syndrome: severe intellectual disability, absent speech, ataxia, seizures, and a characteristically happy demeanour.

Same chromosomal region, same size of lesion, opposite parent of origin, two entirely different diseases. This is the clearest demonstration in human genetics that Mendel's assumption — that it does not matter which parent an allele came from — is not universally true.

25.14 Amara has had a coronary stent placed and is on clopidogrel. Her CYP2C19 genotype is reported as 1/2. Explain the clinical significance and what should be considered.

Model answer

Clopidogrel is a prodrug. It has no antiplatelet activity as administered; it must be oxidised, in two sequential steps that both depend heavily on CYP2C19, into an active thiol metabolite that irreversibly blocks the platelet P2Y₁₂ ADP receptor and prevents platelet activation and aggregation (Chapter 17).

The 2 allele is a loss-of-function variant. A 1/2 genotype makes Amara an intermediate metabolizer*: she converts less of the prodrug, achieves less platelet inhibition, and — in large studies of patients treated with percutaneous coronary intervention — carries a measurably higher rate of stent thrombosis and recurrent ischemic events than normal metabolizers. Roughly 30% of people of European ancestry, and a higher proportion of East Asian ancestry, carry at least one loss-of-function allele.

What should be considered: switching to an antiplatelet agent that does not require CYP2C19 activation — prasugrel or ticagrelor — since neither is affected by CYP2C19 genotype. Guidelines increasingly support genotype-guided selection after PCI for this reason.

Two broader points. First, note the general principle: a pharmacogenomic variant matters most for prodrugs, because the genotype controls whether the drug is ever converted to its active form at all. Second, this is a place where a genetic result changes a decision made about a different organ system in an earlier chapter — the cardiology decision to stent, and the hematology decision about antiplatelet therapy, both turn out to depend on a hepatic enzyme genotype.

Level 4 · Integration and Synthesis

25.15 Construct an argument, using material from at least four chapters, that Amara's coronary artery disease was neither "genetic" nor "lifestyle" but the predictable output of a system. Identify the point at which intervention would have had the largest effect and justify your choice quantitatively.

Model answer

The system. Amara's arterial wall integrates the outputs of at least five subsystems, each of which has a genetic term and an environmental term.

  1. Lipoprotein handling (Chapter 24 and §29.8). Polygenic variants modestly reducing LDL receptor expression and impairing triglyceride clearance set her LDL near 168 mg/dL and her HDL near 38. Diet and adiposity modulate that number but do not set its baseline.
  2. Insulin sensitivity (Chapter 16, §28.6). A polygenic predisposition plus twenty years of night-shift circadian disruption plus a BMI of 29.3 produced insulin resistance, which raises triglycerides, lowers HDL, and generates small dense LDL particles — the exact lipid pattern she has.
  3. Blood pressure control (Chapters 19, 26). Polygenic variants in renal sodium handling and the renin–angiotensin system, plus sympathetic activation from chronic sleep deprivation, produced hypertension, which raises endothelial shear stress and accelerates plaque.
  4. Endothelial function (Chapter 19). Reduced nitric oxide availability from hyperglycemia, dyslipidemia, and hypertension permits lipoprotein retention and monocyte adhesion — and, perhaps, was already compromised by a developmental substrate (§28.9).
  5. Inflammation (Chapter 21, plus clonal hematopoiesis from §29.5). Macrophage-driven IL-1β and IL-6 signalling converts a lipid deposit into an active, rupture-prone plaque.

Why "genetic versus lifestyle" is the wrong frame. No single subsystem is broken. Each is tilted a few percent, and the arterial wall computes the sum over decades. The polygenic score captures the tilt; it does not capture the decades. The lifestyle history captures the decades; it does not capture the tilt. Neither alone predicts the outcome, and the interaction data show they are roughly additive on the absolute risk scale — favourable lifestyle roughly halved ten-year event rates within the highest genetic risk quintile (10.7% → 5.1%).

Where intervention would have mattered most, and why. Not at 45, when she presented, and not at 42, when her blood pressure was first noted to be high. In her twenties, and specifically on LDL. The quantitative justification is the Mendelian randomization comparison: a lifelong 1 mmol/L lower LDL is associated with an 80–88% lower coronary risk, while the same 1 mmol/L reduction achieved with a drug started at 60 lowers risk by 20–25%. Benefit scales with the area under the LDL-versus-time curve, so a year of exposure removed at 25 is worth several years removed at 60. A secondary, nearly as strong argument applies to the night shift: twenty years of circadian disruption acted on insulin sensitivity, blood pressure, and appetite regulation simultaneously, and it began in her twenties too.

The uncomfortable conclusion. The intervention with the largest possible effect on Amara's coronary disease would have been made two decades before she had any symptom, any abnormal number, or any reason to see a physician — which is precisely why the person in the family who should be acting on this information is Nia.

25.16 A colleague argues: "Amara's polygenic risk score is 92nd percentile and coronary disease is about 50% heritable, so roughly half of her disease is genetic and unchangeable, and half is lifestyle." Identify every error in that sentence and write a corrected version.

Model answer

There are four distinct errors, and they compound.

Error 1 — heritability applied to an individual. Heritability is the proportion of variance in a population attributable to genetic variance. It cannot be partitioned within one person. "Half of Amara's disease is genetic" is not a claim that can be true or false; it is a category error, like asking what percentage of a rectangle's area is due to its length.

Error 2 — heritability treated as fixed. h² depends on how much environmental variation the studied population contains. In a population with uniform behaviour, heritability of coronary disease would approach 1; in one with wildly variable smoking and diet, it falls. The number partly measures the environment, not just the genes.

Error 3 — heritability equated with unmodifiability. This is the clinically dangerous one. Phenylketonuria is essentially 100% heritable and completely preventable by diet; adult height is about 80% heritable and rose 20 cm in the Netherlands in 150 years. A trait's heritability places no upper bound on how much an intervention can change it.

Error 4 — the percentile treated as a proportion of causation. A 92nd-percentile score means her position in a distribution of allele burden, not that 92% of anything is genetic. And the score's calibration is uncertain here anyway, since it was derived in a European-ancestry reference population and Amara is of West African descent.

A corrected version: "Amara's genetic burden of common coronary risk variants is higher than that of about 92% of people in the reference population, which in large cohorts corresponds to roughly three- to four-fold the event rate of someone in the middle — although that estimate is less reliable across the ancestry boundary. That burden is fixed, but it sets a slope rather than an outcome: in people at similarly high genetic risk, favourable lifestyle was associated with roughly half the ten-year event rate. Her disease is the accumulated product of both, and the part that remains modifiable is substantial."

Concept Map to Complete

Copy onto blank paper and fill every bracket from memory before checking.

                            DNA  (3.1 × 10⁹ bp, ~[ ______ ] genes)
                                        │
              ┌─────────────────────────┼──────────────────────────┐
              ▼                         ▼                          ▼
        REPLICATION              TRANSCRIPTION              MUTATION
        semiconservative         │                          │
        leading / [ _______ ]    ▼                    ┌─────┼──────┬────────┐
        fidelity from:           pre-mRNA         point  [ _______ ]  repeat
         ① base pairing          │  processing:   │      │           expansion
         ② [ ___________ ]       │  5' [ ___ ]    ├ silent│  ÷3? YES→ one aa
         ③ [ ___________ ]       │  3' [ ___ ]    ├ [ ____]│  ÷3? NO → [ ___ ]
              │                  │  [ _______ ]   └ nonsense
        failure → [ ______ ]     ▼
                            mature mRNA ──► TRANSLATION at the [ _______ ]
                            ALTERNATIVE                A / P / E sites
                            [ _________ ]                    │
                            = 1 gene → many proteins    POST-TRANSLATIONAL
                                                        [ _____________ ]

   ══ INHERITANCE ═══════════════════════════════════════════════════════
   MEIOSIS: 3 sources of variation
      ① [ ______________ ] in prophase I
      ② [ ______________ ] at metaphase I → 2^[ __ ] combinations
      ③ [ ______________ ]
   failure to separate = [ ______________ ] → aneuploidy

   FOUR MENDELIAN PATTERNS — the three sorting questions
      skips generations?   no → [ ________ ]   yes → [ _________ ]
      male-to-male?        yes → [ ________ ]  (excludes [ ________ ])
      sexes equal?         males ≫ → [ _____________________ ]

   BEYOND MENDEL
      [ ______________ ] one gene, many effects (Marfan)
      [ ______________ ] one gene masks another (Bombay)
      [ ______________ ] parent of origin decides (Prader-Willi/Angelman)
      [ ______________ ] strictly maternal, heteroplasmic (MELAS)

   ══ POLYGENIC RISK ════════════════════════════════════════════════════
      millions of variants × [ ___ ]–[ ___ ]% each  →  a [ __________ ]
      heritability = proportion of [ ________ ] in a [ ___________ ]
                     — NOT a statement about [ ______________ ]
                     — NOT a statement about [ ______________ ]
      EPIGENETICS: [ ___________ ] of CpG islands → gene [ _________ ]
                   histone [ ___________ ] → chromatin OPEN
      ⇒ sequence is [ ______ ]; expression is [ ______ ]

Lab / Self-Exploration

  1. Draw your own pedigree. Three generations, standard symbols — squares for males, circles for females, filled for affected. Mark any condition appearing more than once: hypertension, diabetes, cancer, early cardiac events, hearing loss. Then apply the three sorting questions from Figure 29.5 and decide whether what you see is Mendelian or polygenic. Most families produce the second answer, and recognizing that is the skill.
  2. Find your "premature" relatives. Using the clinical definition — a first-degree relative with cardiovascular disease before 55 in men, 65 in women — determine whether you have a positive family history. This single binary variable appears in most clinical risk calculators, and you now know why.
  3. Survey observable variation. In a group of ten people, record earlobe attachment, the ability to roll the tongue, and hair whorl direction. Then look up how each is actually inherited. All three are traditionally taught as simple dominants and all three are, in fact, polygenic or poorly characterized — which is a useful demonstration of how much of classical "human Mendelian genetics" was wishful.
  4. Do the Bayes calculation yourself. Take the worked example in Figure 29.4 and redo it with a panel detection rate of 95% instead of 90%, and then with a carrier frequency of 1 in 50 instead of 1 in 25. Notice which input the final number is most sensitive to.
  5. Read a real consumer genetics report. If you or someone you know has one, find the methods section. Determine whether it is a genotyping array or sequencing, how many positions it interrogates, and what proportion of known pathogenic variants in any single gene it can detect. Then read its wording on risk and identify every place where an association is being presented as a prediction.
  6. Trace one trait through Figure 29.2. Pick any protein you have met in this book — insulin, collagen, hemoglobin, dystrophin — and write out the seven numbered stages from gene to finished molecule, naming what would go wrong at each stage if it failed. You will find a real human disease at almost every one.

Key Terms

allele · One of the alternative DNA sequences that can occupy a given locus.

alternative splicing · Joining different combinations of exons from one pre-mRNA, so that a single gene yields multiple proteins; occurs in ~95% of human multi-exon genes.

aneuploidy · An abnormal chromosome number, produced by nondisjunction.

anticipation · Earlier onset and greater severity in successive generations, caused by expansion of an unstable repeat.

autosome · Any chromosome other than X or Y; humans have 22 pairs.

carrier · A clinically unaffected heterozygote for a recessive allele.

central dogma · The flow of information from DNA to RNA to protein.

clonal hematopoiesis · Age-related expansion of a mutant hematopoietic stem cell clone; raises risks of hematological malignancy and, independently, of coronary disease.

codominance · Both alleles fully expressed in the heterozygote, as in blood type AB.

complementary base pairing · A with T (two hydrogen bonds), G with C (three); the basis of both information storage and replication.

epigenetics · Heritable changes in gene expression without change in DNA sequence, mediated principally by DNA methylation and histone modification.

epistasis · One gene masking the expression of another at a different locus.

expressivity · How severely a phenotype is expressed among those who express it at all.

frameshift · An insertion or deletion not divisible by three, misreading every downstream codon and usually creating a premature stop.

gene · A DNA sequence specifying a functional product; ~19,000–20,000 protein-coding genes occupy about 1.5% of the human genome.

genetic code · The triplet, degenerate, unambiguous, non-overlapping, nearly universal correspondence between codons and amino acids.

genotype / phenotype · The alleles carried / the observable characteristic.

heritability (h²) · The proportion of phenotypic variance in a population, in a given environment, attributable to genetic variance. Not a statement about an individual and not a statement about modifiability.

heteroplasmy · Coexistence of mutant and normal mitochondrial genomes in one cell; symptoms appear above a tissue-specific threshold.

homologous chromosomes · The maternal and paternal members of a chromosome pair, carrying the same loci in the same order but not necessarily the same alleles.

imprinting · Parent-of-origin-specific silencing of a gene by germline methylation; produces Prader-Willi and Angelman syndromes from the same chromosomal region.

independent assortment · Random orientation of each homologous pair at metaphase I, giving 2²³ combinations per gamete.

karyotype · An ordered display of an individual's metaphase chromosomes; resolution 5–10 Mb.

locus · The physical position of a gene on a chromosome.

Lyonization (X-inactivation) · Random permanent silencing of one X in each female somatic cell, making every female a mosaic; the silenced X is the Barr body.

meiosis · Two divisions producing four genetically distinct haploid gametes; meiosis I is reductional, meiosis II equational.

missense mutation · A single base change substituting one amino acid for another.

mismatch repair · Post-replication correction of mispaired bases; when inherited defective, causes Lynch syndrome.

mutation · A heritable change in DNA sequence.

nondisjunction · Failure of chromosomes (meiosis I) or sister chromatids (meiosis II) to separate, producing gametes with an extra or missing chromosome.

nonsense mutation · A base change creating a premature stop codon.

penetrance · The proportion of individuals with a genotype who show any phenotype at all.

pleiotropy · One gene producing multiple apparently unrelated effects, as FBN1 does in Marfan syndrome.

polygenic risk score · A weighted sum of risk alleles across many loci, expressed as a percentile; predictive, not diagnostic, and calibrated to a specific ancestry.

Punnett square · A grid enumerating the possible allele combinations from a cross.

repeat expansion · Increase in the copy number of a short tandem repeat; the mechanism of Huntington disease, fragile X, and myotonic dystrophy.

semiconservative replication · Each daughter duplex retains one parental strand.

silent mutation · A base change that leaves the encoded amino acid unaltered.

transcription / translation · DNA to RNA / RNA to protein.

variant of uncertain significance (VUS) · A sequence change whose clinical effect is unknown; must not be used to guide management.


Next: Chapter 30 · Aging and the Body Systems — where Adwoa at 78, Amara at 45, and Nia at 24 are examined as the same systems at three points on one curve, and where the accumulated exposure of this chapter becomes visible in every organ.