Affiliate disclosure

Book titles on this page link to Amazon. As an Amazon Associate, DataField.Dev earns from qualifying purchases — at no additional cost to you.

Bibliography

Every source recommended anywhere in this book, in one list, with the chapters that recommend it. This file is generated by scripts/assemble.py from each chapter's further-reading.md — edit the chapter file, not this one.

Sources are tagged Tier 1 (we are confident the work exists and recommend it without reservation) or Tier 2 (real and worth seeking, but confirm the current edition, version, or URL yourself — this field's documentation drifts fast).


  • "Hidden Technical Debt in Machine Learning Systems" (Sculley et al., NeurIPS 2015). Nine pages, (Ch. 32)

  • "Monitoring Machine Learning Models in Production" material from the ML-monitoring vendors — Evidently (Ch. 32)

  • "Pagination: You're (probably) doing it wrong" and similar write-ups on keyset versus offset (Ch. 16)

  • "The ML Test Score: A Rubric for ML Production Readiness" (Breck et al., Google, 2017). A checklist (Ch. 32)

  • Advent of Code, solved in SQL. An idiosyncratic but effective (Ch. 18)

  • explain.dalibo.com and pev2. (Ch. 18)

  • pgexercises.com. Free, PostgreSQL-based, and the aggregate and (Ch. 18)

  • 0.30000000000000004.com. A one-page site listing the output of 0.1 + 0.2 in dozens of (Ch. 7)

  • airflow standalone. A scheduler, web server, and SQLite metadata database in one command and (Ch. 24)

  • code/capstone.py. Thirty-eight assertions, including the two that check the result against (Ch. 38)

  • code/catalog_audit.py in this chapter. Eleven findings in a deliberately messy fixture, twenty-two (Ch. 30)

  • code/cost_model.py in this chapter. Three meters, an attributed bill, seven sized wastes, and a (Ch. 33)

  • code/coverage.py in this chapter. Scores a dbt manifest against the six assertions and lists (Ch. 23)

  • code/dag_lint.py in this chapter. Parses DAG source with ast — no Airflow required — and (Ch. 24)

  • code/engine_benchmark.py in this chapter. Generates its own data, runs all four engines in (Ch. 22)

  • code/event_lab.py in this chapter. The dual write measured against the outbox, four projections (Ch. 36)

  • code/health.py in this chapter. No dependencies. Feed it a five-column run history and it (Ch. 25)

  • code/interview_drills.py in this chapter. Twelve problems whose wrong answers are enumerated with (Ch. 39)

  • code/layer_check.py in this chapter. Nine rules, a replay planner with costs, depth and blast (Ch. 34)

  • code/mesh_readiness.py in this chapter. Coupling and fan-out, a classified intake queue, weighted (Ch. 35)

  • code/migration_lab.py in this chapter. An estate scored and ordered, sixty days of classified (Ch. 37)

  • code/pii_scan.py in this chapter. Detection scored three ways, a k-anonymity ladder, and a (Ch. 31)

  • code/pit_join.py in this chapter. Leakage measured, an as-of join with the invariant test, feature (Ch. 32)

  • code/plan_review.py in this chapter. Parses terraform show -json and blocks on stateful (Ch. 28)

  • code/pr_report.py in this chapter. Computes blast radius, exposures, deploy shape, rebuild (Ch. 27)

  • code/scd_lab.py in this chapter. Runs every failure here against SQLite in about a second — (Ch. 20)

  • code/slo.py in this chapter. Attainment, error budget, burn rate over a short window, the (Ch. 26)

  • code/spark_advisor.py in this chapter. Reads a plan and counts the four things §21.3 says to (Ch. 21)

  • code/stream_harness.py in this chapter. Watermarks, tumbling and session windows, allowed (Ch. 29)

  • dbt-project-evaluator — dbt Labs' own package of project-structure checks. It implements (Ch. 34)

  • dbt_artifacts and similar packages that load run_results.json into your warehouse. This is (Ch. 25)

  • elementary and the open-source dbt observability tools. They read run_results.json and (Ch. 19)

  • EXPLAIN (ANALYZE, BUFFERS) and the PostgreSQL documentation on "Using EXPLAIN." §18.10's (Ch. 18)

  • explain.depesz.com and explain.dalibo.com. Two free tools that take an `EXPLAIN (ANALYZE, (Ch. 7)

  • featureform, hopsworks, and the smaller open-source stores. Worth reading the data models of. (Ch. 32)

  • pyspark in local mode. SparkSession.builder.master("local[4]") reproduces skew, coalesce (Ch. 21)

  • sqlfluff, at sqlfluff.com. A SQL linter with a dbt templater, so it (Ch. 19)

  • Accelerate (Forsgren, Humble, Kim) on lead time and deployment frequency as team-level metrics. (Ch. 35)

  • Adrian Cockcroft's writing and talks on evolutionary architecture at Netflix. Useful as a (Ch. 3)

  • Akidau et al., "The Dataflow Model" (VLDB 2015). The paper the book grew out of, free, and about (Ch. 29)

  • Alberto Brandolini's material on Event Storming. A workshop technique for discovering events with (Ch. 36)

  • Alex Xu, System Design Interview (vols. 1–2). Widely used, and written for general software (Ch. 39)

  • And the accounting literature on reconciliation, one more time. It is the oldest idea in this book (Ch. 40)

  • Andreas Andreakis and Ioannis Papapanagiotou, "DBLog: A Watermark Based Change-Data-Capture (Ch. 14)

  • Andrew Jones, Driving Data Quality with Data Contracts (Packt, 2023). A book-length treatment (Ch. 17)

  • Anthony Molinaro and Robert de Graaf, SQL Cookbook (2nd ed., O'Reilly, 2020). Problem-first, (Ch. 18)

  • Anurag Gupta et al., "Amazon Redshift and the Case for Simpler Data Warehouses" (2015), (Ch. 8)

  • Any careful treatment of the small-files problem on object storage. The good sources are usually (Ch. 21)

  • Any careful write-up on embedding model versioning and re-indexing. This is the least (Ch. 12)

  • Any good treatment of confirmation bias in operational data. The specific failure — a small, (Ch. 18)

  • Any honest account of on-call attrition. It is mostly in blog posts and conference talks rather (Ch. 26)

  • Any honest comparison of the three. They are rare; most are written by one of the three. (Ch. 24)

  • Any introductory financial accounting text, on the reconciliation and the trial balance. Two hours, (Ch. 38)

  • Any material on load testing and on chaos engineering, particularly the principle of (Ch. 13)

  • Any of the "why is my Delta/Iceberg table slow" write-ups from practitioners. These are the (Ch. 10)

  • Any postmortem you can find of a replication-slot disk-full incident. They exist, they are (Ch. 14)

  • Any recent benchmark comparing single-node engines to distributed ones — read skeptically. (Ch. 21)

  • Any retrospective on a dead technology written by someone who used it. The Hadoop retrospectives (Ch. 40)

  • Any vendor's "why the data warehouse is dead" or "why the data lake failed" content. Read two (Ch. 3)

  • Any vendor's format comparison. Read two from competing vendors and note what each chose to (Ch. 11)

  • Any write-up of a Kafka rebalance incident. They are numerous, they are consistent, and the (Ch. 15)

  • Anything careful arguing that data mesh is Conway's law with a new name. The critique is fair, it (Ch. 35)

  • Anything careful on "characterization tests." Feathers's term for a test that documents what the (Ch. 37)

  • Anything careful on alert fatigue, from operations rather than from data. The literature is (Ch. 23)

  • Anything careful on clock skew in distributed systems. Kleppmann's Chapter 8 ("The Trouble with (Ch. 20)

  • Anything careful on cloud cost attribution and tagging. The vendor documentation is adequate and (Ch. 28)

  • Anything careful on CSV as an interchange format. RFC 4180 exists, describes what most people (Ch. 22)

  • Anything careful on data masking and synthetic generation. The tooling here is immature and (Ch. 27)

  • Anything careful on idempotent consumers and effectively-once processing. The Kafka documentation's (Ch. 36)

  • Anything careful on property-based testing. The capstone's strongest assertions are properties, not (Ch. 38)

  • Anything careful on requirements elicitation. Case Study 2's four questions are a small, specific (Ch. 29, Ch. 39)

  • Anything careful on retrospective and revisable statistics. Economics has the best-developed (Ch. 26)

  • Anything careful on the "documentation is written by the wrong person" problem. The (Ch. 30)

  • Anything careful on the economics of sunk cost and its behavioural effects. §33.10's claim — that (Ch. 33)

  • Anything careful on weak supervision and programmatic labeling — the Snorkel line of work is the (Ch. 32)

  • Anything on "bus factor" and knowledge transfer. §38.12's test — hand the README to someone who has (Ch. 38)

  • Anything on "dark launching" and "shadow traffic." The service-side equivalent, better documented (Ch. 37)

  • Anything on "the ratchet" in operational process. Case Study 2's central mechanism — every (Ch. 25)

  • Anything on "the summary line problem" in tooling output. There is no canonical source and the (Ch. 28)

  • Anything on alert fatigue — the same recommendation as Chapter 23, for the same reason. Case (Ch. 24)

  • Anything on architectural fitness functions. The idea — an automated test that a system still has (Ch. 34)

  • Anything on blameless postmortems, particularly Google's SRE material. §17.7's fourth fix — that (Ch. 17)

  • Anything on controls that fail open, from safety engineering. Case Study 2's "a check whose (Ch. 27)

  • Anything on Conway's law and on the relationship between team boundaries and system (Ch. 17)

  • Anything on CRDTs (conflict-free replicated data types). §36.9's commutativity analysis is the (Ch. 36)

  • Anything on crypto-shredding / cryptographic erasure. NIST SP 800-88 (Guidelines for Media (Ch. 31)*

  • Anything on database migration tooling: Flyway, Liquibase, Alembic, or dbt snapshots. The (Ch. 27)

  • Anything on evaluating engineering culture from outside — the "questions to ask your interviewer" (Ch. 39)

  • Anything on event sourcing's "the log is the source of truth." Chapter 36 takes this seriously. (Ch. 34)

  • Anything on graph modularity and community detection — the Louvain and Leiden algorithms are the (Ch. 35)

  • Anything on property-based testing applied to event ordering. §29.9's second case — out-of-order (Ch. 29)*

  • Anything on reading the SQL a GUI ETL tool generates. Informatica, DataStage, SSIS, and Talend all (Ch. 37)

  • Anything on reproducible builds — the Reproducible Builds project in software packaging is the (Ch. 38)

  • Anything on revenue recognition, particularly the treatment of gift cards and refunds. §38.5's R3 (Ch. 38)

  • Anything on runbooks and operational readiness reviews. Google's SRE material has the best-known (Ch. 38)

  • Anything on SaaS unit economics and cost of goods sold. §33.11's cost-per-order and percent-of-GMV (Ch. 33)

  • Anything on shadow IT discovery. The information-systems literature has thirty years on finding (Ch. 37)

  • Anything on shadow IT, which has a thirty-year literature. The consistent finding — shadow (Ch. 35)

  • Anything on structured behavioral interviewing, from the interviewer's side rather than the (Ch. 39)

  • Anything on Team Topologies (Skelton and Pais). The single most useful non-data book for this (Ch. 35)

  • Anything on the "manager versus IC" fork, written by someone who went one way and observed the (Ch. 40)

  • Anything on the "outbox vs. CDC vs. dual write" decision written by someone who has run all three. (Ch. 36)

  • Anything on the base-rate problem in screening. §31.2's false-positive argument is an instance of (Ch. 31)

  • Anything on the failure rate of decentralization initiatives generally — microservices adoption is (Ch. 35)

  • Anything on the history of OLAP semantic layers — Business Objects universes, Cognos frameworks, (Ch. 30)

  • Anything on three-valued logic and NULL semantics. The NOT IN failure (§39.6) is one instance of (Ch. 39)

  • Anything on time-series cross-validation — forward-chaining, walk-forward validation, purged (Ch. 32)

  • Anything on Zipf and power-law distributions in operational data. §21.7 draws a line between a (Ch. 21)

  • Anything sober on organizational change and technical projects that produce no features. The (Ch. 37)

  • Anything you can find on semantic drift detection. §17.9's first residual — a field's meaning (Ch. 17)

  • Apache Pulsar's documentation on its segment-based architecture, read as a comparison. Pulsar (Ch. 15)

  • Ask "what is this for, and is that still true?" of five things you maintain. Exercise 37.14. Two (Ch. 37)

  • Ask for an instance. Exercise 40.5. One meeting, four minutes, and it converts a criterion into a (Ch. 40)

  • Ask the Case Study 1 question of one green indicator today. What number was thresholded to (Ch. 25)*

  • Ask the question that started Case Study 1. "If a customer asks us to delete everything, how long (Ch. 31)*

  • Ask, of every compute resource you run: what fraction of its billed time is it working? §33.7. (Ch. 33)

  • Attempt a full rebuild. Exercise 34.12. If it succeeds on the first try, check that you rebuilt (Ch. 34)

  • Audit your mutes today. Exercise 23.16. If your alerting tool cannot tell you how many (Ch. 23)

  • AWS S3 documentation on request rates and performance. The primary source for the small-files (Ch. 2)

  • AWS S3 documentation: "Best practices design patterns: optimizing Amazon S3 performance," and (Ch. 9)

  • AWS S3 Lifecycle configuration documentation, and the equivalent for GCS and ADLS. §30.7's missing (Ch. 30)

  • Baron Schwartz, Peter Zaitsev, and Vadim Tkachenko, High Performance MySQL (O'Reilly). The (Ch. 7)

  • Barr Moses, Lior Gavish, and Molly Vorwerck, Data Quality Fundamentals (O'Reilly, 2022). The (Ch. 23, Ch. 25)

  • Bas Harenslak and Julian de Ruiter, Data Pipelines with Apache Airflow (2nd ed., Manning, (Ch. 24)

  • Benn Stancil's newsletter. The most consistently thoughtful writing about the data tooling (Ch. 5)

  • Benn Stancil's writing on metrics layers and the semantic layer problem. The clearest (Ch. 2)

  • Benoit Dageville et al., "The Snowflake Elastic Data Warehouse" (2016), SIGMOD. How separated (Ch. 8)

  • Betsy Beyer et al., The Site Reliability Workbook (O'Reilly, 2018). Free online, and the (Ch. 26)

  • BigQuery's dry-run documentation and Snowflake's EXPLAIN. §33.6's pre-flight estimate. (Ch. 33)

  • BigQuery: "Introduction to partitioned tables," "Introduction to clustered tables," and (Ch. 8)

  • Bill Inmon and Ralph Kimball's original disagreement, best encountered through Kimball's The (Ch. 3)*

  • Bill Inmon on the operational data store and the atomic warehouse layer. The other half of the (Ch. 34)

  • Bill Inmon, Building the Data Warehouse, 4th edition (Wiley, 2005). The other side of the (Ch. 6)

  • Break your CI on purpose. Exercise 27.14. Point state:modified+ at a missing manifest and watch (Ch. 27)

  • Brendan Gregg, Systems Performance (2nd ed., Addison-Wesley, 2020), on the USE method. Not a (Ch. 18)

  • Bruce Momjian's presentations (momjian.us/presentations). Free, extensive, and unusually (Ch. 7)

  • Build an SCD2 dimension by hand once, without a snapshot macro. Exercise 20.15. Snapshot tools (Ch. 20)

  • Burrow (LinkedIn's consumer lag monitoring tool) and its design rationale. Burrow's argument — (Ch. 15)

  • Camille Fournier, The Manager's Path (O'Reilly, 2017). Nominally about management, and the (Ch. 40)

  • Case Study 1 and Chapter 33 §33.12 of this book. Delivering a finding about somebody else's (Ch. 37)

  • Cassie Kozyrkov's writing on decision-making and metrics. Not a data engineering source, and (Ch. 6)

  • CCPA/CPRA, and the California Privacy Protection Agency's regulations. Different structure, same (Ch. 31)

  • Chad Sanderson's writing on data contracts, and the surrounding discourse of the mid-2020s. He (Ch. 17)

  • Chapter 1 §1.7 — the acceptance criterion, and the frozen anchors that §38.7's verification (Ch. 38)

  • Chapter 1's Case Study 1 of this book, and the writing on metric layers referenced in Chapter (Ch. 6)

  • Chapter 11 §11.6 of this book. The honesty checklist for a format benchmark applies unchanged to (Ch. 22)

  • Chapter 17 of this book, on data contracts. Case Study 2 is a schema change with no schema, from (Ch. 22)

  • Chapter 20 of this book. Not a deflection: a Type 2 slowly changing dimension is a (Ch. 32)

  • Chapter 20 — incremental models and the deterministic tie-break Case Study 2's first defect (Ch. 38)

  • Chapter 23 and Chapter 26 of this book. Not a deflection. The register of assertions is the most (Ch. 30)

  • Chapter 23 §23.11 of this book, and Case Study 1. Reconciliation independence is the (Ch. 36)

  • Chapter 23 — the assertion register, and §38.11's honest reading of what it caught. (Ch. 38)

  • Chapter 25 and Chapter 26 of this book. Case Study 2's finding is that the entire monitoring stack (Ch. 33)

  • Chapter 25 Case Study 1 and Chapter 26 §26.4 of this book. Case Study 2's review is the same move (Ch. 29)

  • Chapter 25 of this book. Feature freshness is freshness; materialization lag is pipeline lag. The (Ch. 32)

  • Chapter 25 §25.12, Chapter 30 Case Study 2, and Chapter 33 Case Study 1 of this book. Measuring (Ch. 39)

  • Chapter 26 of this book. On-call design, and the observation that a rota that pages more than twice (Ch. 40)

  • Chapter 26 — the SLO and the slack that §38.10 publishes. (Ch. 38)

  • Chapter 27 §27.9 and Chapter 32 §32.7 of this book. Shadow running as a deployment technique, and (Ch. 37)

  • Chapter 27 — the CI that runs the models, and where purity.py belongs. (Ch. 38)

  • Chapter 3's further reading on the lakehouse papers. Where warehouses are going: the Armbrust (Ch. 8)

  • Chapter 30 Case Study 2 and Chapter 25 §25.12 of this book. Query-log analysis and reviewing (Ch. 37)

  • Chapter 30 Case Study 2 of this book. The access review that could not say no. §31.6's policies (Ch. 31)

  • Chapter 30 §30.9 and Chapter 33 §33.12 of this book. Instrumenting the asking, and delivering a (Ch. 35)

  • Chapter 30 — the catalog, without which §38.12's second question needs a person. (Ch. 38)

  • Chapter 31 §31.5 of this book. The Iceberg migration that made row-level deletes possible was (Ch. 34, Ch. 36)

  • Chapter 31 — the deletion manifest, without which the platform is not finished. (Ch. 38)

  • Chapter 33 — the rate card §38.10 prices against, and the slack alert. (Ch. 38)

  • Chapter 34 §34.9 and Chapter 36 §36.10 of this book. The rebuild and the replay. §38.9 is the (Ch. 38)

  • Chapter 34 §34.9 — the rebuild. (Ch. 38)

  • Chapter 36 Case Study 1 and Chapter 38 Case Study 1. Reconciliation independence, and four failed (Ch. 39)

  • Chapter 36 Case Study 1 — reconciliation independence, which is §38.7's and §38.8's whole argument. (Ch. 38)

  • Chapter 37 §37.7 — a difference is usually a rule nobody wrote down, which is Case Study 1 four (Ch. 38)

  • Chapter 4 of this book, Case Study 2, and this chapter's Case Study 1. They are two halves of (Ch. 7)

  • Chapter 40's durable list, read as a bibliography. Kimball (1996), Lamport (1978), and — for (Ch. 40)

  • Chapter 9 of this book, on partitioning and file size. Not a deflection: partition pruning is (Ch. 33)

  • Chapter 9 §9.6 of this book, and its Case Study 1. The small-files problem in its (Ch. 10)

  • Chapters 17, 30, and 34 of this book. Contracts, catalog, ownership, and computationally-enforced (Ch. 35)

  • Chapters 23, 30, 34, 36, and 38 of this book. §40.13's claim is that the durable skill is knowing (Ch. 40)

  • Chapters 3, 4, 11, 29, and 34 of this book. Architecture principles, distributed systems, the (Ch. 39)

  • Charity Majors' writing on operational ownership and on-call. Blunt, experienced, and the best (Ch. 5)

  • Charity Majors, Liz Fong-Jones, and George Miranda, Observability Engineering (O'Reilly, 2022). (Ch. 25)

  • Chip Huyen, Designing Machine Learning Systems (O'Reilly, 2022). The best single book for this (Ch. 32)

  • Chris Richardson's microservices.io patterns, specifically Transactional Outbox, Polling (Ch. 36)

  • Christina Maslach's work on burnout, which is the actual research rather than the folklore. The (Ch. 40)

  • Cindy Sridharan, Distributed Systems Observability (O'Reilly, 2018). Short, free, and the (Ch. 25)

  • Clear a historical task in a local instance, deliberately. Exercise 24.13. It takes five minutes (Ch. 24)

  • Compute depth and blast radius for your own project, and compare against where your assertions (Ch. 34)

  • Compute feature age. Exercise 32.6. Nobody has, and the ratio is usually surprising. (Ch. 32)

  • Compute your unattributed share. Exercise 33.5. Report it honestly, including if it is (Ch. 33)

  • Confluent's and AWS MSK's documentation on monitoring. Both publish the metrics that matter — (Ch. 15)

  • Corey Quinn's Last Week in AWS newsletter and writing. Opinionated, frequently funny, and more (Ch. 33)

  • Count last month's alerts and how many were acted on. Exercise 25.16. If you cannot determine (Ch. 25)

  • Cube's documentation on data modeling. A different take on the same problem, and worth reading (Ch. 30)

  • Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy (2014). The (Ch. 31)

  • Damien Desfontaines' blog (desfontain.es), the "differential privacy" series. The clearest (Ch. 31)

  • Dan Linstedt on Data Vault, particularly the raw-vault / business-vault split. This is the same (Ch. 34)

  • Dan McKinley, "Choose Boring Technology" (2015). The innovation-tokens essay. The argument is (Ch. 5)

  • Dan McKinley, "Choose Boring Technology." Recommended in Chapter 5 and directly applicable (Ch. 12)

  • Daniel Abadi, "Consistency Tradeoffs in Modern Distributed Database System Design" (2012), IEEE (Ch. 4)

  • Daniel Abadi, Samuel Madden, and Nabil Hachem, "Column-Stores vs. Row-Stores: How Different Are (Ch. 8)

  • Daniel Linstedt and Michael Olschimke, Building a Scalable Data Warehouse with Data Vault (Ch. 6)*

  • Daniele Procida's "Diátaxis" framework for technical documentation. A runbook is what Diátaxis (Ch. 26)

  • Data observability vendors. Chapter 23 §23.8's assessment applies unchanged: a net for the (Ch. 25)

  • Databricks Feature Store and SageMaker Feature Store documentation. The managed offerings. Read (Ch. 32)

  • DataHub, Amundsen, and OpenMetadata — the open-source catalogs. All three are real, all three are (Ch. 30)

  • David Goldberg, "What Every Computer Scientist Should Know About Floating-Point Arithmetic" (Ch. 7)

  • David Salomon, Data Compression: The Complete Reference. If you want to understand why (Ch. 11)

  • DBT Labs, "What is analytics engineering?" and the accompanying discourse. Vendor material, (Ch. 1)

  • Dehghani's two original articles, "How to Move Beyond a Monolithic Data Lake to a Distributed Data (Ch. 35)

  • Delta Lake documentation: "Table utility commands" (OPTIMIZE, VACUUM, RESTORE, (Ch. 10)

  • Delta Lake's "Best practices" page and the equivalent Databricks optimization guidance. Read (Ch. 10)

  • Do the twelve drills against your own engine. Exercise 39.4, and check your window-frame default. (Ch. 39)

  • Douglas Terry et al., "Session Guarantees for Weakly Consistent Replicated Data" (1994). Where (Ch. 4)

  • Each vendor's usage and metering views: Snowflake's SNOWFLAKE.ACCOUNT_USAGE schema (Ch. 8)

  • Eric Brewer, "CAP Twelve Years Later: How the 'Rules' Have Changed" (2012), IEEE Computer. (Ch. 4)

  • Eric Evans, Domain-Driven Design (Addison-Wesley, 2003), on aggregates and domain events. An (Ch. 36)

  • Eric Evans, Domain-Driven Design, on bounded contexts and context maps. A context map is (Ch. 35)

  • Estimate three queries before running them. Exercise 33.4. Most people are wrong by more than (Ch. 33)

  • Fabian Reinartz et al. on the Prometheus TSDB design, and the Gorilla paper: Tuan Pham et al., (Ch. 12)

  • Fay Chang et al., "Bigtable: A Distributed Storage System for Structured Data" (2006), OSDI. (Ch. 1, Ch. 12)

  • Find your shadow pipelines. Exercise 35.10. Service accounts with unexplained query patterns, (Ch. 35)

  • For each track in §40.5, this book's own chapters: streaming (29), platform (24, 27, 28), analytics (Ch. 40)

  • Fred Brooks, The Mythical Man-Month (1975, anniversary edition 1995). Fifty years old and mostly (Ch. 40)

  • Fred Brooks, The Mythical Man-Month, on the second-system effect. Directly relevant and usually (Ch. 37)

  • GDPR Articles 12 and 17, and your jurisdiction's equivalent. Article 17 is the right to (Ch. 9)

  • GDPR, Articles 5, 6, 12–22, 25, 32, and 33–34. That is the engineering subset: principles, lawful (Ch. 31)

  • Giuseppe DeCandia et al., "Dynamo: Amazon's Highly Available Key-value Store" (2007), SOSP. (Ch. 4, Ch. 12)

  • Goodhart's law, and Marilyn Strathern's formulation of it"when a measure becomes a target, it (Ch. 26)*

  • Google Cloud DLP / Sensitive Data Protection, and AWS Macie — the managed equivalents. Worth (Ch. 31)

  • Google Cloud, "Data lifecycle" and the equivalent architecture-center pages at AWS and Azure. (Ch. 2)

  • Google SRE Team, Site Reliability Engineering (free at sre.google/books), the chapter on (Ch. 5)

  • Google SRE Team, Site Reliability Engineering (O'Reilly, 2016), free online at sre.google. (Ch. 1)

  • Google SRE Team, Site Reliability Engineering, free at sre.google/books. The DataOps (Ch. 2)

  • Google's "Data Validation for Machine Learning" (Breck et al., SysML 2019) and TensorFlow Data (Ch. 32)

  • Google's Site Reliability Engineering (O'Reilly, 2016), and The Site Reliability Workbook (Ch. 25)

  • Google's Site Reliability Engineering (O'Reilly, 2016), on monitoring and alerting. The (Ch. 23)

  • Google's Site Reliability Engineering (O'Reilly, 2016), the chapters on monitoring and on (Ch. 19)

  • Google's Site Reliability Engineering (O'Reilly, 2016). Free online. Read Chapter 11, "Being (Ch. 26)

  • Google's Site Reliability Engineering, on alerting and on error budgets. Case Study 2's (Ch. 20)

  • Google's Site Reliability Engineering, on monitoring and on "the four golden signals." Case (Ch. 24)

  • Great Expectations, and Chapter 23's assessment of it. Its place in CI is validating inputs (Ch. 27)

  • Greg Young's talks and writing on event sourcing. The most experienced practitioner voice, and (Ch. 36)

  • Greg Young, Versioning in an Event Sourced System (Leanpub). The definitive treatment of §36.10, (Ch. 36)

  • Gregor Hohpe and Bobby Woolf, Enterprise Integration Patterns (Addison-Wesley, 2003). Old, (Ch. 13)

  • Gwen Shapira, Todd Palino, Rajini Sivaram, and Krit Petty, Kafka: The Definitive Guide, (Ch. 15)

  • Hand your README to someone. Exercise 38.9. Watch, do not help, and write down every point at which (Ch. 38)

  • Hironobu Suzuki, The Internals of PostgreSQL (interdb.jp/pg). A free online book on the (Ch. 7)

  • Holden Karau and Rachel Warren, High Performance Spark (2nd ed., O'Reilly, 2023). The book (Ch. 21)

  • Ian Robinson and the consumer-driven contracts pattern (2006), and Martin Fowler's write-up. (Ch. 17)

  • IBM's documentation on COBOL copybooks and packed decimal, if you meet a mainframe. Deeply (Ch. 13)

  • Iceberg's "Maintenance" documentation — expiring snapshots, removing orphan files, rewriting (Ch. 10)

  • Interview a company you are not going to join. Exercise 39.13. Uncomfortable, and it calibrates the (Ch. 39)

  • Interview someone. The fastest way to learn what the rubric is measuring is to sit on the other (Ch. 39)

  • Itzik Ben-Gan, T-SQL Window Functions: For Data Analysis and Beyond (2nd ed., Microsoft Press, (Ch. 18)

  • J.R. Storment and Mike Fuller, Cloud FinOps (O'Reilly, 2nd ed. 2023). The standard text. The (Ch. 33)

  • Jake VanderPlas, Python Data Science Handbook (2nd ed., O'Reilly, 2022). Free online. The (Ch. 22)

  • Jay Kreps, "Questioning the Lambda Architecture" (2014). The essay that named and then argued (Ch. 3, Ch. 29)

  • Jay Kreps, "The Log: What every software engineer should know about real-time data's unifying (Ch. 1, Ch. 14, Ch. 15, Ch. 29, Ch. 36)

  • Jerry Muller, The Tyranny of Metrics (Princeton, 2018). Case Study 1's problem — a measure that (Ch. 26)

  • Jez Humble and David Farley, Continuous Delivery (Addison-Wesley, 2010). Old, foundational, and (Ch. 27)

  • Jim Gray and Andreas Reuter, Transaction Processing: Concepts and Techniques (1992), on (Ch. 4)

  • Job postings, read as market data rather than as opportunities. Count how many roles in your market (Ch. 40)

  • Joe Celko, SQL for Smarties: Advanced SQL Programming (5th ed., Morgan Kaufmann, 2014). The (Ch. 18)

  • Joe Reis and Matt Housley, Fundamentals of Data Engineering (O'Reilly, 2022), Chapter 2. (Ch. 2)

  • Joe Reis and Matt Housley, Fundamentals of Data Engineering (O'Reilly, 2022), on (Ch. 24)

  • Joe Reis and Matt Housley, Fundamentals of Data Engineering (O'Reilly, 2022). The book that (Ch. 1, Ch. 19)

  • Joe Reis and Matt Housley, Fundamentals of Data Engineering, Chapter 8 ("Queries, Modeling, and (Ch. 6)

  • Joel Spolsky, "In Defense of Not-Invented-Here Syndrome" (2001). Twenty-plus years old and the (Ch. 5)

  • John Allspaw and the "blameless postmortem" literature. Case Study 1's known-issues entry was (Ch. 19)

  • John Allspaw on incident analysis, and the resilience-engineering literature generally. Case (Ch. 25)

  • John Allspaw's writing on incident analysis, particularly on how organizations converge on a (Ch. 18)

  • John D. C. Little's original 1961 paper, "A Proof for the Queuing Formula $L = \lambda W$", (Ch. 3)

  • Jordan Tigani, "Big Data Is Dead" (2023), MotherDuck blog. Argues, with data from BigQuery (Ch. 3, Ch. 5)

  • Jordan Tigani, "Big Data Is Dead" (2023). Recommended in Chapter 5 and relevant again: the (Ch. 8)

  • Jules Damji, Brooke Wenig, Tathagata Das, and Denny Lee, Learning Spark (2nd ed., O'Reilly, (Ch. 21)

  • Julien Le Dem and Nong Li, "Parquet: Columnar Storage for Hadoop" (2013) and the associated (Ch. 9)

  • Kapoor and Narayanan, "Leakage and the Reproducibility Crisis in ML-based Science" (2023). (Ch. 32)

  • Kaufman, Rosset, and Perlich, "Leakage in Data Mining" (KDD 2011 / TKDD 2012). The formal treatment, (Ch. 32)

  • Kent Beck and others on "make it work, make it right, make it fast." Relevant backwards: the (Ch. 38)

  • Kief Morris, Infrastructure as Code (2nd ed., O'Reilly, 2020). The standard reference and the (Ch. 28)

  • Kimball Group Design Tips, the archived newsletter. Short, specific, and several of them address (Ch. 20)

  • Kimball's original articles on SCD types, collected in the Reader and widely summarized (Ch. 6)

  • KIP-345 (static membership) and KIP-429 (incremental cooperative rebalancing). Kafka (Ch. 15)

  • KIP-98 (exactly-once delivery and transactional messaging). The design document behind the (Ch. 15)

  • Kyle Kingsbury's Jepsen reports (jepsen.io). Empirical testing of distributed databases' (Ch. 4, Ch. 13)

  • Lamport, "Time, Clocks, and the Ordering of Events in a Distributed System" (1978). The foundation. (Ch. 36)

  • Latanya Sweeney, "k-Anonymity: A Model for Protecting Privacy" (2002), and her earlier work (Ch. 31)

  • Laura Sebastian-Coleman, Measuring Data Quality for Ongoing Improvement (Morgan Kaufmann, (Ch. 23)

  • Lauren Balik and the critical dbt commentary of the mid-2020s. There is a genuine dissenting (Ch. 19)

  • Leslie Lamport, "Time, Clocks, and the Ordering of Events in a Distributed System" (1978), (Ch. 4)

  • Look up one stateful resource's ForceNew attributes and check whether any of them is something (Ch. 28)

  • Machanavajjhala et al., "l-Diversity: Privacy Beyond k-Anonymity" (2006). The homogeneity attack — (Ch. 31)

  • Marc Brooker's blog (brooker.co.za/blog) and the AWS Builders' Library (Ch. 4)

  • Marc Brooker's personal blog (brooker.co.za/blog). The long-form version of the same (Ch. 16)

  • Marc Brooker, "Timeouts, retries, and backoff with jitter" (AWS Builders' Library). The primary (Ch. 16)

  • Mark Raasveldt and Hannes Mühleisen's papers on DuckDB, particularly "DuckDB: an Embeddable (Ch. 22)*

  • Mark Raasveldt et al., "Fair Benchmarking Considered Difficult: Common Pitfalls In Database (Ch. 8, Ch. 11)

  • Markus Winand, SQL Performance Explained (2012), and the companion site (Ch. 18)

  • Markus Winand, SQL Performance Explained and the companion site use-the-index-luke.com. The (Ch. 7)

  • Martin Fowler on "Polyglot Persistence." The essay that named the idea that different parts of (Ch. 12)

  • Martin Fowler on "the two hard things" and, more usefully, anything careful on ubiquitous (Ch. 34)

  • Martin Fowler on ParallelChange (also called expand-contract). The refactoring pattern (Ch. 17)

  • Martin Fowler on the Strangler Fig Application (2004). §37.5, from the person who named it. Short, (Ch. 37)

  • Martin Fowler's writing on "expand and contract" / parallel change. Chapter 17's compatibility (Ch. 27)

  • Martin Fowler's writing on the Outbox pattern and on dual writes. Directly relevant to §13.5's (Ch. 13)

  • Martin Fowler, "Event Sourcing" (2005) and "CQRS" (2011). The canonical write-ups. The event (Ch. 36)

  • Martin Fowler, "Utility vs Strategic Dichotomy" and the related writing on technical (Ch. 5)

  • Martin Fowler, "What do you mean by 'Event-Driven'?" (2017). Short, free, and it is §36.2 — the (Ch. 36)

  • Martin Fowler, "Who Needs an Architect?" (2003), IEEE Software. The source of the (Ch. 3)

  • Martin Kleppmann, "Turning the database inside-out with Apache Samza" (2015), and the (Ch. 14)

  • Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Chapter 10. The (Ch. 21)

  • Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Chapter 11. Streams (Ch. 29, Ch. 36)

  • Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Chapters 7 and 11. (Ch. 20)

  • Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017). The best (Ch. 1, Ch. 4, Ch. 39)

  • Martin Kleppmann, Designing Data-Intensive Applications, Chapter 11 ("Stream Processing"), (Ch. 14, Ch. 15)

  • Martin Kleppmann, Designing Data-Intensive Applications, Chapter 2 ("Data Models and Query (Ch. 12)

  • Martin Kleppmann, Designing Data-Intensive Applications, Chapter 7 ("Transactions"). The (Ch. 2)

  • Matt Harrison, Effective Pandas (2021, and a second edition). Idiomatic pandas: method (Ch. 22)

  • Maxime Beauchemin, "Functional Data Engineering — a modern paradigm for batch data (Ch. 2)

  • Maxime Beauchemin, "Functional Data Engineering — a modern paradigm for batch data processing" (Ch. 13)

  • Maxime Beauchemin, "Functional Data Engineering" (2018). Recommended in Chapter 2 and again (Ch. 6)

  • Maxime Beauchemin, "The Rise of the Data Engineer" (2017) and "The Downfall of the Data (Ch. 1)

  • Measure k on something you actually ship. Exercise 31.6. If your organization sends any extract (Ch. 31)

  • Measure your bottleneck. Exercise 35.4. Split by wait time, not volume, and look at who could (Ch. 35)

  • Measure your own processing_time − event_time distribution. Exercise 29.13. Everyone picks 30 (Ch. 29)

  • Melanie Mitchell's and others' writing on institutional metrics and Goodhart's law, applied to (Ch. 30)

  • Michael Armbrust et al., "Delta Lake: High-Performance ACID Table Storage over Cloud Object (Ch. 3, Ch. 10)

  • Michael Armbrust et al., "Lakehouse: A New Generation of Open Platforms that Unify Data (Ch. 3, Ch. 10)

  • Michael Feathers, Working Effectively with Legacy Code (Prentice Hall, 2004). Twenty years old (Ch. 37)

  • Michael Nygard, "Documenting Architecture Decisions" (2011). The essay that introduced the ADR (Ch. 3)

  • Michael Nygard, Release It! (2nd ed., Pragmatic Bookshelf, 2018). On deploying things that (Ch. 27)

  • Microsoft Presidio — open-source PII detection and anonymization, and a substantially more (Ch. 31)

  • Mike Stonebraker et al., "C-Store: A Column-oriented DBMS" (2005), VLDB. The ancestor of (Ch. 8)

  • Modern SQL's feature tables, at modern-sql.com. Which engines support (Ch. 18)

  • Narayanan and Shmatikov, "Robust De-anonymization of Large Sparse Datasets" (2008) — the Netflix (Ch. 31)

  • Nathan Marz and James Warren, Big Data (Manning, 2015) — the Lambda architecture, from its (Ch. 29)

  • Nathen Harvey and the PagerDuty incident response documentation. PagerDuty publishes its internal (Ch. 26)

  • Neha Narkhede, "Exactly-once Semantics are Possible: Here's How Kafka Does it" (2017), Confluent (Ch. 4)

  • Neha Narkhede, "Exactly-once Semantics are Possible: Here's How Kafka Does it" (2017). (Ch. 15)

  • Neil Gunther, Guerrilla Capacity Planning (Springer, 2007). Capacity planning for people who (Ch. 3)

  • Nickolas Means and the "how they built it" genre, plus Google's Site Reliability Engineering (Ch. 28)

  • Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate (IT Revolution, 2018). The evidence base: (Ch. 27)

  • NIST SP 800-53, control family AC (Access Control), and AC-2(3) specifically. Dry, and the (Ch. 30)

  • Northcutt, Athalye, and Mueller, "Pervasive Label Errors in Test Sets" (2021). Found substantial (Ch. 32)

  • OpenCost and Kubecost, if you run on Kubernetes. Allocation for shared clusters is the hardest (Ch. 33)

  • Pact and the consumer-driven contract testing literature. A more formal approach to §16.9's (Ch. 16)

  • Parts IV and V of this book. Chapter 20 (idempotency, SCD2, the deterministic tie-break), Chapter (Ch. 39)

  • Piethein Strengholt, Data Management at Scale (O'Reilly, 2nd ed. 2023). The most practical book (Ch. 30)

  • PostgreSQL documentation, "Concurrency Control" (Chapter 13 of the manual). The section on (Ch. 2)

  • PostgreSQL documentation, "High Availability, Load Balancing, and Replication" (Chapter 26 of (Ch. 4)

  • PostgreSQL documentation, postgresql.org/docs. Genuinely excellent, unusually well-written (Ch. 1)

  • Postmortem practice — Allspaw, Dekker, and the resilience-engineering literature. Both case (Ch. 23)

  • Practice sites with a data slant — anything offering realistic multi-table problems rather than (Ch. 39)

  • Pramod Sadalage and Martin Fowler, NoSQL Distilled (Addison-Wesley, 2012). Short, and its (Ch. 12)

  • Prometheus + Grafana for metrics, Loki / Elasticsearch / a cloud log service for structured (Ch. 25)

  • Public engineering blogs and incident write-ups from the company you are interviewing with. A (Ch. 39)

  • Public engineering write-ups on cost-per-unit at scale — Dropbox's storage migration, Netflix's (Ch. 33)

  • Publicly published engineering ladders — several companies have opened theirs, and (Ch. 40)

  • Ralph Kimball and Joe Caserta, The Data Warehouse ETL Toolkit (Wiley, 2004). The companion to (Ch. 13)

  • Ralph Kimball and Margy Ross, The Data Warehouse Toolkit (3rd ed., Wiley, 2013), Chapter 5. (Ch. 20)

  • Ralph Kimball and Margy Ross, The Data Warehouse Toolkit (3rd ed., Wiley, 2013), on error event (Ch. 23)

  • Ralph Kimball and Margy Ross, The Data Warehouse Toolkit (3rd ed., Wiley, 2013). Referenced in (Ch. 19)

  • Ralph Kimball and Margy Ross, The Data Warehouse Toolkit, 3rd edition (Wiley, 2013). The (Ch. 1)

  • Ralph Kimball and Margy Ross, The Data Warehouse Toolkit, 3rd edition, Chapter 1. The (Ch. 2)

  • Ralph Kimball and Margy Ross, The Data Warehouse Toolkit: The Definitive Guide to Dimensional (Ch. 6)*

  • Ralph Kimball and Margy Ross, The Kimball Group Reader, 2nd edition (Wiley, 2016). Collected (Ch. 6)

  • Ralph Kimball, The Data Warehouse Toolkit (Wiley, 3rd ed. 2013), the chapters on the back room and (Ch. 34)

  • Read sys/fs/cgroup/memory.peak for one job you own, today. Compare it to the limit. Case Study (Ch. 22)

  • Read one production plan a week. Not to fix anything — just to build the habit of counting (Ch. 21)

  • Rebuild and diff row by row. Exercise 38.5. A rebuild that succeeds first time usually means you (Ch. 38)

  • Reconcile something real. Exercise 38.12. Write down every rule before you run the comparison, (Ch. 38)

  • Redshift: "Distribution styles" and "Sort keys." The two settings that determine whether (Ch. 8)

  • Rehearse a rollback. Exercise 37.9. Time it, and compare against what the plan claims. (Ch. 37)

  • Replay something. Exercise 36.10. It will fail, and the failure is the point. (Ch. 36)

  • Rewrite one production pandas job in DuckDB. Not a large one. The exercise is worth doing for (Ch. 22)

  • RFC 4180, "Common Format and MIME Type for Comma-Separated Values Files" (2005). Two pages, (Ch. 11)

  • RFC 5988 / RFC 8288, "Web Linking." The Link header format, for §16.2's fourth pagination (Ch. 16)

  • RFC 6585 §4 (status 429) and the IETF RateLimit header fields draft. The 429 definition, and (Ch. 16)

  • RFC 6749 (OAuth 2.0) and RFC 6750 (Bearer Token Usage). Read §4.4 (client credentials) and §6 (Ch. 16)

  • RFC 8594, the Sunset HTTP header, and the Deprecation header draft. The mechanism by which (Ch. 16)

  • RFC 9110 §10.2.3 on Retry-After. Two paragraphs, and it specifies both accepted forms — (Ch. 16)

  • RFC 9110, "HTTP Semantics." The current HTTP specification, superseding RFC 7231. Read the (Ch. 16)

  • Richard Cook, "How Complex Systems Fail" (1998). Recommended in Chapter 1 and recommended (Ch. 4)

  • Richard Cook, "How Complex Systems Fail" (1998/2000). Eighteen numbered propositions, four (Ch. 1)

  • Rob Ewaschuk, "My Philosophy on Alerting." An internal Google document, widely circulated, (Ch. 25)

  • Roy Fielding and the REST/hypermedia literature on evolvability, particularly the argument that (Ch. 17)

  • Run pr_report.py against your last ten merged changes and count the definitional ones that were (Ch. 27)

  • Run terraform plan -refresh-only today. Exercise 28.15. Write down your prediction first, (Ch. 28)

  • Run a mock loop. Exercise 39.11. With another person, scored against the rubric, with feedback (Ch. 39)

  • Run one runbook drill this quarter. Exercise 26.18. One runbook, one colleague who did not write (Ch. 26)

  • Run the grep. Case Study 1's is four minutes and finds something in almost every organization. If (Ch. 30)

  • Run the row-level diff on any feature computed in two places. Exercise 32.7. An afternoon, and the (Ch. 32)

  • Run the time-split check on a model you have access to. Exercise 32.12. If the gap is small, (Ch. 32)

  • Run the two queries. Exercise 30.7. Granted versus used, against a warehouse you have access to. (Ch. 30)

  • Run §29.11's four questions against the next "we need real time" request you receive. Case Study (Ch. 29)

  • Ryan Blue and Daniel Weeks' Iceberg papers and talks (Netflix). Iceberg's design rationale, (Ch. 10)

  • Sam Newman, Monolith to Microservices (O'Reilly, 2019). The most practical modern treatment of (Ch. 37)

  • Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung, "The Google File System" (2003), SOSP. (Ch. 1)

  • Scan for ambiguous timestamps. Exercise 34.13. Kestrel found 14 of 41 external timestamp columns (Ch. 34)

  • Score your own team. Exercise 40.8. Honestly, including the signals you do not have. (Ch. 40)

  • Search your codebase for dual writes. Exercise 36.4. A database write followed by a publish, an (Ch. 36)

  • Sergey Melnik et al., "Dremel: Interactive Analysis of Web-Scale Datasets" (2010), VLDB, and (Ch. 8)

  • Shirshanka Das et al. on DataHub's architecture (LinkedIn engineering, and the subsequent talks). (Ch. 30)

  • Sidney Dekker, The Field Guide to Understanding 'Human Error' (3rd ed., CRC Press, 2014). The (Ch. 26)

  • Snowflake's ACCESS_HISTORY view in particular. It reports column-level access, which turns (Ch. 30)

  • Snowflake's masking policy and row access policy documentation; the BigQuery column-level security (Ch. 31)

  • Snowflake: "Understanding Snowflake Table Structures" (micro-partitions and clustering), and (Ch. 8)

  • Soda, Monte Carlo, Elementary, and the observability category generally. Worth an evaluation (Ch. 23)

  • Sort your projections by commutativity. Exercise 36.9. An afternoon, and it may save two (Ch. 36)

  • Start the running document. Exercise 40.1. Today, not when you finish the chapter. (Ch. 40)

  • Studies of how developers and analysts actually find things. The consistent finding across (Ch. 30)

  • Take one inventory. Exercise 37.2. Five jobs, eight fields, without reading the code. The (Ch. 37)

  • Take the unwanted job. Exercise 40.11 — the reconciliation, the metric definitions, the finance (Ch. 40)

  • Tanya Reilly, The Staff Engineer's Path (O'Reilly, 2022). The practical companion to Larson. (Ch. 40)

  • The "Best Practices" page in the same documentation. It says most of §24.2 and §24.10, including (Ch. 24)

  • The "birthday problem" in any elementary probability text. The mechanism behind this chapter's (Ch. 13)

  • The "Data Mesh Architecture" community site and the Data Mesh Learning community. Practitioner (Ch. 35)

  • The "data mesh is not for you" genre. Several thoughtful practitioners have written versions of (Ch. 35)

  • The "error budget" material in SRE. §23.1's argument that a reconciliation tolerance is a budget (Ch. 23)

  • The "Falsehoods Programmers Believe About..." genre, particularly the entries on CSV, on names, (Ch. 11)

  • The awesome-data-engineering style lists on GitHub, and the various "data stack landscape" (Ch. 5)

  • The dbt_expectations package. The bridge between the two worlds: Great Expectations' (Ch. 23)

  • The dbt_utils package README, at github.com/dbt-labs/dbt-utils. (Ch. 19)

  • The delta-rs documentation and repository. The Rust implementation with Python bindings, (Ch. 10)

  • The h2oai/db-benchmark project and its successors. A maintained cross-engine benchmark on (Ch. 22)

  • The lzbench benchmark suite. A comparison of dozens of compressors on standard corpora. (Ch. 11)

  • The manifest.json schema documentation. dbt publishes a JSON schema for the artifacts it (Ch. 19)

  • The parquet-format repository's README and LogicalTypes.md on GitHub. More precise than (Ch. 9)

  • The psycopg 3 documentation on server-side cursors. Directly relevant to this chapter's first (Ch. 7)

  • The WITH / WITH RECURSIVE documentation for your engine. PostgreSQL's is the reference (Ch. 18)

  • The Airbyte protocol documentation. A more recent formalization of the same problem, with a (Ch. 13)

  • The Airflow documentation on "Data Interval," "Catchup," and "Backfill." Even before Chapter (Ch. 13)

  • The Airflow documentation on airflow db clean and database maintenance. Case Study 2. It is (Ch. 24)

  • The Airflow documentation on executors, comparing LocalExecutor, CeleryExecutor, and (Ch. 28)

  • The Airflow documentation on metrics (StatsD and OpenTelemetry). §25.13. Four lines of (Ch. 25)

  • The Airflow source, airflow/jobs/scheduler_job_runner.py. An unusual recommendation, and worth (Ch. 24)

  • The Apache Airflow documentation, and specifically these pages read end to end rather than (Ch. 24)

  • The Apache Airflow documentation, especially "Best Practices" and "Concepts." Airflow's docs (Ch. 5)

  • The Apache Airflow, Kafka, Spark, and Iceberg documentation sets. Each is the primary source (Ch. 1)

  • The Apache Arrow columnar format specification. The in-memory counterpart to Parquet's on-disk (Ch. 11)

  • The Apache Arrow documentation, particularly on the Arrow/Parquet relationship. Arrow is the (Ch. 9)

  • The Apache Arrow specification and the Arrow Columnar Format page. §22.5's material. You do (Ch. 22)

  • The Apache Avro specification (avro.apache.org/docs/), particularly "Schema Resolution." (Ch. 11)

  • The Apache Avro specification, "Schema Resolution." Recommended in Chapter 11 and mandatory (Ch. 17)

  • The Apache Beam documentation on watermarks and triggers. Beam's model is the most careful (Ch. 20)

  • The Apache Cassandra documentation on data modelling, particularly "Basic Rules of Cassandra (Ch. 12)

  • The Apache Flink documentation on event time, watermarks, and idleness. §29.5's (Ch. 29)

  • The Apache Hive documentation on partitioning, which is where key=value path partitioning (Ch. 9)

  • The Apache Iceberg and Delta Lake documentation on row-level deletes, deletion vectors, and (Ch. 31)

  • The Apache Iceberg and Delta Lake documentation on time travel and schema evolution. Bronze's (Ch. 34)

  • The Apache Iceberg specification (iceberg.apache.org/spec/). Read alongside the Delta paper. (Ch. 10)

  • The Apache Iceberg specification (iceberg.apache.org/spec/). Worth reading alongside the (Ch. 3)

  • The Apache Kafka documentation (kafka.apache.org/documentation). Read three parts properly: (Ch. 15)

  • The Apache ORC specification (orc.apache.org/specification/). Worth skimming alongside (Ch. 11)

  • The Apache Parquet documentation and format specification (parquet.apache.org/docs/). Read (Ch. 9)

  • The Apache Parquet format specification (parquet.apache.org/docs/file-format/). Short and (Ch. 2, Ch. 11)

  • The Apache Spark documentation on adaptive query execution (AQE) and skew join optimization. (Ch. 4)

  • The Apache XTable project (formerly OneTable). Translates metadata between Delta, Iceberg, and (Ch. 10)

  • The Astronomer documentation and guides. A vendor, and the best free Airflow teaching material (Ch. 24)

  • The AWS announcement of strong read-after-write consistency (December 2020). Worth reading (Ch. 9)

  • The AWS Builders' Library more generally, particularly "Avoiding fallback in distributed (Ch. 16)

  • The AWS Kinesis and Google Pub/Sub documentation, if you are on those platforms. Both solve the (Ch. 15)

  • The AWS S3 documentation on request rate and performance. The reason Case Study 2's per-file (Ch. 21)

  • The AWS, GCP, and Azure documentation on Spot / Preemptible / Spot VMs, particularly the (Ch. 33)

  • The Cassandra CDC documentation, read specifically to understand why most teams do not use it — (Ch. 12)

  • The checklist literature: Atul Gawande, The Checklist Manifesto (Metropolitan Books, 2009). (Ch. 26)

  • The clinical alarm-fatigue literature. Genuinely worth reading, and almost nobody in software (Ch. 25)

  • The Confluent Schema Registry documentation on compatibility modes. BACKWARD, FORWARD, and (Ch. 36)

  • The Confluent Schema Registry documentation on compatibility types. The operational (Ch. 17)

  • The Dagster documentation, particularly on Software-Defined Assets. §24.7's datasets, taken all (Ch. 24)

  • The DAMA-DMBOK (Data Management Body of Knowledge), 2nd edition. Dry, comprehensive, and the (Ch. 9)

  • The DAMA-DMBOK (Data Management Body of Knowledge, 2nd ed.). The reference work, and it is a (Ch. 30)

  • The Databricks documentation on the medallion architecture. Where the bronze/silver/gold naming (Ch. 34)

  • The dbt "Best Practices" guide on CI/CD, and the community writing around slim CI. Read it for (Ch. 27)

  • The dbt "Best Practices" guides, in the same documentation. The staging/intermediate/marts (Ch. 19)

  • The dbt --select and --exclude graph selectors documentation. Not obviously about this (Ch. 34)

  • The dbt community's writing on "one big table" versus star schemas. A genuine, current (Ch. 6)

  • The dbt documentation (docs.getdbt.com). Among the best documentation in this field, and the (Ch. 5)

  • The dbt documentation at docs.getdbt.com. Unusually good, and (Ch. 19)

  • The dbt documentation on docs, meta, and exposures. §30.11's argument in software form: (Ch. 30)

  • The dbt documentation on incremental models, including incremental_strategy, (Ch. 20)

  • The dbt documentation on snapshots and incremental models. §32.9's seven-of-nine answer. Snapshots (Ch. 32)

  • The dbt documentation on snapshots. The check versus timestamp strategy discussion, the (Ch. 20)

  • The dbt documentation on source freshness, and then go and check whether your project runs it. (Ch. 19)

  • The dbt documentation on staging, intermediate, and marts. dbt's naming maps to silver/gold with (Ch. 34)

  • The dbt documentation on state comparison, --defer, and artifacts. The primary source for (Ch. 27)

  • The dbt documentation on tests, severity, error_if/warn_if, and --store-failures. The (Ch. 23)

  • The dbt Learn courses at courses.getdbt.com. Free, official, and (Ch. 19)

  • The dbt Semantic Layer documentation, and the MetricFlow specification. Case Study 1's problem, (Ch. 30)

  • The dbt snapshots documentation (docs.getdbt.com, "Snapshots"). dbt implements Type 2 (Ch. 6)

  • The Debezium blog post on incremental snapshots. The implementation write-up alongside the (Ch. 14)

  • The Debezium documentation on the outbox event router. The CDC-on-the-outbox pattern Kestrel (Ch. 36)

  • The Debezium documentation, "PostgreSQL Connector." Read the section on replication slots and (Ch. 7)

  • The Debezium documentation, especially the PostgreSQL connector page. Read three sections (Ch. 14)

  • The Debezium FAQ on REPLICA IDENTITY. Short, and it is the fix for this chapter's second case (Ch. 14)

  • The Delta Lake and Apache Iceberg documentation on OPTIMIZE / compaction and on ZORDER / (Ch. 9)

  • The Delta Lake documentation on OPTIMIZE and Z-ordering. Case Study 2's actual fix. Chapter 10 (Ch. 21)

  • The Delta Lake UniForm / Iceberg compatibility work, and the various "one format to read them (Ch. 10)

  • The Docker Compose specification (docs.docker.com/compose/compose-file/). The reference for (Ch. 5)

  • The Docker documentation on multi-stage builds and on content-addressable image identifiers. (Ch. 28)

  • The DuckDB and Polars documentation, and Chapter 22. §21.1's argument depends on knowing what a (Ch. 21)

  • The DuckDB blog. Unusually honest performance write-ups, including ones where DuckDB is not (Ch. 8)

  • The DuckDB documentation (duckdb.org/docs). Short, well-written, and the "Guides" section is (Ch. 5)

  • The DuckDB documentation on read_parquet, parquet_metadata, and Hive partitioning. DuckDB (Ch. 9)

  • The DuckDB documentation. Unusually good, unusually short, and the pages worth reading end to (Ch. 22)

  • The EDPB's opinions and guidelines, particularly anything on anonymisation and pseudonymisation. (Ch. 31)

  • The Elasticsearch guide's "Getting Started" and the relevance/scoring chapters, particularly on (Ch. 12)

  • The Feast documentation, particularly get_historical_features and the entity-dataframe concept. (Ch. 32)

  • The FinOps Foundation's framework and its "capabilities" list. Free, vendor-neutral, and useful as (Ch. 33)

  • The FinOps Foundation's materials (finops.org). Vendor-neutral practice for cloud cost (Ch. 8)

  • The Flink TestHarness documentation, and Spark's MemoryStream. The real versions of (Ch. 29)

  • The Flink documentation on state, savepoints, and operator UIDs. §29.7 and §29.10. The operator (Ch. 29)

  • The GDPR text itself, Articles 12 and 17 (gdpr-info.eu or the official EUR-Lex text). (Ch. 2)

  • The GitHub Actions, GitLab CI, or Buildkite documentation for whichever you use — specifically (Ch. 27)

  • The GitHub Engineering "Scientist" library and its write-up (2016). A small library for running old (Ch. 37)

  • The Google Cloud Storage and Azure Blob Storage documentation on consistency and performance. (Ch. 9)

  • The Google SRE book's chapter on access and the "Building Secure and Reliable Systems" companion (Ch. 30)

  • The Google SRE book's chapter on eliminating toil, and its 50% rule. §40.7's first failure mode, (Ch. 40)

  • The Great Expectations documentation. The concepts pages first — Expectations, Suites, (Ch. 23)

  • The IANA time zone database documentation, and any careful treatment of timestamp handling. Case (Ch. 34)

  • The IAPP's practitioner material and the CIPT certification syllabus. The syllabus itself is a (Ch. 31)

  • The ICO's guidance on data retention (UK) and the equivalent from your regulator. Written for (Ch. 30)

  • The ICO's guidance (UK) — the best practitioner-facing writing on this subject anywhere, and it is (Ch. 31)

  • The idempotence and convergence literature from configuration management — Puppet's and Chef's (Ch. 28)

  • The Kafka Connect documentation on offsets, errors.tolerance, and dead-letter queues. (Ch. 14)

  • The Kafka documentation on log compaction. Read it specifically to understand why compaction and (Ch. 36)

  • The Kafka documentation on partitioning and ordering guarantees. Read the exact wording: Kafka (Ch. 36)

  • The Kafka documentation on transactions and exactly-once semantics, and the original KIP-98 (Ch. 29)

  • The Kubernetes documentation on operational overhead, and more usefully, any postmortem (Ch. 28)

  • The Linux kernel documentation on cgroup v2 memory control. Case Study 1's material from the (Ch. 22)

  • The literature on documentation that gets read — Diátaxis is the most useful framework, because it (Ch. 38)

  • The literature on requirements phrasing, or failing that, the discipline of writing acceptance (Ch. 27)

  • The literature on shift work, sleep disruption, and cognitive performance. §26.12's "nights (Ch. 26)

  • The literature on spreadsheet errors — Panko's work is the standard reference, and the reported (Ch. 37)

  • The MinIO documentation on S3 API compatibility. Specifically the compatibility matrix, which (Ch. 5)

  • The MongoDB change streams documentation. The best non-relational change feed in this chapter, (Ch. 12, Ch. 14)

  • The MySQL documentation on the binary log, particularly binlog_format and why ROW is the (Ch. 14)

  • The MySQL Reference Manual, "InnoDB Storage Engine" and "The Binary Log." The clustered-index (Ch. 7)

  • The Neo4j documentation on Cypher, and any comparison of Cypher to recursive SQL. Worth an (Ch. 12)

  • The OAuth 2.1 draft, which consolidates a decade of security guidance. Worth knowing exists; (Ch. 16)

  • The OpenLineage specification and Marquez. An open standard for emitting lineage events from (Ch. 25)

  • The OpenLineage specification, and Marquez as a reference implementation. The vendor-neutral (Ch. 30)

  • The OpenTofu documentation, if that is what your organization uses. The divergence from Terraform (Ch. 28)

  • The Oracle GoldenGate and SQL Server CDC documentation, if you meet them. SQL Server's built-in (Ch. 14)

  • The original "data lake" coinage — James Dixon's 2010 blog post — and the "data swamp" (Ch. 9)

  • The Pact documentation and its "consumer-driven" framing. The service-testing implementation of (Ch. 17)

  • The pandas documentation's "Scaling to large datasets" page, and the Copy-on-Write page. (Ch. 22)

  • The PayPal data contract template, published openly, and the Open Data Contract Standard (Ch. 17)

  • The pgvector repository README. Short, practical, and honest about limits. It documents both (Ch. 12)

  • The Polars user guide, particularly the "Lazy API" and "Expressions" sections. §22.3's (Ch. 22)

  • The PostgreSQL JSONB documentation and the GIN index chapter. The honest comparison point for (Ch. 12)

  • The PostgreSQL pg_replication_slots view documentation. Every column, particularly (Ch. 14)

  • The PostgreSQL pg_stat_replication and pg_last_xact_replay_timestamp documentation. The (Ch. 13)

  • The PostgreSQL documentation (postgresql.org/docs). Read these chapters, in this order, and (Ch. 7)

  • The PostgreSQL documentation on partitioning, parallel query, and BRIN indexes. Read as a (Ch. 5)

  • The PostgreSQL documentation on table sizes and pg_total_relation_size. Two queries from (Ch. 24)

  • The PostgreSQL documentation on transaction isolation and on logical replication slots. The (Ch. 20)

  • The PostgreSQL documentation, "Concurrency Control" (Chapter 13 of the manual). Recommended in (Ch. 13)

  • The PostgreSQL documentation, "Logical Decoding" and "Logical Replication." Recommended in (Ch. 14)

  • The PostgreSQL full-text search chapter, and the pg_trgm extension documentation. The (Ch. 12)

  • The Prefect documentation. Python-first, far less ceremony, and dynamic behaviour that Airflow (Ch. 24)

  • The Prometheus documentation on "Instrumentation" and "Naming", and the Robust Perception blog (Ch. 12)

  • The property-based testing literature — Hypothesis (Python) and QuickCheck. Not usually applied to (Ch. 27)

  • The Protocol Buffers documentation on "Updating A Message Type." A different evolution model — (Ch. 17)

  • The Protocol Buffers language guide and encoding documentation (protobuf.dev). Read the (Ch. 11)

  • The Python csv module documentation, particularly the Dialect class and the Sniffer. The (Ch. 11)

  • The Redpanda documentation, particularly on Kafka API compatibility and on what differs. Worth (Ch. 15)

  • The Singer and Airbyte connector specifications, recommended in Chapter 13 and relevant again: (Ch. 16)

  • The Singer specification (singer.io) and the Meltano project's documentation. Singer defines (Ch. 13)

  • The Slack, Stripe, and GitHub API documentation on pagination. Three well-designed APIs with (Ch. 16)

  • The Spark configuration reference. Worth skimming once so you know what exists. Most of it you (Ch. 21)

  • The Spark SQL Performance Tuning guide, in the official documentation. Short, and the single (Ch. 21)

  • The Spark Structured Streaming Programming Guide, particularly the output modes and watermarking (Ch. 29)

  • The Spark Web UI documentation. Undersold and rarely read. It explains what every column on the (Ch. 21)

  • The Tecton engineering blog, and Uber's Michelangelo papers. Michelangelo is where much of this (Ch. 32)

  • The Terraform documentation on lifecycle, import, moved, and check blocks. Four short (Ch. 28)

  • The TimescaleDB documentation. The strongest argument for §12.10's absorb-first position on (Ch. 12)

  • The TPC-H and TPC-DS specifications (tpc.org). The standard analytical benchmarks. Worth (Ch. 8)

  • The Unicode Consortium's material on encodings, and the "UTF-8 Everywhere" manifesto. Failure 6 (Ch. 11)

  • The US Census Bureau's material on their 2020 disclosure-avoidance system. The largest real (Ch. 31)

  • The VCR / vcrpy / betamax family of libraries. Record-and-replay HTTP for tests. Read the (Ch. 16)

  • The Zstandard documentation and Yann Collet's benchmarks (facebook.github.io/zstd/). The (Ch. 11)

  • There is very little good writing on runbooks specifically, which is worth knowing so you do not (Ch. 26)

  • Tyler Akidau et al., Streaming Systems (O'Reilly, 2018), Chapters 1–3. Event time versus (Ch. 4)

  • Tyler Akidau, "Streaming 101" and "Streaming 102" (2015), on the O'Reilly Radar blog. The (Ch. 3)

  • Tyler Akidau, Slava Chernyak, and Reuven Lax, Streaming Systems (O'Reilly, 2018). The (Ch. 3, Ch. 20, Ch. 29)

  • Vaughn Vernon, Implementing Domain-Driven Design, the chapters on domain events and event (Ch. 36)

  • Vinoth Chandar et al. on Apache Hudi. Hudi's record-level indexing and incremental query model (Ch. 10)

  • Walk the register against one mart. Twenty-two items, an hour, and the output is a list rather (Ch. 23)

  • Warehouse vendors' Iceberg support announcements (Snowflake, BigQuery, Redshift). The (Ch. 10)

  • Wes McKinney, Python for Data Analysis (3rd ed., O'Reilly, 2022). By pandas' author, and the (Ch. 22)

  • Will Larson, Staff Engineer (2021), and An Elegant Puzzle (2019). The best available writing (Ch. 40)

  • Woodrow Hartzog, Privacy's Blueprint (Harvard, 2018). On design as a privacy decision, from the (Ch. 31)

  • Write the "not done" list. Exercise 38.10. If it is empty, you have not looked. (Ch. 38)

  • Write the falsifiable recommendation. Exercise 35.15. "Not yet" is not a recommendation; "not (Ch. 35)*

  • Write the five trade-off sentences. Exercise 39.6 — and notice which technologies you cannot (Ch. 39)

  • Write the five-year letter. Exercise 40.14, with a calendar reminder. (Ch. 40)

  • Write the manifest. Exercise 31.8. Fifteen entries minimum, including logs and vendors. The (Ch. 31)

  • Write your paging list, and show it to somebody who would be woken by it. Exercise 26.15. The (Ch. 26)

  • Yevgeniy Brikman, Terraform: Up & Running (3rd ed., O'Reilly, 2022). The practical companion. (Ch. 28)

  • Your cloud provider's pricing pages, for the three or four services that dominate your bill. (Ch. 33)

  • Your database's documentation on migrating stored procedures. Postgres, SQL Server, and Oracle all (Ch. 37)

  • Your database's documentation on window functions, specifically the frame clause. ROWS versus (Ch. 39)

  • Your engine's documentation on reading a query plan. Spark's EXPLAIN FORMATTED, Snowflake's (Ch. 33)

  • Your engine's documentation on window functions. For PostgreSQL this is the "Window Functions" (Ch. 18)

  • Your linter's custom-rule documentation, whatever it is. SQLFluff for SQL, a dbt package, or forty (Ch. 34)

  • Your own bill, at line-item granularity, exported to somewhere you can query it. AWS Cost and (Ch. 33)

  • Your own company's levelling rubric, read carefully and then set aside. §40.2 and Case Study 2: (Ch. 40)

  • Your own finance team. An hour, and it is worth more than anything on this list. Bring the four (Ch. 38)

  • Your own incident write-ups. §39.9's best source. If you have written postmortems (Chapter 26), (Ch. 39)

  • Your own legal or privacy counsel. Said in §31.6 and worth repeating: an hour with the person (Ch. 31)

  • Your own legal team. §30.6's point is that classification is a legal decision and engineering's job (Ch. 30)

  • Your own ticket system. §35.3's analysis is one query and an afternoon of categorization, and it (Ch. 35)

  • Your provider's budget and anomaly-detection services — AWS Cost Anomaly Detection, GCP budget (Ch. 33)

  • Your provider's documentation on which attributes are ForceNew. This is not a general document (Ch. 28)

  • Your provider's Savings Plan / Committed Use Discount calculator, run twice: once against current (Ch. 33)

  • Your warehouse's ASOF JOIN support. Snowflake, DuckDB, ClickHouse, and Databricks all have one (Ch. 32)

  • Your warehouse's access history, again: Snowflake ACCESS_HISTORY, BigQuery Data Access logs, (Ch. 37)

  • Your warehouse's documentation on time travel, fail-safe, and backup retention. Snowflake's Time (Ch. 31)

  • Your warehouse's own metadata schema. Snowflake ACCOUNT_USAGE (GRANTS_TO_ROLES, (Ch. 30)

  • Your warehouse's usage views. Snowflake ACCOUNT_USAGE.WAREHOUSE_METERING_HISTORY and (Ch. 33)

  • Yu Malkov and Dmitry Yashunin, "Efficient and robust approximate nearest neighbor search using (Ch. 12)

  • Zhamak Dehghani's data mesh writing (covered properly in Chapter 35). Her diagnosis of why (Ch. 9)

  • Zhamak Dehghani, Data Mesh: Delivering Data-Driven Value at Scale (O'Reilly, 2022). The book, by (Ch. 35)


625 distinct sources.