Affiliate disclosure
Book titles on this page link to Amazon. As an Amazon Associate, DataField.Dev earns from qualifying purchases — at no additional cost to you.
Bibliography
Every source recommended anywhere in this book, in one list, with the
chapters that recommend it. This file is generated by
scripts/assemble.py from each chapter's further-reading.md — edit
the chapter file, not this one.
Sources are tagged Tier 1 (we are confident the work exists and recommend it without reservation) or Tier 2 (real and worth seeking, but confirm the current edition, version, or URL yourself — this field's documentation drifts fast).
-
"Hidden Technical Debt in Machine Learning Systems" (Sculley et al., NeurIPS 2015). Nine pages, (Ch. 32)
-
"Monitoring Machine Learning Models in Production" material from the ML-monitoring vendors — Evidently (Ch. 32)
-
"Pagination: You're (probably) doing it wrong" and similar write-ups on keyset versus offset (Ch. 16)
-
"The ML Test Score: A Rubric for ML Production Readiness" (Breck et al., Google, 2017). A checklist (Ch. 32)
-
Advent of Code, solved in SQL. An idiosyncratic but effective (Ch. 18)
-
pgexercises.com. Free, PostgreSQL-based, and the aggregate and (Ch. 18)
-
0.30000000000000004.com. A one-page site listing the output of0.1 + 0.2in dozens of (Ch. 7) -
airflow standalone. A scheduler, web server, and SQLite metadata database in one command and (Ch. 24) -
code/capstone.py. Thirty-eight assertions, including the two that check the result against (Ch. 38) -
code/catalog_audit.pyin this chapter. Eleven findings in a deliberately messy fixture, twenty-two (Ch. 30) -
code/cost_model.pyin this chapter. Three meters, an attributed bill, seven sized wastes, and a (Ch. 33) -
code/coverage.pyin this chapter. Scores a dbt manifest against the six assertions and lists (Ch. 23) -
code/dag_lint.pyin this chapter. Parses DAG source withast— no Airflow required — and (Ch. 24) -
code/engine_benchmark.pyin this chapter. Generates its own data, runs all four engines in (Ch. 22) -
code/event_lab.pyin this chapter. The dual write measured against the outbox, four projections (Ch. 36) -
code/health.pyin this chapter. No dependencies. Feed it a five-column run history and it (Ch. 25) -
code/interview_drills.pyin this chapter. Twelve problems whose wrong answers are enumerated with (Ch. 39) -
code/layer_check.pyin this chapter. Nine rules, a replay planner with costs, depth and blast (Ch. 34) -
code/mesh_readiness.pyin this chapter. Coupling and fan-out, a classified intake queue, weighted (Ch. 35) -
code/migration_lab.pyin this chapter. An estate scored and ordered, sixty days of classified (Ch. 37) -
code/pii_scan.pyin this chapter. Detection scored three ways, a k-anonymity ladder, and a (Ch. 31) -
code/pit_join.pyin this chapter. Leakage measured, an as-of join with the invariant test, feature (Ch. 32) -
code/plan_review.pyin this chapter. Parsesterraform show -jsonand blocks on stateful (Ch. 28) -
code/pr_report.pyin this chapter. Computes blast radius, exposures, deploy shape, rebuild (Ch. 27) -
code/scd_lab.pyin this chapter. Runs every failure here against SQLite in about a second — (Ch. 20) -
code/slo.pyin this chapter. Attainment, error budget, burn rate over a short window, the (Ch. 26) -
code/spark_advisor.pyin this chapter. Reads a plan and counts the four things §21.3 says to (Ch. 21) -
code/stream_harness.pyin this chapter. Watermarks, tumbling and session windows, allowed (Ch. 29) -
dbt-project-evaluator— dbt Labs' own package of project-structure checks. It implements (Ch. 34) -
dbt_artifactsand similar packages that loadrun_results.jsoninto your warehouse. This is (Ch. 25) -
elementaryand the open-source dbt observability tools. They readrun_results.jsonand (Ch. 19) -
EXPLAIN (ANALYZE, BUFFERS)and the PostgreSQL documentation on "Using EXPLAIN." §18.10's (Ch. 18) -
explain.depesz.comandexplain.dalibo.com. Two free tools that take an `EXPLAIN (ANALYZE, (Ch. 7) -
featureform,hopsworks, and the smaller open-source stores. Worth reading the data models of. (Ch. 32) -
pysparkin local mode.SparkSession.builder.master("local[4]")reproduces skew,coalesce(Ch. 21) -
sqlfluff, at sqlfluff.com. A SQL linter with a dbt templater, so it (Ch. 19) -
Accelerate (Forsgren, Humble, Kim) on lead time and deployment frequency as team-level metrics. (Ch. 35)
-
Adrian Cockcroft's writing and talks on evolutionary architecture at Netflix. Useful as a (Ch. 3)
-
Akidau et al., "The Dataflow Model" (VLDB 2015). The paper the book grew out of, free, and about (Ch. 29)
-
Alberto Brandolini's material on Event Storming. A workshop technique for discovering events with (Ch. 36)
-
Alex Xu, System Design Interview (vols. 1–2). Widely used, and written for general software (Ch. 39)
-
And the accounting literature on reconciliation, one more time. It is the oldest idea in this book (Ch. 40)
-
Andreas Andreakis and Ioannis Papapanagiotou, "DBLog: A Watermark Based Change-Data-Capture (Ch. 14)
-
Andrew Jones, Driving Data Quality with Data Contracts (Packt, 2023). A book-length treatment (Ch. 17)
-
Anthony Molinaro and Robert de Graaf, SQL Cookbook (2nd ed., O'Reilly, 2020). Problem-first, (Ch. 18)
-
Anurag Gupta et al., "Amazon Redshift and the Case for Simpler Data Warehouses" (2015), (Ch. 8)
-
Any careful treatment of the small-files problem on object storage. The good sources are usually (Ch. 21)
-
Any careful write-up on embedding model versioning and re-indexing. This is the least (Ch. 12)
-
Any good treatment of confirmation bias in operational data. The specific failure — a small, (Ch. 18)
-
Any honest account of on-call attrition. It is mostly in blog posts and conference talks rather (Ch. 26)
-
Any honest comparison of the three. They are rare; most are written by one of the three. (Ch. 24)
-
Any introductory financial accounting text, on the reconciliation and the trial balance. Two hours, (Ch. 38)
-
Any material on load testing and on chaos engineering, particularly the principle of (Ch. 13)
-
Any of the "why is my Delta/Iceberg table slow" write-ups from practitioners. These are the (Ch. 10)
-
Any postmortem you can find of a replication-slot disk-full incident. They exist, they are (Ch. 14)
-
Any recent benchmark comparing single-node engines to distributed ones — read skeptically. (Ch. 21)
-
Any retrospective on a dead technology written by someone who used it. The Hadoop retrospectives (Ch. 40)
-
Any vendor's "why the data warehouse is dead" or "why the data lake failed" content. Read two (Ch. 3)
-
Any vendor's format comparison. Read two from competing vendors and note what each chose to (Ch. 11)
-
Any write-up of a Kafka rebalance incident. They are numerous, they are consistent, and the (Ch. 15)
-
Anything careful arguing that data mesh is Conway's law with a new name. The critique is fair, it (Ch. 35)
-
Anything careful on "characterization tests." Feathers's term for a test that documents what the (Ch. 37)
-
Anything careful on alert fatigue, from operations rather than from data. The literature is (Ch. 23)
-
Anything careful on clock skew in distributed systems. Kleppmann's Chapter 8 ("The Trouble with (Ch. 20)
-
Anything careful on cloud cost attribution and tagging. The vendor documentation is adequate and (Ch. 28)
-
Anything careful on CSV as an interchange format. RFC 4180 exists, describes what most people (Ch. 22)
-
Anything careful on data masking and synthetic generation. The tooling here is immature and (Ch. 27)
-
Anything careful on idempotent consumers and effectively-once processing. The Kafka documentation's (Ch. 36)
-
Anything careful on property-based testing. The capstone's strongest assertions are properties, not (Ch. 38)
-
Anything careful on requirements elicitation. Case Study 2's four questions are a small, specific (Ch. 29, Ch. 39)
-
Anything careful on retrospective and revisable statistics. Economics has the best-developed (Ch. 26)
-
Anything careful on the "documentation is written by the wrong person" problem. The (Ch. 30)
-
Anything careful on the economics of sunk cost and its behavioural effects. §33.10's claim — that (Ch. 33)
-
Anything careful on weak supervision and programmatic labeling — the Snorkel line of work is the (Ch. 32)
-
Anything on "bus factor" and knowledge transfer. §38.12's test — hand the README to someone who has (Ch. 38)
-
Anything on "dark launching" and "shadow traffic." The service-side equivalent, better documented (Ch. 37)
-
Anything on "the ratchet" in operational process. Case Study 2's central mechanism — every (Ch. 25)
-
Anything on "the summary line problem" in tooling output. There is no canonical source and the (Ch. 28)
-
Anything on alert fatigue — the same recommendation as Chapter 23, for the same reason. Case (Ch. 24)
-
Anything on architectural fitness functions. The idea — an automated test that a system still has (Ch. 34)
-
Anything on blameless postmortems, particularly Google's SRE material. §17.7's fourth fix — that (Ch. 17)
-
Anything on controls that fail open, from safety engineering. Case Study 2's "a check whose (Ch. 27)
-
Anything on Conway's law and on the relationship between team boundaries and system (Ch. 17)
-
Anything on CRDTs (conflict-free replicated data types). §36.9's commutativity analysis is the (Ch. 36)
-
Anything on crypto-shredding / cryptographic erasure. NIST SP 800-88 (Guidelines for Media (Ch. 31)*
-
Anything on database migration tooling: Flyway, Liquibase, Alembic, or
dbtsnapshots. The (Ch. 27) -
Anything on evaluating engineering culture from outside — the "questions to ask your interviewer" (Ch. 39)
-
Anything on event sourcing's "the log is the source of truth." Chapter 36 takes this seriously. (Ch. 34)
-
Anything on graph modularity and community detection — the Louvain and Leiden algorithms are the (Ch. 35)
-
Anything on property-based testing applied to event ordering. §29.9's second case — out-of-order (Ch. 29)*
-
Anything on reading the SQL a GUI ETL tool generates. Informatica, DataStage, SSIS, and Talend all (Ch. 37)
-
Anything on reproducible builds — the Reproducible Builds project in software packaging is the (Ch. 38)
-
Anything on revenue recognition, particularly the treatment of gift cards and refunds. §38.5's R3 (Ch. 38)
-
Anything on runbooks and operational readiness reviews. Google's SRE material has the best-known (Ch. 38)
-
Anything on SaaS unit economics and cost of goods sold. §33.11's cost-per-order and percent-of-GMV (Ch. 33)
-
Anything on shadow IT discovery. The information-systems literature has thirty years on finding (Ch. 37)
-
Anything on shadow IT, which has a thirty-year literature. The consistent finding — shadow (Ch. 35)
-
Anything on structured behavioral interviewing, from the interviewer's side rather than the (Ch. 39)
-
Anything on Team Topologies (Skelton and Pais). The single most useful non-data book for this (Ch. 35)
-
Anything on the "manager versus IC" fork, written by someone who went one way and observed the (Ch. 40)
-
Anything on the "outbox vs. CDC vs. dual write" decision written by someone who has run all three. (Ch. 36)
-
Anything on the base-rate problem in screening. §31.2's false-positive argument is an instance of (Ch. 31)
-
Anything on the failure rate of decentralization initiatives generally — microservices adoption is (Ch. 35)
-
Anything on the history of OLAP semantic layers — Business Objects universes, Cognos frameworks, (Ch. 30)
-
Anything on three-valued logic and NULL semantics. The
NOT INfailure (§39.6) is one instance of (Ch. 39) -
Anything on time-series cross-validation — forward-chaining, walk-forward validation, purged (Ch. 32)
-
Anything on Zipf and power-law distributions in operational data. §21.7 draws a line between a (Ch. 21)
-
Anything sober on organizational change and technical projects that produce no features. The (Ch. 37)
-
Anything you can find on semantic drift detection. §17.9's first residual — a field's meaning (Ch. 17)
-
Apache Pulsar's documentation on its segment-based architecture, read as a comparison. Pulsar (Ch. 15)
-
Ask "what is this for, and is that still true?" of five things you maintain. Exercise 37.14. Two (Ch. 37)
-
Ask for an instance. Exercise 40.5. One meeting, four minutes, and it converts a criterion into a (Ch. 40)
-
Ask the Case Study 1 question of one green indicator today. What number was thresholded to (Ch. 25)*
-
Ask the question that started Case Study 1. "If a customer asks us to delete everything, how long (Ch. 31)*
-
Ask, of every compute resource you run: what fraction of its billed time is it working? §33.7. (Ch. 33)
-
Attempt a full rebuild. Exercise 34.12. If it succeeds on the first try, check that you rebuilt (Ch. 34)
-
Audit your mutes today. Exercise 23.16. If your alerting tool cannot tell you how many (Ch. 23)
-
AWS S3 documentation on request rates and performance. The primary source for the small-files (Ch. 2)
-
AWS S3 documentation: "Best practices design patterns: optimizing Amazon S3 performance," and (Ch. 9)
-
AWS S3 Lifecycle configuration documentation, and the equivalent for GCS and ADLS. §30.7's missing (Ch. 30)
-
Baron Schwartz, Peter Zaitsev, and Vadim Tkachenko, High Performance MySQL (O'Reilly). The (Ch. 7)
-
Barr Moses, Lior Gavish, and Molly Vorwerck, Data Quality Fundamentals (O'Reilly, 2022). The (Ch. 23, Ch. 25)
-
Bas Harenslak and Julian de Ruiter, Data Pipelines with Apache Airflow (2nd ed., Manning, (Ch. 24)
-
Benn Stancil's newsletter. The most consistently thoughtful writing about the data tooling (Ch. 5)
-
Benn Stancil's writing on metrics layers and the semantic layer problem. The clearest (Ch. 2)
-
Benoit Dageville et al., "The Snowflake Elastic Data Warehouse" (2016), SIGMOD. How separated (Ch. 8)
-
Betsy Beyer et al., The Site Reliability Workbook (O'Reilly, 2018). Free online, and the (Ch. 26)
-
BigQuery's dry-run documentation and Snowflake's
EXPLAIN. §33.6's pre-flight estimate. (Ch. 33) -
BigQuery: "Introduction to partitioned tables," "Introduction to clustered tables," and (Ch. 8)
-
Bill Inmon and Ralph Kimball's original disagreement, best encountered through Kimball's The (Ch. 3)*
-
Bill Inmon on the operational data store and the atomic warehouse layer. The other half of the (Ch. 34)
-
Bill Inmon, Building the Data Warehouse, 4th edition (Wiley, 2005). The other side of the (Ch. 6)
-
Break your CI on purpose. Exercise 27.14. Point
state:modified+at a missing manifest and watch (Ch. 27) -
Brendan Gregg, Systems Performance (2nd ed., Addison-Wesley, 2020), on the USE method. Not a (Ch. 18)
-
Bruce Momjian's presentations (
momjian.us/presentations). Free, extensive, and unusually (Ch. 7) -
Build an SCD2 dimension by hand once, without a snapshot macro. Exercise 20.15. Snapshot tools (Ch. 20)
-
Burrow (LinkedIn's consumer lag monitoring tool) and its design rationale. Burrow's argument — (Ch. 15)
-
Camille Fournier, The Manager's Path (O'Reilly, 2017). Nominally about management, and the (Ch. 40)
-
Case Study 1 and Chapter 33 §33.12 of this book. Delivering a finding about somebody else's (Ch. 37)
-
Cassie Kozyrkov's writing on decision-making and metrics. Not a data engineering source, and (Ch. 6)
-
CCPA/CPRA, and the California Privacy Protection Agency's regulations. Different structure, same (Ch. 31)
-
Chad Sanderson's writing on data contracts, and the surrounding discourse of the mid-2020s. He (Ch. 17)
-
Chapter 1 §1.7 — the acceptance criterion, and the frozen anchors that §38.7's verification (Ch. 38)
-
Chapter 1's Case Study 1 of this book, and the writing on metric layers referenced in Chapter (Ch. 6)
-
Chapter 11 §11.6 of this book. The honesty checklist for a format benchmark applies unchanged to (Ch. 22)
-
Chapter 17 of this book, on data contracts. Case Study 2 is a schema change with no schema, from (Ch. 22)
-
Chapter 20 of this book. Not a deflection: a Type 2 slowly changing dimension is a (Ch. 32)
-
Chapter 20 — incremental models and the deterministic tie-break Case Study 2's first defect (Ch. 38)
-
Chapter 23 and Chapter 26 of this book. Not a deflection. The register of assertions is the most (Ch. 30)
-
Chapter 23 §23.11 of this book, and Case Study 1. Reconciliation independence is the (Ch. 36)
-
Chapter 23 — the assertion register, and §38.11's honest reading of what it caught. (Ch. 38)
-
Chapter 25 and Chapter 26 of this book. Case Study 2's finding is that the entire monitoring stack (Ch. 33)
-
Chapter 25 Case Study 1 and Chapter 26 §26.4 of this book. Case Study 2's review is the same move (Ch. 29)
-
Chapter 25 of this book. Feature freshness is freshness; materialization lag is pipeline lag. The (Ch. 32)
-
Chapter 25 §25.12, Chapter 30 Case Study 2, and Chapter 33 Case Study 1 of this book. Measuring (Ch. 39)
-
Chapter 26 of this book. On-call design, and the observation that a rota that pages more than twice (Ch. 40)
-
Chapter 26 — the SLO and the slack that §38.10 publishes. (Ch. 38)
-
Chapter 27 §27.9 and Chapter 32 §32.7 of this book. Shadow running as a deployment technique, and (Ch. 37)
-
Chapter 27 — the CI that runs the models, and where
purity.pybelongs. (Ch. 38) -
Chapter 3's further reading on the lakehouse papers. Where warehouses are going: the Armbrust (Ch. 8)
-
Chapter 30 Case Study 2 and Chapter 25 §25.12 of this book. Query-log analysis and reviewing (Ch. 37)
-
Chapter 30 Case Study 2 of this book. The access review that could not say no. §31.6's policies (Ch. 31)
-
Chapter 30 §30.9 and Chapter 33 §33.12 of this book. Instrumenting the asking, and delivering a (Ch. 35)
-
Chapter 30 — the catalog, without which §38.12's second question needs a person. (Ch. 38)
-
Chapter 31 §31.5 of this book. The Iceberg migration that made row-level deletes possible was (Ch. 34, Ch. 36)
-
Chapter 31 — the deletion manifest, without which the platform is not finished. (Ch. 38)
-
Chapter 33 — the rate card §38.10 prices against, and the slack alert. (Ch. 38)
-
Chapter 34 §34.9 and Chapter 36 §36.10 of this book. The rebuild and the replay. §38.9 is the (Ch. 38)
-
Chapter 34 §34.9 — the rebuild. (Ch. 38)
-
Chapter 36 Case Study 1 and Chapter 38 Case Study 1. Reconciliation independence, and four failed (Ch. 39)
-
Chapter 36 Case Study 1 — reconciliation independence, which is §38.7's and §38.8's whole argument. (Ch. 38)
-
Chapter 37 §37.7 — a difference is usually a rule nobody wrote down, which is Case Study 1 four (Ch. 38)
-
Chapter 4 of this book, Case Study 2, and this chapter's Case Study 1. They are two halves of (Ch. 7)
-
Chapter 40's durable list, read as a bibliography. Kimball (1996), Lamport (1978), and — for (Ch. 40)
-
Chapter 9 of this book, on partitioning and file size. Not a deflection: partition pruning is (Ch. 33)
-
Chapter 9 §9.6 of this book, and its Case Study 1. The small-files problem in its (Ch. 10)
-
Chapters 17, 30, and 34 of this book. Contracts, catalog, ownership, and computationally-enforced (Ch. 35)
-
Chapters 23, 30, 34, 36, and 38 of this book. §40.13's claim is that the durable skill is knowing (Ch. 40)
-
Chapters 3, 4, 11, 29, and 34 of this book. Architecture principles, distributed systems, the (Ch. 39)
-
Charity Majors' writing on operational ownership and on-call. Blunt, experienced, and the best (Ch. 5)
-
Charity Majors, Liz Fong-Jones, and George Miranda, Observability Engineering (O'Reilly, 2022). (Ch. 25)
-
Chip Huyen, Designing Machine Learning Systems (O'Reilly, 2022). The best single book for this (Ch. 32)
-
Chris Richardson's microservices.io patterns, specifically Transactional Outbox, Polling (Ch. 36)
-
Christina Maslach's work on burnout, which is the actual research rather than the folklore. The (Ch. 40)
-
Cindy Sridharan, Distributed Systems Observability (O'Reilly, 2018). Short, free, and the (Ch. 25)
-
Clear a historical task in a local instance, deliberately. Exercise 24.13. It takes five minutes (Ch. 24)
-
Compute depth and blast radius for your own project, and compare against where your assertions (Ch. 34)
-
Compute feature age. Exercise 32.6. Nobody has, and the ratio is usually surprising. (Ch. 32)
-
Compute your unattributed share. Exercise 33.5. Report it honestly, including if it is (Ch. 33)
-
Confluent's and AWS MSK's documentation on monitoring. Both publish the metrics that matter — (Ch. 15)
-
Corey Quinn's Last Week in AWS newsletter and writing. Opinionated, frequently funny, and more (Ch. 33)
-
Count last month's alerts and how many were acted on. Exercise 25.16. If you cannot determine (Ch. 25)
-
Cube's documentation on data modeling. A different take on the same problem, and worth reading (Ch. 30)
-
Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy (2014). The (Ch. 31)
-
Damien Desfontaines' blog (
desfontain.es), the "differential privacy" series. The clearest (Ch. 31) -
Dan Linstedt on Data Vault, particularly the raw-vault / business-vault split. This is the same (Ch. 34)
-
Dan McKinley, "Choose Boring Technology" (2015). The innovation-tokens essay. The argument is (Ch. 5)
-
Dan McKinley, "Choose Boring Technology." Recommended in Chapter 5 and directly applicable (Ch. 12)
-
Daniel Abadi, "Consistency Tradeoffs in Modern Distributed Database System Design" (2012), IEEE (Ch. 4)
-
Daniel Abadi, Samuel Madden, and Nabil Hachem, "Column-Stores vs. Row-Stores: How Different Are (Ch. 8)
-
Daniel Linstedt and Michael Olschimke, Building a Scalable Data Warehouse with Data Vault (Ch. 6)*
-
Daniele Procida's "Diátaxis" framework for technical documentation. A runbook is what Diátaxis (Ch. 26)
-
Data observability vendors. Chapter 23 §23.8's assessment applies unchanged: a net for the (Ch. 25)
-
Databricks Feature Store and SageMaker Feature Store documentation. The managed offerings. Read (Ch. 32)
-
DataHub, Amundsen, and OpenMetadata — the open-source catalogs. All three are real, all three are (Ch. 30)
-
David Goldberg, "What Every Computer Scientist Should Know About Floating-Point Arithmetic" (Ch. 7)
-
David Salomon, Data Compression: The Complete Reference. If you want to understand why (Ch. 11)
-
DBT Labs, "What is analytics engineering?" and the accompanying discourse. Vendor material, (Ch. 1)
-
Dehghani's two original articles, "How to Move Beyond a Monolithic Data Lake to a Distributed Data (Ch. 35)
-
Delta Lake documentation: "Table utility commands" (
OPTIMIZE,VACUUM,RESTORE, (Ch. 10) -
Delta Lake's "Best practices" page and the equivalent Databricks optimization guidance. Read (Ch. 10)
-
Do the twelve drills against your own engine. Exercise 39.4, and check your window-frame default. (Ch. 39)
-
Douglas Terry et al., "Session Guarantees for Weakly Consistent Replicated Data" (1994). Where (Ch. 4)
-
Each vendor's usage and metering views: Snowflake's
SNOWFLAKE.ACCOUNT_USAGEschema (Ch. 8) -
Eric Brewer, "CAP Twelve Years Later: How the 'Rules' Have Changed" (2012), IEEE Computer. (Ch. 4)
-
Eric Evans, Domain-Driven Design (Addison-Wesley, 2003), on aggregates and domain events. An (Ch. 36)
-
Eric Evans, Domain-Driven Design, on bounded contexts and context maps. A context map is (Ch. 35)
-
Estimate three queries before running them. Exercise 33.4. Most people are wrong by more than (Ch. 33)
-
Fabian Reinartz et al. on the Prometheus TSDB design, and the Gorilla paper: Tuan Pham et al., (Ch. 12)
-
Fay Chang et al., "Bigtable: A Distributed Storage System for Structured Data" (2006), OSDI. (Ch. 1, Ch. 12)
-
Find your shadow pipelines. Exercise 35.10. Service accounts with unexplained query patterns, (Ch. 35)
-
For each track in §40.5, this book's own chapters: streaming (29), platform (24, 27, 28), analytics (Ch. 40)
-
Fred Brooks, The Mythical Man-Month (1975, anniversary edition 1995). Fifty years old and mostly (Ch. 40)
-
Fred Brooks, The Mythical Man-Month, on the second-system effect. Directly relevant and usually (Ch. 37)
-
GDPR Articles 12 and 17, and your jurisdiction's equivalent. Article 17 is the right to (Ch. 9)
-
GDPR, Articles 5, 6, 12–22, 25, 32, and 33–34. That is the engineering subset: principles, lawful (Ch. 31)
-
Giuseppe DeCandia et al., "Dynamo: Amazon's Highly Available Key-value Store" (2007), SOSP. (Ch. 4, Ch. 12)
-
Goodhart's law, and Marilyn Strathern's formulation of it — "when a measure becomes a target, it (Ch. 26)*
-
Google Cloud DLP / Sensitive Data Protection, and AWS Macie — the managed equivalents. Worth (Ch. 31)
-
Google Cloud, "Data lifecycle" and the equivalent architecture-center pages at AWS and Azure. (Ch. 2)
-
Google SRE Team, Site Reliability Engineering (free at
sre.google/books), the chapter on (Ch. 5) -
Google SRE Team, Site Reliability Engineering (O'Reilly, 2016), free online at
sre.google. (Ch. 1) -
Google SRE Team, Site Reliability Engineering, free at
sre.google/books. The DataOps (Ch. 2) -
Google's "Data Validation for Machine Learning" (Breck et al., SysML 2019) and TensorFlow Data (Ch. 32)
-
Google's Site Reliability Engineering (O'Reilly, 2016), and The Site Reliability Workbook (Ch. 25)
-
Google's Site Reliability Engineering (O'Reilly, 2016), on monitoring and alerting. The (Ch. 23)
-
Google's Site Reliability Engineering (O'Reilly, 2016), the chapters on monitoring and on (Ch. 19)
-
Google's Site Reliability Engineering (O'Reilly, 2016). Free online. Read Chapter 11, "Being (Ch. 26)
-
Google's Site Reliability Engineering, on alerting and on error budgets. Case Study 2's (Ch. 20)
-
Google's Site Reliability Engineering, on monitoring and on "the four golden signals." Case (Ch. 24)
-
Great Expectations, and Chapter 23's assessment of it. Its place in CI is validating inputs (Ch. 27)
-
Greg Young's talks and writing on event sourcing. The most experienced practitioner voice, and (Ch. 36)
-
Greg Young, Versioning in an Event Sourced System (Leanpub). The definitive treatment of §36.10, (Ch. 36)
-
Gregor Hohpe and Bobby Woolf, Enterprise Integration Patterns (Addison-Wesley, 2003). Old, (Ch. 13)
-
Gwen Shapira, Todd Palino, Rajini Sivaram, and Krit Petty, Kafka: The Definitive Guide, (Ch. 15)
-
Hand your README to someone. Exercise 38.9. Watch, do not help, and write down every point at which (Ch. 38)
-
Hironobu Suzuki, The Internals of PostgreSQL (
interdb.jp/pg). A free online book on the (Ch. 7) -
Holden Karau and Rachel Warren, High Performance Spark (2nd ed., O'Reilly, 2023). The book (Ch. 21)
-
Ian Robinson and the consumer-driven contracts pattern (2006), and Martin Fowler's write-up. (Ch. 17)
-
IBM's documentation on COBOL copybooks and packed decimal, if you meet a mainframe. Deeply (Ch. 13)
-
Iceberg's "Maintenance" documentation — expiring snapshots, removing orphan files, rewriting (Ch. 10)
-
Interview a company you are not going to join. Exercise 39.13. Uncomfortable, and it calibrates the (Ch. 39)
-
Interview someone. The fastest way to learn what the rubric is measuring is to sit on the other (Ch. 39)
-
Itzik Ben-Gan, T-SQL Window Functions: For Data Analysis and Beyond (2nd ed., Microsoft Press, (Ch. 18)
-
J.R. Storment and Mike Fuller, Cloud FinOps (O'Reilly, 2nd ed. 2023). The standard text. The (Ch. 33)
-
Jake VanderPlas, Python Data Science Handbook (2nd ed., O'Reilly, 2022). Free online. The (Ch. 22)
-
Jay Kreps, "Questioning the Lambda Architecture" (2014). The essay that named and then argued (Ch. 3, Ch. 29)
-
Jay Kreps, "The Log: What every software engineer should know about real-time data's unifying (Ch. 1, Ch. 14, Ch. 15, Ch. 29, Ch. 36)
-
Jerry Muller, The Tyranny of Metrics (Princeton, 2018). Case Study 1's problem — a measure that (Ch. 26)
-
Jez Humble and David Farley, Continuous Delivery (Addison-Wesley, 2010). Old, foundational, and (Ch. 27)
-
Jim Gray and Andreas Reuter, Transaction Processing: Concepts and Techniques (1992), on (Ch. 4)
-
Job postings, read as market data rather than as opportunities. Count how many roles in your market (Ch. 40)
-
Joe Celko, SQL for Smarties: Advanced SQL Programming (5th ed., Morgan Kaufmann, 2014). The (Ch. 18)
-
Joe Reis and Matt Housley, Fundamentals of Data Engineering (O'Reilly, 2022), Chapter 2. (Ch. 2)
-
Joe Reis and Matt Housley, Fundamentals of Data Engineering (O'Reilly, 2022), on (Ch. 24)
-
Joe Reis and Matt Housley, Fundamentals of Data Engineering (O'Reilly, 2022). The book that (Ch. 1, Ch. 19)
-
Joe Reis and Matt Housley, Fundamentals of Data Engineering, Chapter 8 ("Queries, Modeling, and (Ch. 6)
-
Joel Spolsky, "In Defense of Not-Invented-Here Syndrome" (2001). Twenty-plus years old and the (Ch. 5)
-
John Allspaw and the "blameless postmortem" literature. Case Study 1's known-issues entry was (Ch. 19)
-
John Allspaw on incident analysis, and the resilience-engineering literature generally. Case (Ch. 25)
-
John Allspaw's writing on incident analysis, particularly on how organizations converge on a (Ch. 18)
-
John D. C. Little's original 1961 paper, "A Proof for the Queuing Formula $L = \lambda W$", (Ch. 3)
-
Jordan Tigani, "Big Data Is Dead" (2023), MotherDuck blog. Argues, with data from BigQuery (Ch. 3, Ch. 5)
-
Jordan Tigani, "Big Data Is Dead" (2023). Recommended in Chapter 5 and relevant again: the (Ch. 8)
-
Jules Damji, Brooke Wenig, Tathagata Das, and Denny Lee, Learning Spark (2nd ed., O'Reilly, (Ch. 21)
-
Julien Le Dem and Nong Li, "Parquet: Columnar Storage for Hadoop" (2013) and the associated (Ch. 9)
-
Kapoor and Narayanan, "Leakage and the Reproducibility Crisis in ML-based Science" (2023). (Ch. 32)
-
Kaufman, Rosset, and Perlich, "Leakage in Data Mining" (KDD 2011 / TKDD 2012). The formal treatment, (Ch. 32)
-
Kent Beck and others on "make it work, make it right, make it fast." Relevant backwards: the (Ch. 38)
-
Kief Morris, Infrastructure as Code (2nd ed., O'Reilly, 2020). The standard reference and the (Ch. 28)
-
Kimball Group Design Tips, the archived newsletter. Short, specific, and several of them address (Ch. 20)
-
Kimball's original articles on SCD types, collected in the Reader and widely summarized (Ch. 6)
-
KIP-345 (static membership) and KIP-429 (incremental cooperative rebalancing). Kafka (Ch. 15)
-
KIP-98 (exactly-once delivery and transactional messaging). The design document behind the (Ch. 15)
-
Kyle Kingsbury's Jepsen reports (
jepsen.io). Empirical testing of distributed databases' (Ch. 4, Ch. 13) -
Lamport, "Time, Clocks, and the Ordering of Events in a Distributed System" (1978). The foundation. (Ch. 36)
-
Latanya Sweeney, "k-Anonymity: A Model for Protecting Privacy" (2002), and her earlier work (Ch. 31)
-
Laura Sebastian-Coleman, Measuring Data Quality for Ongoing Improvement (Morgan Kaufmann, (Ch. 23)
-
Lauren Balik and the critical dbt commentary of the mid-2020s. There is a genuine dissenting (Ch. 19)
-
Leslie Lamport, "Time, Clocks, and the Ordering of Events in a Distributed System" (1978), (Ch. 4)
-
Look up one stateful resource's
ForceNewattributes and check whether any of them is something (Ch. 28) -
Machanavajjhala et al., "l-Diversity: Privacy Beyond k-Anonymity" (2006). The homogeneity attack — (Ch. 31)
-
Marc Brooker's blog (
brooker.co.za/blog) and the AWS Builders' Library (Ch. 4) -
Marc Brooker's personal blog (
brooker.co.za/blog). The long-form version of the same (Ch. 16) -
Marc Brooker, "Timeouts, retries, and backoff with jitter" (AWS Builders' Library). The primary (Ch. 16)
-
Mark Raasveldt and Hannes Mühleisen's papers on DuckDB, particularly "DuckDB: an Embeddable (Ch. 22)*
-
Mark Raasveldt et al., "Fair Benchmarking Considered Difficult: Common Pitfalls In Database (Ch. 8, Ch. 11)
-
Markus Winand, SQL Performance Explained (2012), and the companion site (Ch. 18)
-
Markus Winand, SQL Performance Explained and the companion site
use-the-index-luke.com. The (Ch. 7) -
Martin Fowler on "Polyglot Persistence." The essay that named the idea that different parts of (Ch. 12)
-
Martin Fowler on "the two hard things" and, more usefully, anything careful on ubiquitous (Ch. 34)
-
Martin Fowler on ParallelChange (also called expand-contract). The refactoring pattern (Ch. 17)
-
Martin Fowler on the Strangler Fig Application (2004). §37.5, from the person who named it. Short, (Ch. 37)
-
Martin Fowler's writing on "expand and contract" / parallel change. Chapter 17's compatibility (Ch. 27)
-
Martin Fowler's writing on the Outbox pattern and on dual writes. Directly relevant to §13.5's (Ch. 13)
-
Martin Fowler, "Event Sourcing" (2005) and "CQRS" (2011). The canonical write-ups. The event (Ch. 36)
-
Martin Fowler, "Utility vs Strategic Dichotomy" and the related writing on technical (Ch. 5)
-
Martin Fowler, "What do you mean by 'Event-Driven'?" (2017). Short, free, and it is §36.2 — the (Ch. 36)
-
Martin Fowler, "Who Needs an Architect?" (2003), IEEE Software. The source of the (Ch. 3)
-
Martin Kleppmann, "Turning the database inside-out with Apache Samza" (2015), and the (Ch. 14)
-
Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Chapter 10. The (Ch. 21)
-
Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Chapter 11. Streams (Ch. 29, Ch. 36)
-
Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Chapters 7 and 11. (Ch. 20)
-
Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017). The best (Ch. 1, Ch. 4, Ch. 39)
-
Martin Kleppmann, Designing Data-Intensive Applications, Chapter 11 ("Stream Processing"), (Ch. 14, Ch. 15)
-
Martin Kleppmann, Designing Data-Intensive Applications, Chapter 2 ("Data Models and Query (Ch. 12)
-
Martin Kleppmann, Designing Data-Intensive Applications, Chapter 7 ("Transactions"). The (Ch. 2)
-
Matt Harrison, Effective Pandas (2021, and a second edition). Idiomatic pandas: method (Ch. 22)
-
Maxime Beauchemin, "Functional Data Engineering — a modern paradigm for batch data (Ch. 2)
-
Maxime Beauchemin, "Functional Data Engineering — a modern paradigm for batch data processing" (Ch. 13)
-
Maxime Beauchemin, "Functional Data Engineering" (2018). Recommended in Chapter 2 and again (Ch. 6)
-
Maxime Beauchemin, "The Rise of the Data Engineer" (2017) and "The Downfall of the Data (Ch. 1)
-
Measure k on something you actually ship. Exercise 31.6. If your organization sends any extract (Ch. 31)
-
Measure your bottleneck. Exercise 35.4. Split by wait time, not volume, and look at who could (Ch. 35)
-
Measure your own
processing_time − event_timedistribution. Exercise 29.13. Everyone picks 30 (Ch. 29) -
Melanie Mitchell's and others' writing on institutional metrics and Goodhart's law, applied to (Ch. 30)
-
Michael Armbrust et al., "Delta Lake: High-Performance ACID Table Storage over Cloud Object (Ch. 3, Ch. 10)
-
Michael Armbrust et al., "Lakehouse: A New Generation of Open Platforms that Unify Data (Ch. 3, Ch. 10)
-
Michael Feathers, Working Effectively with Legacy Code (Prentice Hall, 2004). Twenty years old (Ch. 37)
-
Michael Nygard, "Documenting Architecture Decisions" (2011). The essay that introduced the ADR (Ch. 3)
-
Michael Nygard, Release It! (2nd ed., Pragmatic Bookshelf, 2018). On deploying things that (Ch. 27)
-
Microsoft Presidio — open-source PII detection and anonymization, and a substantially more (Ch. 31)
-
Mike Stonebraker et al., "C-Store: A Column-oriented DBMS" (2005), VLDB. The ancestor of (Ch. 8)
-
Modern SQL's feature tables, at modern-sql.com. Which engines support (Ch. 18)
-
Narayanan and Shmatikov, "Robust De-anonymization of Large Sparse Datasets" (2008) — the Netflix (Ch. 31)
-
Nathan Marz and James Warren, Big Data (Manning, 2015) — the Lambda architecture, from its (Ch. 29)
-
Nathen Harvey and the PagerDuty incident response documentation. PagerDuty publishes its internal (Ch. 26)
-
Neha Narkhede, "Exactly-once Semantics are Possible: Here's How Kafka Does it" (2017), Confluent (Ch. 4)
-
Neha Narkhede, "Exactly-once Semantics are Possible: Here's How Kafka Does it" (2017). (Ch. 15)
-
Neil Gunther, Guerrilla Capacity Planning (Springer, 2007). Capacity planning for people who (Ch. 3)
-
Nickolas Means and the "how they built it" genre, plus Google's Site Reliability Engineering (Ch. 28)
-
Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate (IT Revolution, 2018). The evidence base: (Ch. 27)
-
NIST SP 800-53, control family AC (Access Control), and AC-2(3) specifically. Dry, and the (Ch. 30)
-
Northcutt, Athalye, and Mueller, "Pervasive Label Errors in Test Sets" (2021). Found substantial (Ch. 32)
-
OpenCost and Kubecost, if you run on Kubernetes. Allocation for shared clusters is the hardest (Ch. 33)
-
Pact and the consumer-driven contract testing literature. A more formal approach to §16.9's (Ch. 16)
-
Parts IV and V of this book. Chapter 20 (idempotency, SCD2, the deterministic tie-break), Chapter (Ch. 39)
-
Piethein Strengholt, Data Management at Scale (O'Reilly, 2nd ed. 2023). The most practical book (Ch. 30)
-
PostgreSQL documentation, "Concurrency Control" (Chapter 13 of the manual). The section on (Ch. 2)
-
PostgreSQL documentation, "High Availability, Load Balancing, and Replication" (Chapter 26 of (Ch. 4)
-
PostgreSQL documentation,
postgresql.org/docs. Genuinely excellent, unusually well-written (Ch. 1) -
Postmortem practice — Allspaw, Dekker, and the resilience-engineering literature. Both case (Ch. 23)
-
Practice sites with a data slant — anything offering realistic multi-table problems rather than (Ch. 39)
-
Pramod Sadalage and Martin Fowler, NoSQL Distilled (Addison-Wesley, 2012). Short, and its (Ch. 12)
-
Prometheus + Grafana for metrics, Loki / Elasticsearch / a cloud log service for structured (Ch. 25)
-
Public engineering blogs and incident write-ups from the company you are interviewing with. A (Ch. 39)
-
Public engineering write-ups on cost-per-unit at scale — Dropbox's storage migration, Netflix's (Ch. 33)
-
Publicly published engineering ladders — several companies have opened theirs, and (Ch. 40)
-
Ralph Kimball and Joe Caserta, The Data Warehouse ETL Toolkit (Wiley, 2004). The companion to (Ch. 13)
-
Ralph Kimball and Margy Ross, The Data Warehouse Toolkit (3rd ed., Wiley, 2013), Chapter 5. (Ch. 20)
-
Ralph Kimball and Margy Ross, The Data Warehouse Toolkit (3rd ed., Wiley, 2013), on error event (Ch. 23)
-
Ralph Kimball and Margy Ross, The Data Warehouse Toolkit (3rd ed., Wiley, 2013). Referenced in (Ch. 19)
-
Ralph Kimball and Margy Ross, The Data Warehouse Toolkit, 3rd edition (Wiley, 2013). The (Ch. 1)
-
Ralph Kimball and Margy Ross, The Data Warehouse Toolkit, 3rd edition, Chapter 1. The (Ch. 2)
-
Ralph Kimball and Margy Ross, The Data Warehouse Toolkit: The Definitive Guide to Dimensional (Ch. 6)*
-
Ralph Kimball and Margy Ross, The Kimball Group Reader, 2nd edition (Wiley, 2016). Collected (Ch. 6)
-
Ralph Kimball, The Data Warehouse Toolkit (Wiley, 3rd ed. 2013), the chapters on the back room and (Ch. 34)
-
Read
sys/fs/cgroup/memory.peakfor one job you own, today. Compare it to the limit. Case Study (Ch. 22) -
Read one production plan a week. Not to fix anything — just to build the habit of counting (Ch. 21)
-
Rebuild and diff row by row. Exercise 38.5. A rebuild that succeeds first time usually means you (Ch. 38)
-
Reconcile something real. Exercise 38.12. Write down every rule before you run the comparison, (Ch. 38)
-
Redshift: "Distribution styles" and "Sort keys." The two settings that determine whether (Ch. 8)
-
Rehearse a rollback. Exercise 37.9. Time it, and compare against what the plan claims. (Ch. 37)
-
Replay something. Exercise 36.10. It will fail, and the failure is the point. (Ch. 36)
-
Rewrite one production pandas job in DuckDB. Not a large one. The exercise is worth doing for (Ch. 22)
-
RFC 4180, "Common Format and MIME Type for Comma-Separated Values Files" (2005). Two pages, (Ch. 11)
-
RFC 5988 / RFC 8288, "Web Linking." The
Linkheader format, for §16.2's fourth pagination (Ch. 16) -
RFC 6585 §4 (status 429) and the IETF
RateLimitheader fields draft. The 429 definition, and (Ch. 16) -
RFC 6749 (OAuth 2.0) and RFC 6750 (Bearer Token Usage). Read §4.4 (client credentials) and §6 (Ch. 16)
-
RFC 8594, the
SunsetHTTP header, and theDeprecationheader draft. The mechanism by which (Ch. 16) -
RFC 9110 §10.2.3 on
Retry-After. Two paragraphs, and it specifies both accepted forms — (Ch. 16) -
RFC 9110, "HTTP Semantics." The current HTTP specification, superseding RFC 7231. Read the (Ch. 16)
-
Richard Cook, "How Complex Systems Fail" (1998). Recommended in Chapter 1 and recommended (Ch. 4)
-
Richard Cook, "How Complex Systems Fail" (1998/2000). Eighteen numbered propositions, four (Ch. 1)
-
Rob Ewaschuk, "My Philosophy on Alerting." An internal Google document, widely circulated, (Ch. 25)
-
Roy Fielding and the REST/hypermedia literature on evolvability, particularly the argument that (Ch. 17)
-
Run
pr_report.pyagainst your last ten merged changes and count the definitional ones that were (Ch. 27) -
Run
terraform plan -refresh-onlytoday. Exercise 28.15. Write down your prediction first, (Ch. 28) -
Run a mock loop. Exercise 39.11. With another person, scored against the rubric, with feedback (Ch. 39)
-
Run one runbook drill this quarter. Exercise 26.18. One runbook, one colleague who did not write (Ch. 26)
-
Run the grep. Case Study 1's is four minutes and finds something in almost every organization. If (Ch. 30)
-
Run the row-level diff on any feature computed in two places. Exercise 32.7. An afternoon, and the (Ch. 32)
-
Run the time-split check on a model you have access to. Exercise 32.12. If the gap is small, (Ch. 32)
-
Run the two queries. Exercise 30.7. Granted versus used, against a warehouse you have access to. (Ch. 30)
-
Run §29.11's four questions against the next "we need real time" request you receive. Case Study (Ch. 29)
-
Ryan Blue and Daniel Weeks' Iceberg papers and talks (Netflix). Iceberg's design rationale, (Ch. 10)
-
Sam Newman, Monolith to Microservices (O'Reilly, 2019). The most practical modern treatment of (Ch. 37)
-
Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung, "The Google File System" (2003), SOSP. (Ch. 1)
-
Scan for ambiguous timestamps. Exercise 34.13. Kestrel found 14 of 41 external timestamp columns (Ch. 34)
-
Score your own team. Exercise 40.8. Honestly, including the signals you do not have. (Ch. 40)
-
Search your codebase for dual writes. Exercise 36.4. A database write followed by a publish, an (Ch. 36)
-
Sergey Melnik et al., "Dremel: Interactive Analysis of Web-Scale Datasets" (2010), VLDB, and (Ch. 8)
-
Shirshanka Das et al. on DataHub's architecture (LinkedIn engineering, and the subsequent talks). (Ch. 30)
-
Sidney Dekker, The Field Guide to Understanding 'Human Error' (3rd ed., CRC Press, 2014). The (Ch. 26)
-
Snowflake's
ACCESS_HISTORYview in particular. It reports column-level access, which turns (Ch. 30) -
Snowflake's masking policy and row access policy documentation; the BigQuery column-level security (Ch. 31)
-
Snowflake: "Understanding Snowflake Table Structures" (micro-partitions and clustering), and (Ch. 8)
-
Soda, Monte Carlo, Elementary, and the observability category generally. Worth an evaluation (Ch. 23)
-
Sort your projections by commutativity. Exercise 36.9. An afternoon, and it may save two (Ch. 36)
-
Start the running document. Exercise 40.1. Today, not when you finish the chapter. (Ch. 40)
-
Studies of how developers and analysts actually find things. The consistent finding across (Ch. 30)
-
Take one inventory. Exercise 37.2. Five jobs, eight fields, without reading the code. The (Ch. 37)
-
Take the unwanted job. Exercise 40.11 — the reconciliation, the metric definitions, the finance (Ch. 40)
-
Tanya Reilly, The Staff Engineer's Path (O'Reilly, 2022). The practical companion to Larson. (Ch. 40)
-
The "Best Practices" page in the same documentation. It says most of §24.2 and §24.10, including (Ch. 24)
-
The "birthday problem" in any elementary probability text. The mechanism behind this chapter's (Ch. 13)
-
The "Data Mesh Architecture" community site and the Data Mesh Learning community. Practitioner (Ch. 35)
-
The "data mesh is not for you" genre. Several thoughtful practitioners have written versions of (Ch. 35)
-
The "error budget" material in SRE. §23.1's argument that a reconciliation tolerance is a budget (Ch. 23)
-
The "Falsehoods Programmers Believe About..." genre, particularly the entries on CSV, on names, (Ch. 11)
-
The
awesome-data-engineeringstyle lists on GitHub, and the various "data stack landscape" (Ch. 5) -
The
dbt_expectationspackage. The bridge between the two worlds: Great Expectations' (Ch. 23) -
The
dbt_utilspackage README, at github.com/dbt-labs/dbt-utils. (Ch. 19) -
The
delta-rsdocumentation and repository. The Rust implementation with Python bindings, (Ch. 10) -
The
h2oai/db-benchmarkproject and its successors. A maintained cross-engine benchmark on (Ch. 22) -
The
lzbenchbenchmark suite. A comparison of dozens of compressors on standard corpora. (Ch. 11) -
The
manifest.jsonschema documentation. dbt publishes a JSON schema for the artifacts it (Ch. 19) -
The
parquet-formatrepository'sREADMEandLogicalTypes.mdon GitHub. More precise than (Ch. 9) -
The
psycopg3 documentation on server-side cursors. Directly relevant to this chapter's first (Ch. 7) -
The
WITH/WITH RECURSIVEdocumentation for your engine. PostgreSQL's is the reference (Ch. 18) -
The Airbyte protocol documentation. A more recent formalization of the same problem, with a (Ch. 13)
-
The Airflow documentation on "Data Interval," "Catchup," and "Backfill." Even before Chapter (Ch. 13)
-
The Airflow documentation on
airflow db cleanand database maintenance. Case Study 2. It is (Ch. 24) -
The Airflow documentation on executors, comparing
LocalExecutor,CeleryExecutor, and (Ch. 28) -
The Airflow documentation on metrics (StatsD and OpenTelemetry). §25.13. Four lines of (Ch. 25)
-
The Airflow source,
airflow/jobs/scheduler_job_runner.py. An unusual recommendation, and worth (Ch. 24) -
The Apache Airflow documentation, and specifically these pages read end to end rather than (Ch. 24)
-
The Apache Airflow documentation, especially "Best Practices" and "Concepts." Airflow's docs (Ch. 5)
-
The Apache Airflow, Kafka, Spark, and Iceberg documentation sets. Each is the primary source (Ch. 1)
-
The Apache Arrow columnar format specification. The in-memory counterpart to Parquet's on-disk (Ch. 11)
-
The Apache Arrow documentation, particularly on the Arrow/Parquet relationship. Arrow is the (Ch. 9)
-
The Apache Arrow specification and the Arrow Columnar Format page. §22.5's material. You do (Ch. 22)
-
The Apache Avro specification (
avro.apache.org/docs/), particularly "Schema Resolution." (Ch. 11) -
The Apache Avro specification, "Schema Resolution." Recommended in Chapter 11 and mandatory (Ch. 17)
-
The Apache Beam documentation on watermarks and triggers. Beam's model is the most careful (Ch. 20)
-
The Apache Cassandra documentation on data modelling, particularly "Basic Rules of Cassandra (Ch. 12)
-
The Apache Flink documentation on event time, watermarks, and idleness. §29.5's (Ch. 29)
-
The Apache Hive documentation on partitioning, which is where
key=valuepath partitioning (Ch. 9) -
The Apache Iceberg and Delta Lake documentation on row-level deletes, deletion vectors, and (Ch. 31)
-
The Apache Iceberg and Delta Lake documentation on time travel and schema evolution. Bronze's (Ch. 34)
-
The Apache Iceberg specification (
iceberg.apache.org/spec/). Read alongside the Delta paper. (Ch. 10) -
The Apache Iceberg specification (
iceberg.apache.org/spec/). Worth reading alongside the (Ch. 3) -
The Apache Kafka documentation (
kafka.apache.org/documentation). Read three parts properly: (Ch. 15) -
The Apache ORC specification (
orc.apache.org/specification/). Worth skimming alongside (Ch. 11) -
The Apache Parquet documentation and format specification (
parquet.apache.org/docs/). Read (Ch. 9) -
The Apache Parquet format specification (
parquet.apache.org/docs/file-format/). Short and (Ch. 2, Ch. 11) -
The Apache Spark documentation on adaptive query execution (AQE) and skew join optimization. (Ch. 4)
-
The Apache XTable project (formerly OneTable). Translates metadata between Delta, Iceberg, and (Ch. 10)
-
The Astronomer documentation and guides. A vendor, and the best free Airflow teaching material (Ch. 24)
-
The AWS announcement of strong read-after-write consistency (December 2020). Worth reading (Ch. 9)
-
The AWS Builders' Library more generally, particularly "Avoiding fallback in distributed (Ch. 16)
-
The AWS Kinesis and Google Pub/Sub documentation, if you are on those platforms. Both solve the (Ch. 15)
-
The AWS S3 documentation on request rate and performance. The reason Case Study 2's per-file (Ch. 21)
-
The AWS, GCP, and Azure documentation on Spot / Preemptible / Spot VMs, particularly the (Ch. 33)
-
The Cassandra CDC documentation, read specifically to understand why most teams do not use it — (Ch. 12)
-
The checklist literature: Atul Gawande, The Checklist Manifesto (Metropolitan Books, 2009). (Ch. 26)
-
The clinical alarm-fatigue literature. Genuinely worth reading, and almost nobody in software (Ch. 25)
-
The Confluent Schema Registry documentation on compatibility modes.
BACKWARD,FORWARD, and (Ch. 36) -
The Confluent Schema Registry documentation on compatibility types. The operational (Ch. 17)
-
The Dagster documentation, particularly on Software-Defined Assets. §24.7's datasets, taken all (Ch. 24)
-
The DAMA-DMBOK (Data Management Body of Knowledge), 2nd edition. Dry, comprehensive, and the (Ch. 9)
-
The DAMA-DMBOK (Data Management Body of Knowledge, 2nd ed.). The reference work, and it is a (Ch. 30)
-
The Databricks documentation on the medallion architecture. Where the bronze/silver/gold naming (Ch. 34)
-
The dbt "Best Practices" guide on CI/CD, and the community writing around slim CI. Read it for (Ch. 27)
-
The dbt "Best Practices" guides, in the same documentation. The staging/intermediate/marts (Ch. 19)
-
The dbt
--selectand--excludegraph selectors documentation. Not obviously about this (Ch. 34) -
The dbt community's writing on "one big table" versus star schemas. A genuine, current (Ch. 6)
-
The dbt documentation (
docs.getdbt.com). Among the best documentation in this field, and the (Ch. 5) -
The dbt documentation at docs.getdbt.com. Unusually good, and (Ch. 19)
-
The dbt documentation on
docs,meta, and exposures. §30.11's argument in software form: (Ch. 30) -
The dbt documentation on incremental models, including
incremental_strategy, (Ch. 20) -
The dbt documentation on snapshots and incremental models. §32.9's seven-of-nine answer. Snapshots (Ch. 32)
-
The dbt documentation on snapshots. The
checkversustimestampstrategy discussion, the (Ch. 20) -
The dbt documentation on source freshness, and then go and check whether your project runs it. (Ch. 19)
-
The dbt documentation on staging, intermediate, and marts. dbt's naming maps to silver/gold with (Ch. 34)
-
The dbt documentation on state comparison,
--defer, and artifacts. The primary source for (Ch. 27) -
The dbt documentation on tests,
severity,error_if/warn_if, and--store-failures. The (Ch. 23) -
The dbt Learn courses at courses.getdbt.com. Free, official, and (Ch. 19)
-
The dbt Semantic Layer documentation, and the MetricFlow specification. Case Study 1's problem, (Ch. 30)
-
The dbt snapshots documentation (
docs.getdbt.com, "Snapshots"). dbt implements Type 2 (Ch. 6) -
The Debezium blog post on incremental snapshots. The implementation write-up alongside the (Ch. 14)
-
The Debezium documentation on the outbox event router. The CDC-on-the-outbox pattern Kestrel (Ch. 36)
-
The Debezium documentation, "PostgreSQL Connector." Read the section on replication slots and (Ch. 7)
-
The Debezium documentation, especially the PostgreSQL connector page. Read three sections (Ch. 14)
-
The Debezium FAQ on
REPLICA IDENTITY. Short, and it is the fix for this chapter's second case (Ch. 14) -
The Delta Lake and Apache Iceberg documentation on
OPTIMIZE/ compaction and onZORDER/ (Ch. 9) -
The Delta Lake documentation on
OPTIMIZEand Z-ordering. Case Study 2's actual fix. Chapter 10 (Ch. 21) -
The Delta Lake UniForm / Iceberg compatibility work, and the various "one format to read them (Ch. 10)
-
The Docker Compose specification (
docs.docker.com/compose/compose-file/). The reference for (Ch. 5) -
The Docker documentation on multi-stage builds and on content-addressable image identifiers. (Ch. 28)
-
The DuckDB and Polars documentation, and Chapter 22. §21.1's argument depends on knowing what a (Ch. 21)
-
The DuckDB blog. Unusually honest performance write-ups, including ones where DuckDB is not (Ch. 8)
-
The DuckDB documentation (
duckdb.org/docs). Short, well-written, and the "Guides" section is (Ch. 5) -
The DuckDB documentation on
read_parquet,parquet_metadata, and Hive partitioning. DuckDB (Ch. 9) -
The DuckDB documentation. Unusually good, unusually short, and the pages worth reading end to (Ch. 22)
-
The EDPB's opinions and guidelines, particularly anything on anonymisation and pseudonymisation. (Ch. 31)
-
The Elasticsearch guide's "Getting Started" and the relevance/scoring chapters, particularly on (Ch. 12)
-
The Feast documentation, particularly
get_historical_featuresand the entity-dataframe concept. (Ch. 32) -
The FinOps Foundation's framework and its "capabilities" list. Free, vendor-neutral, and useful as (Ch. 33)
-
The FinOps Foundation's materials (
finops.org). Vendor-neutral practice for cloud cost (Ch. 8) -
The Flink
TestHarnessdocumentation, and Spark'sMemoryStream. The real versions of (Ch. 29) -
The Flink documentation on state, savepoints, and operator UIDs. §29.7 and §29.10. The operator (Ch. 29)
-
The GDPR text itself, Articles 12 and 17 (
gdpr-info.euor the official EUR-Lex text). (Ch. 2) -
The GitHub Actions, GitLab CI, or Buildkite documentation for whichever you use — specifically (Ch. 27)
-
The GitHub Engineering "Scientist" library and its write-up (2016). A small library for running old (Ch. 37)
-
The Google Cloud Storage and Azure Blob Storage documentation on consistency and performance. (Ch. 9)
-
The Google SRE book's chapter on access and the "Building Secure and Reliable Systems" companion (Ch. 30)
-
The Google SRE book's chapter on eliminating toil, and its 50% rule. §40.7's first failure mode, (Ch. 40)
-
The Great Expectations documentation. The concepts pages first — Expectations, Suites, (Ch. 23)
-
The IANA time zone database documentation, and any careful treatment of timestamp handling. Case (Ch. 34)
-
The IAPP's practitioner material and the CIPT certification syllabus. The syllabus itself is a (Ch. 31)
-
The ICO's guidance on data retention (UK) and the equivalent from your regulator. Written for (Ch. 30)
-
The ICO's guidance (UK) — the best practitioner-facing writing on this subject anywhere, and it is (Ch. 31)
-
The idempotence and convergence literature from configuration management — Puppet's and Chef's (Ch. 28)
-
The Kafka Connect documentation on offsets,
errors.tolerance, and dead-letter queues. (Ch. 14) -
The Kafka documentation on log compaction. Read it specifically to understand why compaction and (Ch. 36)
-
The Kafka documentation on partitioning and ordering guarantees. Read the exact wording: Kafka (Ch. 36)
-
The Kafka documentation on transactions and exactly-once semantics, and the original KIP-98 (Ch. 29)
-
The Kubernetes documentation on operational overhead, and more usefully, any postmortem (Ch. 28)
-
The Linux kernel documentation on cgroup v2 memory control. Case Study 1's material from the (Ch. 22)
-
The literature on documentation that gets read — Diátaxis is the most useful framework, because it (Ch. 38)
-
The literature on requirements phrasing, or failing that, the discipline of writing acceptance (Ch. 27)
-
The literature on shift work, sleep disruption, and cognitive performance. §26.12's "nights (Ch. 26)
-
The literature on spreadsheet errors — Panko's work is the standard reference, and the reported (Ch. 37)
-
The MinIO documentation on S3 API compatibility. Specifically the compatibility matrix, which (Ch. 5)
-
The MongoDB change streams documentation. The best non-relational change feed in this chapter, (Ch. 12, Ch. 14)
-
The MySQL documentation on the binary log, particularly
binlog_formatand whyROWis the (Ch. 14) -
The MySQL Reference Manual, "InnoDB Storage Engine" and "The Binary Log." The clustered-index (Ch. 7)
-
The Neo4j documentation on Cypher, and any comparison of Cypher to recursive SQL. Worth an (Ch. 12)
-
The OAuth 2.1 draft, which consolidates a decade of security guidance. Worth knowing exists; (Ch. 16)
-
The OpenLineage specification and Marquez. An open standard for emitting lineage events from (Ch. 25)
-
The OpenLineage specification, and Marquez as a reference implementation. The vendor-neutral (Ch. 30)
-
The OpenTofu documentation, if that is what your organization uses. The divergence from Terraform (Ch. 28)
-
The Oracle GoldenGate and SQL Server CDC documentation, if you meet them. SQL Server's built-in (Ch. 14)
-
The original "data lake" coinage — James Dixon's 2010 blog post — and the "data swamp" (Ch. 9)
-
The Pact documentation and its "consumer-driven" framing. The service-testing implementation of (Ch. 17)
-
The pandas documentation's "Scaling to large datasets" page, and the Copy-on-Write page. (Ch. 22)
-
The PayPal data contract template, published openly, and the Open Data Contract Standard (Ch. 17)
-
The pgvector repository README. Short, practical, and honest about limits. It documents both (Ch. 12)
-
The Polars user guide, particularly the "Lazy API" and "Expressions" sections. §22.3's (Ch. 22)
-
The PostgreSQL
JSONBdocumentation and the GIN index chapter. The honest comparison point for (Ch. 12) -
The PostgreSQL
pg_replication_slotsview documentation. Every column, particularly (Ch. 14) -
The PostgreSQL
pg_stat_replicationandpg_last_xact_replay_timestampdocumentation. The (Ch. 13) -
The PostgreSQL documentation (
postgresql.org/docs). Read these chapters, in this order, and (Ch. 7) -
The PostgreSQL documentation on partitioning, parallel query, and
BRINindexes. Read as a (Ch. 5) -
The PostgreSQL documentation on table sizes and
pg_total_relation_size. Two queries from (Ch. 24) -
The PostgreSQL documentation on transaction isolation and on logical replication slots. The (Ch. 20)
-
The PostgreSQL documentation, "Concurrency Control" (Chapter 13 of the manual). Recommended in (Ch. 13)
-
The PostgreSQL documentation, "Logical Decoding" and "Logical Replication." Recommended in (Ch. 14)
-
The PostgreSQL full-text search chapter, and the
pg_trgmextension documentation. The (Ch. 12) -
The Prefect documentation. Python-first, far less ceremony, and dynamic behaviour that Airflow (Ch. 24)
-
The Prometheus documentation on "Instrumentation" and "Naming", and the Robust Perception blog (Ch. 12)
-
The property-based testing literature — Hypothesis (Python) and QuickCheck. Not usually applied to (Ch. 27)
-
The Protocol Buffers documentation on "Updating A Message Type." A different evolution model — (Ch. 17)
-
The Protocol Buffers language guide and encoding documentation (
protobuf.dev). Read the (Ch. 11) -
The Python
csvmodule documentation, particularly theDialectclass and theSniffer. The (Ch. 11) -
The Redpanda documentation, particularly on Kafka API compatibility and on what differs. Worth (Ch. 15)
-
The Singer and Airbyte connector specifications, recommended in Chapter 13 and relevant again: (Ch. 16)
-
The Singer specification (
singer.io) and the Meltano project's documentation. Singer defines (Ch. 13) -
The Slack, Stripe, and GitHub API documentation on pagination. Three well-designed APIs with (Ch. 16)
-
The Spark configuration reference. Worth skimming once so you know what exists. Most of it you (Ch. 21)
-
The Spark SQL Performance Tuning guide, in the official documentation. Short, and the single (Ch. 21)
-
The Spark Structured Streaming Programming Guide, particularly the output modes and watermarking (Ch. 29)
-
The Spark Web UI documentation. Undersold and rarely read. It explains what every column on the (Ch. 21)
-
The Tecton engineering blog, and Uber's Michelangelo papers. Michelangelo is where much of this (Ch. 32)
-
The Terraform documentation on
lifecycle,import,moved, andcheckblocks. Four short (Ch. 28) -
The TimescaleDB documentation. The strongest argument for §12.10's absorb-first position on (Ch. 12)
-
The TPC-H and TPC-DS specifications (
tpc.org). The standard analytical benchmarks. Worth (Ch. 8) -
The Unicode Consortium's material on encodings, and the "UTF-8 Everywhere" manifesto. Failure 6 (Ch. 11)
-
The US Census Bureau's material on their 2020 disclosure-avoidance system. The largest real (Ch. 31)
-
The VCR /
vcrpy/betamaxfamily of libraries. Record-and-replay HTTP for tests. Read the (Ch. 16) -
The Zstandard documentation and Yann Collet's benchmarks (
facebook.github.io/zstd/). The (Ch. 11) -
There is very little good writing on runbooks specifically, which is worth knowing so you do not (Ch. 26)
-
Tyler Akidau et al., Streaming Systems (O'Reilly, 2018), Chapters 1–3. Event time versus (Ch. 4)
-
Tyler Akidau, "Streaming 101" and "Streaming 102" (2015), on the O'Reilly Radar blog. The (Ch. 3)
-
Tyler Akidau, Slava Chernyak, and Reuven Lax, Streaming Systems (O'Reilly, 2018). The (Ch. 3, Ch. 20, Ch. 29)
-
Vaughn Vernon, Implementing Domain-Driven Design, the chapters on domain events and event (Ch. 36)
-
Vinoth Chandar et al. on Apache Hudi. Hudi's record-level indexing and incremental query model (Ch. 10)
-
Walk the register against one mart. Twenty-two items, an hour, and the output is a list rather (Ch. 23)
-
Warehouse vendors' Iceberg support announcements (Snowflake, BigQuery, Redshift). The (Ch. 10)
-
Wes McKinney, Python for Data Analysis (3rd ed., O'Reilly, 2022). By pandas' author, and the (Ch. 22)
-
Will Larson, Staff Engineer (2021), and An Elegant Puzzle (2019). The best available writing (Ch. 40)
-
Woodrow Hartzog, Privacy's Blueprint (Harvard, 2018). On design as a privacy decision, from the (Ch. 31)
-
Write the "not done" list. Exercise 38.10. If it is empty, you have not looked. (Ch. 38)
-
Write the falsifiable recommendation. Exercise 35.15. "Not yet" is not a recommendation; "not (Ch. 35)*
-
Write the five trade-off sentences. Exercise 39.6 — and notice which technologies you cannot (Ch. 39)
-
Write the five-year letter. Exercise 40.14, with a calendar reminder. (Ch. 40)
-
Write the manifest. Exercise 31.8. Fifteen entries minimum, including logs and vendors. The (Ch. 31)
-
Write your paging list, and show it to somebody who would be woken by it. Exercise 26.15. The (Ch. 26)
-
Yevgeniy Brikman, Terraform: Up & Running (3rd ed., O'Reilly, 2022). The practical companion. (Ch. 28)
-
Your cloud provider's pricing pages, for the three or four services that dominate your bill. (Ch. 33)
-
Your database's documentation on migrating stored procedures. Postgres, SQL Server, and Oracle all (Ch. 37)
-
Your database's documentation on window functions, specifically the frame clause.
ROWSversus (Ch. 39) -
Your engine's documentation on reading a query plan. Spark's
EXPLAIN FORMATTED, Snowflake's (Ch. 33) -
Your engine's documentation on window functions. For PostgreSQL this is the "Window Functions" (Ch. 18)
-
Your linter's custom-rule documentation, whatever it is. SQLFluff for SQL, a dbt package, or forty (Ch. 34)
-
Your own bill, at line-item granularity, exported to somewhere you can query it. AWS Cost and (Ch. 33)
-
Your own company's levelling rubric, read carefully and then set aside. §40.2 and Case Study 2: (Ch. 40)
-
Your own finance team. An hour, and it is worth more than anything on this list. Bring the four (Ch. 38)
-
Your own incident write-ups. §39.9's best source. If you have written postmortems (Chapter 26), (Ch. 39)
-
Your own legal or privacy counsel. Said in §31.6 and worth repeating: an hour with the person (Ch. 31)
-
Your own legal team. §30.6's point is that classification is a legal decision and engineering's job (Ch. 30)
-
Your own ticket system. §35.3's analysis is one query and an afternoon of categorization, and it (Ch. 35)
-
Your provider's budget and anomaly-detection services — AWS Cost Anomaly Detection, GCP budget (Ch. 33)
-
Your provider's documentation on which attributes are
ForceNew. This is not a general document (Ch. 28) -
Your provider's Savings Plan / Committed Use Discount calculator, run twice: once against current (Ch. 33)
-
Your warehouse's
ASOF JOINsupport. Snowflake, DuckDB, ClickHouse, and Databricks all have one (Ch. 32) -
Your warehouse's access history, again: Snowflake
ACCESS_HISTORY, BigQuery Data Access logs, (Ch. 37) -
Your warehouse's documentation on time travel, fail-safe, and backup retention. Snowflake's Time (Ch. 31)
-
Your warehouse's own metadata schema. Snowflake
ACCOUNT_USAGE(GRANTS_TO_ROLES, (Ch. 30) -
Your warehouse's usage views. Snowflake
ACCOUNT_USAGE.WAREHOUSE_METERING_HISTORYand (Ch. 33) -
Yu Malkov and Dmitry Yashunin, "Efficient and robust approximate nearest neighbor search using (Ch. 12)
-
Zhamak Dehghani's data mesh writing (covered properly in Chapter 35). Her diagnosis of why (Ch. 9)
-
Zhamak Dehghani, Data Mesh: Delivering Data-Driven Value at Scale (O'Reilly, 2022). The book, by (Ch. 35)