Part 2: Exploring Data
Before you can analyze data, you need to see it.
That might sound obvious, but it's one of the most overlooked steps in statistics. People get excited about hypothesis tests and p-values and regression models, and they skip right past the part where you just... look at the data. Make some graphs. Calculate some summaries. Ask basic questions like "what's the typical value?" and "how spread out is this?" and "are there any weird outliers that might mess everything up?"
This is called exploratory data analysis, and it's not just a warm-up act. It's where you discover the story the data is trying to tell you. It's where you catch errors before they derail your conclusions. It's where you develop the intuition that separates someone who runs tests mechanically from someone who actually understands what they're doing.
Here's a useful analogy: imagine you're a detective arriving at a crime scene. You wouldn't immediately start running DNA tests. First, you'd walk around, observe, take notes, look for things that seem out of place. That's what this part of the book is about — walking around your data before you start testing theories.
Chapter 5 is about visualization. You'll learn to create and interpret histograms, bar charts, box plots, and scatterplots. More importantly, you'll learn to read these graphs — to see the shape of a distribution, spot outliers, and choose the right graph for the right situation. A well-chosen graph can reveal patterns that no table of numbers ever would.
Chapter 6 puts numbers on what your eyes have already noticed. You'll calculate measures of center (mean, median, mode) and spread (standard deviation, IQR, range). You'll learn why the mean and standard deviation are the power couple of statistics, but also why the median sometimes tells a more honest story. And you'll meet the Empirical Rule — the 68-95-99.7 guideline that will become one of your most-used mental tools.
Chapter 7 is the chapter nobody wants to do but everybody needs. Real data is messy. It has missing values, inconsistencies, typos, and formatting nightmares. Data wrangling — cleaning and preparing data for analysis — is where most analysts spend the majority of their time. You'll learn to handle missing data, transform variables, and document every decision you make so your work is reproducible.
By the end of Part 2, you'll be able to take a raw dataset and turn it into something you genuinely understand. You'll know its shape, its center, its spread, its quirks. And that understanding will be the foundation for everything that follows.