Before you build a model, write a report, or make a recommendation, you need to understand what you're working with. That's what exploratory data analysis (EDA) is for.

EDA is the practice of examining a dataset before any formal modelling or hypothesis testing begins. The goal is to develop an intuition for the data: its shape, its quality, its quirks. What you learn in EDA determines every decision that follows.

Here are five steps to work through every time you encounter a new dataset.

Step 1: Profile the dataset

Start with the basics. How many rows? How many columns? What data type is each column — numerical, categorical, date, text? Are there any columns you don't understand?

A data profile gives you the shape of the dataset. It tells you immediately if a column is storing numbers as text (a common import issue), if dates are in inconsistent formats, or if a column that should have unique values has duplicates.

Look at the first and last few rows. Scroll through a sample of the middle. You're not analysing yet — you're getting familiar.

Step 2: Examine distributions

For numerical columns, calculate summary statistics and create histograms. The statistics to look at first: minimum, maximum, mean, median, and standard deviation.

Is the distribution symmetric or skewed? A heavily right-skewed distribution (lots of small values, a few very large ones) is common for revenue and order size data. A roughly normal (bell-shaped) distribution is common for physical measurements. A bimodal distribution (two peaks) often indicates there are two distinct groups in the data that shouldn't be mixed.

For categorical columns, look at frequency counts. How many distinct values are there? Is one value dominant? Are there typos or inconsistencies — "United States" vs "US" vs "USA" — that need to be normalised?

Step 3: Identify missing values

Missing data is nearly universal in real-world datasets. The question isn't whether values are missing, but how many, in which columns, and whether the absence is random or systematic.

A column with 1% missing values can usually be filled with a median or dropped with little impact. A column with 40% missing values is a different problem — it may need to be excluded from analysis entirely, or its missingness may itself be informative.

Ask why values are missing. Survey questions that were optional will have different missing patterns than system fields that should always be populated. Missing values in one column correlated with missing values in another can reveal important structure in the data.

Step 4: Look for correlations

Once you understand each column individually, look at how they relate to each other.

A correlation matrix shows how strongly pairs of numerical variables move together. A correlation close to +1 means they rise and fall together. A correlation close to -1 means one rises when the other falls. A correlation near 0 means they're largely independent.

High correlations are useful — they can simplify models by showing which variables carry the same information. But correlation doesn't imply causation, and it can also flag data quality issues (a column that's essentially a calculated copy of another).

Scatter plots between pairs of variables reveal non-linear relationships that correlation coefficients miss. If revenue and marketing spend have a correlation of 0.3 overall but a scatter plot shows a strong relationship for spends below £10,000 and no relationship above, that's analytically important and numerically invisible.

Step 5: Investigate outliers

Outliers deserve attention before anything else in the analysis happens. A single outlier can completely distort a mean, a regression, or a chart scale.

Start with the extremes: What are the top 5 and bottom 5 values in each numerical column? Do they make sense? A customer age of 247 is a data entry error. A transaction amount of £500,000 in a dataset where the average is £80 might be a legitimate bulk order — or a fraud signal.

The standard statistical definition of an outlier is a value more than 1.5 × IQR beyond the quartiles (where IQR is the interquartile range, Q3 minus Q1). This is useful as a starting point but not the final word. Domain knowledge matters more than a formula.

Decide on each outlier individually: fix it, remove it, flag it, or leave it in with a note. Don't silently drop outliers just because they're inconvenient.

Turning EDA into action

EDA isn't just a checklist — it's a thinking process. As you work through these steps, write down what you notice. Surprising findings, questions you can't answer, decisions you made about data quality. These notes become the methodology section of any report you produce.

Good EDA also reshapes your original question. The patterns you find often reveal that you were asking a slightly wrong question, or that there's a more interesting question hiding in the data you didn't know to ask.

Tools like Amridata automate the mechanical parts of EDA — profiling columns, detecting missing values, calculating statistics, flagging outliers — so you can spend your time on the thinking rather than the counting.