Data Cleaning with Python for Better Visualizations
Why Data Cleaning is Important Before Visualization
When we work with data, the first thing most of us want to do is make a chart. But here’s the truth: if the data is messy, the chart will also be messy — and it might even tell the wrong story.
This is where data cleaning comes in. Data cleaning means fixing problems in your dataset before you create visualizations. It’s like preparing your ingredients before cooking. If the ingredients are spoiled, the dish won’t taste good
Common Problems in Raw Data
Here are some issues you’ll often see in raw datasets:
· Missing values – blank cells or “NaN” values.
· Duplicates – the same row repeated more than once.
· Wrong data types – numbers stored as text, or dates written as strings.
· Outliers – extreme values that don’t look right.
Example: Customer Email Data
Let’s take the customer table. Suppose we check emails:
SELECT customer_id, first_name, last_name, email
FROM customer
LIMIT 10;
We might notice:
· Some customers don’t have an email (NULL).
· A few emails are duplicated.
· Some emails look invalid (like “test@abc” instead of “test@abc.com"
Key Takeaway
Data cleaning may not be the most exciting part of data analysis, but it’s the most important step. Without clean data, even the best-looking chart can mislead us.
Handling Missing Data in Python
Missing data means some values in a dataset are not available or not recorded. They usually appear as:
· Blank cells → empty boxes in a table or Excel sheet.
· NaN (Not a Number) → a marker used in Python, R, or databases for missing numeric values.
· NULL / None → common in SQL or programming to show missing information.
How to fix missing data:
· Remove missing rows (dropna()).
Example : df.dropna()
· Fill missing values (fillna()).
Example : df.fillna(0)
· Replace with mean/median for numbers.
Example: df["Age"].fillna(df["Age"].mean(), inplace=True)
Dealing with Duplicates and Outliers
What are duplicates?
Duplicates are rows that appear more than once in your dataset. They don’t add value and can make your analysis wrong.
How to find duplicates in pandas: df.duplicated()
Count duplicates: df.duplicated().sum()
Remove duplicates: df.drop_duplicates(inplace=True)
What are outliers?Outliers are values that are much higher or lower than the rest of the data. They may be real, or they may be errors.
Most ages are around 20–40, but one row says 200. That’s probably a mistake.
How to detect outliers: df["Age"].describe()
Cleaning Data Types and Strings
Sometimes our data looks fine but is actually in the wrong format. For example:
· Dates are stored as text (string).
· Numbers are stored as strings instead of integers.
· Text values may have extra spaces or different cases (New York, new york, NEW YORK).
If we don’t fix this, our analysis and plots can give the wrong result.
Standardize Data Formats and Inconsistencies:
Even after fixing missing values, duplicates, and outliers, our dataset can still have inconsistent formats. These don’t always look like “errors,” but they can create problems when we analyze or visualize data.


