top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

Data Cleaning with Python for Better Visualizations

Nov 21, 2025
2 min read

Why Data Cleaning is Important Before Visualization

When we work with data, the first thing most of us want to do is make a chart. But here’s the truth: if the data is messy, the chart will also be messy — and it might even tell the wrong story.

This is where data cleaning comes in. Data cleaning means fixing problems in your dataset before you create visualizations. It’s like preparing your ingredients before cooking. If the ingredients are spoiled, the dish won’t taste good

Common Problems in Raw Data

Here are some issues you’ll often see in raw datasets:

·       Missing values – blank cells or “NaN” values.

·       Duplicates – the same row repeated more than once.

·       Wrong data types – numbers stored as text, or dates written as strings.

·       Outliers – extreme values that don’t look right.

 

Example: Customer Email Data

Let’s take the customer table. Suppose we check emails:

SELECT customer_id, first_name, last_name, email

FROM customer

LIMIT 10;

 

We might notice:

·       Some customers don’t have an email (NULL).

·       A few emails are duplicated.

·       Some emails look invalid (like “test@abc” instead of “test@abc.com"

 

Key Takeaway

Data cleaning may not be the most exciting part of data analysis, but it’s the most important step. Without clean data, even the best-looking chart can mislead us.

  

Handling Missing Data in Python

Missing data means some values in a dataset are not available or not recorded. They usually appear as:

·       Blank cells → empty boxes in a table or Excel sheet.

·       NaN (Not a Number) → a marker used in Python, R, or databases for missing numeric values.

·       NULL / None → common in SQL or programming to show missing information.

 

How to fix missing data:

 

·       Remove missing rows (dropna()).

Example : df.dropna()

 

·       Fill missing values (fillna()).

Example : df.fillna(0)

 

·       Replace with mean/median for numbers.

Example: df["Age"].fillna(df["Age"].mean(), inplace=True)

 

 Dealing with Duplicates and Outliers

 

What are duplicates?

Duplicates are rows that appear more than once in your dataset. They don’t add value and can make your analysis wrong.

How to find duplicates in pandas: df.duplicated()

Count duplicates: df.duplicated().sum()

Remove duplicates: df.drop_duplicates(inplace=True)

 

What are outliers?Outliers are values that are much higher or lower than the rest of the data. They may be real, or they may be errors.

Most ages are around 20–40, but one row says 200. That’s probably a mistake.

How to detect outliers: df["Age"].describe()

 

Cleaning Data Types and Strings

Sometimes our data looks fine but is actually in the wrong format. For example:

·       Dates are stored as text (string).

·       Numbers are stored as strings instead of integers.

·       Text values may have extra spaces or different cases (New York, new york, NEW YORK).

If we don’t fix this, our analysis and plots can give the wrong result.

 

Standardize Data Formats and Inconsistencies:

 

Even after fixing missing values, duplicates, and outliers, our dataset can still have inconsistent formats. These don’t always look like “errors,” but they can create problems when we analyze or visualize data.

 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page