top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

Your Data Deserves a Profile - Here's How to Do It in Seconds

Apr 23, 2025
4 min read

Updated: May 28, 2025



Before diving deeper into Data Profiling, let's understand brief about pandas in python. If you are not beginner you can skip below content and jump to YData profiling


What is Pandas in Python?


Pandas is an open-source Python package that provides numerous tools for data analysis. The package comes with several data structures that can be used for many different data manipulation tasks.


Why do we need pandas in python?


Pandas is the most popular Python library that is used for data analysis. It provides highly optimized performance with back-end source code that is purely written in C or Python. It is very easy to write and understand.



import pandas as pd
df = pd.read_csv("subject-info.csv")
df.info()
df.head()
df.head(10)

Code Explanation:


  • import pandas as pd

    Import pandas library and gives it the alias pd, which is common convention. Pandas is essential for data manipulation and analysis in python

  • df = pd.read_csv("subject-info.csv")

    Reads CSV file named subject-info.csv into DataFrame called df. A DataFrame is like a table with rows and columns, similar to an Excel sheet. Make sure that CSV is in the current working directory, or provide full file path to avoid file not found errors.

  • .info() Displays the number rows and columns, data types, number of non-null values and memory usage

  • df.head()

    displays the first 5 rows of a DataFrame, This is useful for quickly previewing your dataset

  • df.head(10)

    To display first 10 rows, pass the number as an argument to .head()






Let's Start by Diving Deeper!


As a Data Analyst and data Enthusiast, I often find myself constantly exploring datasets. Performing Exploratory Data Analysis(EDA) is a critical step in understanding the structure and quality of any data.

However, I used to spend hours manually profiling data. which often slowed down my workflow. But what if i told you that you could reduce those hours to just seconds - and get ahead of your colleagues by working smarter?

Instead of profiling data using traditional SQL method, this entire process can be automated with the help of Pandas YData Profiling(formerly known as panadas-profiling)


"Automate Exploratory Data analysis with YData Profiling in Python"


Why YData Profiling?


Exploratory Data Analysis (EDA) is a crucial first step in any data science or analytics project. But let's be honest - manually generating statistics, Visualization, and correlation heatmaps can be time consuming. That's where YData profiling (formerly known as panadas-profiling) comes in a powerful Python library that automates EDA in just a few lines of code


What is YData Profiling?


YData Profiling is a fast and easy way to generate detailed reports from pandas DataFrames. it analyzes your data, find missing values, detects data types, calculate statistics, and even suggests warnings about potential data issues.


What is Exploratory Data Analysis(EDA)?


EDA is an analysis approach that identifies general pattern in the data. These patterns include outliers and features of the data that might be unexpected. EDA is an important first step in any data analysis


What happens During EDA?


Here are typical things done during EDA:


  • Understanding data structure : Checking rows, columns, data types, etc


  • Identifying missing values


  • Detecting outliers or anomalies


  • Exploring distributions of variables


  • Finding relationships and correlations between features


  • Validating assumptions for statistical modeling or machine learning



Why EDA Is Important?


EDA helps you:


  • Detect data quality issues early


  • Decide on feature engineering steps


  • choose appropriate models


  • Avoid misleading insights



Common Tools Used in EDA


  • Python Libraries: Pandas, Matplotlib, seaborn, plotly, ydata-profiling

  • R packages : ggplot2, dplyr, DataExplorer



Installation


You can install it via pip in Jupyter Notebook


!pip install ydata-profiling
!pip install pandas



How to use YData Profiling


Step 1: Import necessary Libraries


import pandas as pd
from ydata_profiling import ProfileReport


Step 2: Load a Dataset


Note: I have subject-info.csv, you can use any csv


df = pd.read_csv("subject-info.csv")


Step 3: Create the Profile Report


profile = ProfileReport(df, title = "profiling Report")


Step 4: Display Profile Report


profile.to_notebook_iframe()



Disclaimer: While displaying profile report in notebook if kernel crash follow below steps


Common Cause of kernel crashes with YData Profiling


  1. Large Dataset size


YData Profiling loads the entire DataFrame into memory and performs deep analysis. If your dataset is too large, It can:


  • Consume all available RAM


  • Trigger memory overflows


  • Cause the notebook kernel to crash or freeze


Fixes:

  • Use a sample of the dataset for profiling


  • Filter unnecessary columns or rows before profiling


2.      Jupyter Notebook Rendering Limitations

 

The output of YData Profiling (especially .to_notebbok_iframe() can be very large HTML, which Jupyter sometimes struggles to render.

 

Fixes:

  • Use .to_file() instead, Then open the report in a browser instead of inline


To open .html file in browser

  • Go to the Home in your jupyter notebook

  • You will find subject_info_report.html generated.

  • Now click on it to open in the browser

  • You can also upload by selecting file and clicking on upload button



profile.to_file("subject_info_report.html")



  • Or use .to_widgets() of you’re in Jupyetr Lab


Profile.to_widgets(“subject_info_report.html”)




  1. System resource Limitations


    If you're running on a local machine with limited RAM( e.g, <8GB), or using browser based platform(like Google Colab), large profiling tasks may exceed the limits.


    Fixes:


    • Upgrade your environment (RAM, cloud instance).


    • Try running a Python script instead of a notebook.


    • Use lightweight profiling alternatives like sweetviz or dtale




    Final Output report looks like this:

    • First image represent the Overview of dataset statistics, such as Number of column(variables), Number of rows(Observations) , Null values, Duplicate rows, Column data types etc.

    • Second image represent the Variables(column) "Age" details like Distinct, Mean , Max, Mini and missing values etc. You can select other columns in the Variable dropdown to get insights

    • Third Image represent the interaction of "Age" column with "ID" column

    • Fourth image represent the correlation between other columns With "Age" column














  • Conclusion:


    Whether you're a data scientist, analyst, or ML engineer, YData Profiling can save you hours in your data exploration phase. In a world where time = insights = money, why not let your code do the heavy lifting?


Thank You!












 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page