DataSets + Visualization Tools + ML Model Training: A Step by Step Guide Towards Building a Diabetes Prediction Model
As a student of NumpyNinja, my aim is to apply the concepts that I have been learning in my classes in building real-life applications.
As per a World Health Organization (WHO) survey, in India, there are estimated 77 million people above the age of 18 years who are suffering from diabetes (type 2), and nearly 25 million are prediabetics (at a higher risk of developing diabetes in the near future) [1]. More than 50% of people are unaware of their diabetic status, which leads to health complications if not detected and treated early. Adults with diabetes have a two-to three-fold increased risk of heart attacks and strokes. To make matters worse, access to healthcare is not readily available to most of people suffering from diabetes.
Under the circumstances, it would be very helpful to build a machine learning model that can predict if a patient is diabetic or pre-diabetic based on the patient’s health information when access to doctors and primary healthcare is unavailable.
Goal
Build a machine learning model that can predict the probability of a patient being diabetic given the patient’s health data.
How Can We Achieve Our Goal?
Building a machine learning model involves the following steps:
Raw dataset collection
Dataset cleaning
Understanding the important features in the data to use for model training
Choice of ML model architecture to use for training
DataSet Collection
The first step in any ML model training is data collection. As part of the NumpyNinja coursework, we were already provided access to the diabetic dataset. So we did not need to perform the data collection step.
DataSet Cleaning
The next step involves cleaning the raw dataset. Most of the time, when data is collected, there are discrepancies in the table of values that need to be addressed. For instance, the diabetic dataset in its original form has 121 rows and 145 columns, which correspond to 121 patient entries with 145 different fields.
Some of the observations made on the raw dataset:
Fields with trailing _x characters
Missing or null values
Typos in text
Duplicate records
Variation in test names
Non-standard units used for test results
Inconsistency across record entries
How did I address the above issues and help clean up the dataset?
1. Remove (_x) from column names and blank values. Rename the column to "Patient ID." Merge "Patient ID" and "Visit" columns to create "PatientVisit_ID
2. Renaming Column Name to' HYPERTENSION' AND Replacing values HTN as HYPERTENSION and ntn as NO HYPERTENSION, Text Formatted to UPPERCASE.
3. Change the datatype to Whole Number. Replace 57 null values with zero (no duration for control groups).
4. Created Custom Column Race2 =if [PatientID] = "S0301" then 4 else if [PatientID] = "S0544" then 3 else [Race], so that S0301 is replaced with Asian and S0544 is replaced with Latino for both visit ID 2 and Now created a New conditional Column as Race Category and given option for 1- White, 2 -Black or African American, 3- Latino, and 4- Asian and Others for Unknown and Changed DataType to Text
5. Rounded off to 2 Decimals and Height(m) can be converted to Height (cm) Height(cm)= Height(m) * 100
Visualization Tools: Power BI
Once we have a cleaned dataset, the next step is to understand the underlying patterns in the data before we start training a ML model. Our dataset has 145 fields, out of which almost 100+ fields can be thought of as contributing to the final outcome of whether a patient is diabetic or not. However, it is important to note that not all of the 100+ fields contribute in a similar manner to the final model prediction, and we need to gain insights into which fields (or, in other words, “features”) contribute prominently to the model outcome.
This is where visualization tools play a very important role. In our coursework, we used Power BI to analyze the diabetic dataset, and from the analysis, we can see that there is a strong correlation between some of the features and whether a patient is diabetic or not. For the sake of simplicity and first pass implementation, we will use these features in training our ML model and omit the rest from our training data.
To make the insights clearer, I created the following visualization in Power BI.
We create new columns, Diabetes-category and BMI-category
Daibeties-category = SWITCH(
TRUE(),
Lab[Hb A1C%] <= 5.7, "Normal",
Lab[Hb A1C%] <= 6.4 && Lab[Hb A1C%] > 5.7, "Predaibetes",
Lab[Hb A1C%] >=6.5, "Daibetis"
)BMI-Category =
IF(Demographics[BMI2] < 18.5, " Under Weight",
IF(Demographics[BMI2] >=18.5 && Demographics[BMI2]<24.99," Normal Weight",
IF(Demographics[BMI2]>=25 && Demographics[BMI2]<29.99,"Over weight ",
IF(Demographics[BMI2]>=30,"Obesity"))))

Now, let’s dive into a Power BI visual that helped me understand the patterns better.
As BMI increases, the risk of developing, diabetes also increases.People who are overweight or obese are much more likely to become insulin resistant, which leads to high blood sugar levels
Fasting blood sugar is directly related to diabetes.
Patients with diabetes tend to have elevated fasting glucose levels, whereas patients without diabetes usually have normal levels.
ML Model Training
Once we have identified the significant features that have a high correlation (either positive or negative) to the model prediction, it’s now time to build the ML model. This part is something I am in the process of exploring and is yet to be implemented. Some of the key concepts related to training an ML model that I would need to learn can be enumerated as follows [2] (resource on the web):
Imagine a dataset of patients, where each patient is labeled as either "Diabetic" (1) or "Not Diabetic" (0).
1. Data Preparation:
Feature Extraction: Extract relevant features from the patient data, such as the BMI, fasting glucose, HbA1c, weight, or age from the PowerBI visualization tools above.
Splitting Data: Divide the dataset into a training set (e.g., 80% of the data) and a test set (e.g., 20% of the data).
Label Encoding: Ensure that the class labels are in a numerical format (e.g., 0 and 1).
2. Model Training:
Choose an Algorithm: Select an appropriate binary classification algorithm, such as logistic regression.
Train the Model: Use the training data to fit the chosen model to the data, learning the relationships between the features and the class labels.
Hyperparameter Tuning: Optimize the model's performance by adjusting its hyperparameters (e.g., regularization strength).
3. Model Evaluation:
Predict on Test Data: Use the trained model to predict the class labels for the unseen data in the test set.
Evaluate Performance: Measure the model's accuracy, precision, recall, and F1-score to assess its performance.
4. Prediction:
New Data: Once the model is trained and evaluated, it can be used to predict the class label for new, unseen data points.
References


