top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

Beyond the Numbers: How Smarter Feature Selection Can Supercharge Diabetes Predictions

Apr 30, 2025
4 min read

Why Feature Selection Actually Matters (More Than You Think)


When you dive into machine learning for healthcare, it’s easy to think:More data = better predictions, right?


Not exactly.


It’s more like trying to diagnose a patient while drowning in random lab results. Instead of helping, too much irrelevant data overwhelms both the doctor and the model — and you end up prioritizing something like skin thickness over glucose levels when predicting diabetes. (Seriously.)


In healthcare, where trust and interpretability matter deeply, it’s not about using more data — it’s about using the right data.


Each feature tells a story:


  • BMI → lifestyle risks

  • Glucose → real-time metabolic snapshot

  • Insulin → pancreas performance

  • Age → the silent risk factor


Some features? Honestly, just noise. Feature selection is how we decide what’s meaningful — and what’s just clutter.


Meet the Dataset: Pima Indians Diabetes


For this project, I worked with the well-known Pima Indians Diabetes Dataset, a classic dataset often used for healthcare ML experimentation.


Some quick facts:


  • 768 female patients (all 21+ years old)

  • All of Pima Indian heritage

  • Features include:

    • Pregnancies

    • Glucose

    • Blood Pressure

    • Skin Thickness

    • Insulin

    • BMI

    • Diabetes Pedigree Function

    • Age

    • Outcome (0 = no diabetes, 1 = diabetes)

The goal:

Build a smarter diabetes prediction model by choosing only the features that actually matter.


How I Approached Feature Selection


I explored three different techniques to trim down unnecessary features:


1. Filter Method (Correlation Heatmap)


import pandas as pd

import seaborn as sns

import matplotlib.pyplot as plt


# Load the dataset

dataFrame = pd.read_csv('diabetes.csv')

dataFrame.head()


# Plot the correlation heatmap

corr = dataFrame.corr()

sns.heatmap(corr, annot=True, cmap='plasma')

plt.title('Feature Correlation with Outcome')




Observations:


  • Glucose had the strongest correlation with diabetes outcome.

  • BMI, Age, and Diabetes Pedigree Function also showed some connection.

  • Skin Thickness barely correlated at all.


2. Wrapper Method (Recursive Feature Elimination)


Next, I used RFE to automatically pick the best features by iteratively removing the weakest ones:


from sklearn.model_selection import train_test_split

from sklearn.ensemble import RandomForestClassifier

from sklearn.feature_selection import RFE


X = dataFrame.drop('Outcome', axis=1)

y = dataFrame['Outcome']


model = RandomForestClassifier()

rfe = RFE(estimator=model, n_features_to_select=5)

rfe.fit(X, y)


selected_features_rfe = X.columns[rfe.support_]

print("Top 5 Features by RFE:", selected_features_rfe.tolist())



Results:

Top 5 Features by RFE: ['Glucose', 'BloodPressure', 'BMI', 'DiabetesPedigreeFunction', 'Age']
  • RFE chose: Glucose, BMI, Age, Insulin, and Diabetes Pedigree Function.


3. Embedded Method (Lasso Regularization)


Finally, I used LassoCV to let the model decide which features to zero out automatically:


from sklearn.linear_model import LassoCV

from sklearn.preprocessing import StandardScaler


scaler = StandardScaler()

X_scaled = scaler.fit_transform(X)


lasso = LassoCV(cv=5)

lasso.fit(X_scaled, y)


coef = pd.Series(lasso.coef_, index=X.columns)

selected_features_lasso = coef[coef != 0].index.tolist()

print("Selected features by Lasso:", selected_features_lasso)


Lasso’s ruthless picks:

Selected features by Lasso: ['Pregnancies', 'Glucose', 'BloodPressure', 'Insulin', 'BMI', 'DiabetesPedigreeFunction', 'Age']

  • Glucose

  • BMI

  • Age

(Lasso kept it tight and efficient.)


Building Two Competing Models


At this point, I built and trained two Logistic Regression models:

  • One using all features.

  • One using only the top 3 features selected by Lasso.


Here’s the training and evaluation:


from sklearn.model_selection import train_test_split # You need this!

from sklearn.linear_model import LogisticRegression

from sklearn.metrics import accuracy_score, f1_score, roc_auc_score



# Full feature model

modelrecord = LogisticRegression(max_iter=1000)

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)


modelrecord.fit(X_train, y_train)

y_pred = modelrecord.predict(X_test)


acc_base = accuracy_score(y_test, y_pred)

f1_base = f1_score(y_test, y_pred)

roc_base = roc_auc_score(y_test, y_pred)


# Lasso-selected feature model


X_small = dataFrame[['Glucose', 'BMI', 'Age']] # Make sure dataFrame is correct

X_small_train, X_test_s, y_train_s, y_test_s = train_test_split(X_small, y, test_size=0.2, random_state=4)


modelrecord.fit(X_small_train, y_train_s)

y_pred_s = modelrecord.predict(X_test_s)


acc_sel = accuracy_score(y_test_s, y_pred_s)

f1_sel = f1_score(y_test_s, y_pred_s)

roc_sel = roc_auc_score(y_test_s, y_pred_s)


Visualizing Performance Differences


import pandas as pd

import matplotlib.cm as cm

import matplotlib.pyplot as plt

import matplotlib.colors as mcolors



# Metrics dynamic grabbing

metricNames = ['Accuracy', 'F1 Score', 'ROC AUC']


base_scores = [acc_base, f1_base, roc_base]


selected_scores = [acc_sel, f1_sel, roc_sel]


improvements = [s - b for s, b in zip(selected_scores, base_scores)]


# Build DataFrame

metrics_df = pd.DataFrame({

'Metric': metricNames,

'Improvement': improvements

})


# Plot

fig, ax = plt.subplots(figsize=(10, 6))


colors = ['green' if imp >= 0 else 'tomato' for imp in metrics_df['Improvement']] #color

bars = ax.bar( metrics_df['Metric'], metrics_df['Improvement'], color=colors, edgecolor='brown') #bars


# Annotate bars

for bar in bars:

height = bar.get_height()

ax.annotate(f'{height:+.2%}',

xy=(bar.get_x() + bar.get_width()/2, height),

xytext=(0, 5 if height > 0 else -15),

textcoords='offset points',

ha='center', va='bottom' if height > 0 else 'top',

fontsize=10)


# Style

ax.set_title('Real Model Improvement (Green = Better, Red = Worse)', fontsize=17, pad=20)

ax.axhline(0, color='gray', linewidth=0.8)

ax.set_ylabel('Improvement (%)')

ax.grid(axis='y', linestyle='--', alpha=0.6)

ax.set_facecolor('whitesmoke')

fig.patch.set_facecolor('white')

ax.spines['top'].set_visible(False)

ax.spines['right'].set_visible(False)


plt.tight_layout()





What the chart shows:

All bars are green → All metrics improved when using the Lasso-selected features (Glucose, BMI, Age).

  • +5.19% in Accuracy, +4.99% in F1 Score, and +4.70% in ROC AUC — solid gains.

  • The new color scheme (green = better) makes your point visually clear and optimistic.


Using just 3 features selected by Lasso, the model actually performed better across the board.


Accuracy, F1 Score, and ROC AUC all improved, especially accuracy (+5.19%).


The green bars below highlight these gains — sometimes, less really is more.


Real-World Lessons from This Experiment


This taught me a key truth in real-world machine learning:

  • Simpler models are easier to explain, faster to train, and more practical.

  • Reducing features improves transparency — but you might sacrifice a tiny bit of predictive performance.

  • Healthcare ML is about trust — doctors prefer slightly less accurate models that are interpretable over black-box ones.

Interpretability beats tiny accuracy gains — especially when lives are involved.


Where I'd Take This Next


If I had more time, I would:

  • Try hybrid feature selection (combining filter + wrapper + embedded).

  • Explore interaction features (like Glucose/Age ratios).

  • Run the same experiments on larger medical datasets.

Feature selection isn't about chasing higher numbers — it's about telling a better story with your data.


 Final Thoughts


Beyond the numbers, smart feature selection is about making models that people can trust.

In healthcare, where every percentage point matters but every explanation matters even more, feature selection isn't a side quest — it’s the main game.

 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page