Beyond the Numbers: How Smarter Feature Selection Can Supercharge Diabetes Predictions
Why Feature Selection Actually Matters (More Than You Think)
When you dive into machine learning for healthcare, it’s easy to think:More data = better predictions, right?
Not exactly.
It’s more like trying to diagnose a patient while drowning in random lab results. Instead of helping, too much irrelevant data overwhelms both the doctor and the model — and you end up prioritizing something like skin thickness over glucose levels when predicting diabetes. (Seriously.)
In healthcare, where trust and interpretability matter deeply, it’s not about using more data — it’s about using the right data.
Each feature tells a story:
BMI → lifestyle risks
Glucose → real-time metabolic snapshot
Insulin → pancreas performance
Age → the silent risk factor
Some features? Honestly, just noise. Feature selection is how we decide what’s meaningful — and what’s just clutter.
Meet the Dataset: Pima Indians Diabetes
For this project, I worked with the well-known Pima Indians Diabetes Dataset, a classic dataset often used for healthcare ML experimentation.
Some quick facts:
768 female patients (all 21+ years old)
All of Pima Indian heritage
Features include:
Pregnancies
Glucose
Blood Pressure
Skin Thickness
Insulin
BMI
Diabetes Pedigree Function
Age
Outcome (0 = no diabetes, 1 = diabetes)
The goal:
Build a smarter diabetes prediction model by choosing only the features that actually matter.
How I Approached Feature Selection
I explored three different techniques to trim down unnecessary features:
1. Filter Method (Correlation Heatmap)
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
# Load the dataset
dataFrame = pd.read_csv('diabetes.csv')
dataFrame.head()
# Plot the correlation heatmap
corr = dataFrame.corr()
sns.heatmap(corr, annot=True, cmap='plasma')
plt.title('Feature Correlation with Outcome')
plt.show()

Observations:
Glucose had the strongest correlation with diabetes outcome.
BMI, Age, and Diabetes Pedigree Function also showed some connection.
Skin Thickness barely correlated at all.
2. Wrapper Method (Recursive Feature Elimination)
Next, I used RFE to automatically pick the best features by iteratively removing the weakest ones:
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.feature_selection import RFE
X = dataFrame.drop('Outcome', axis=1)
y = dataFrame['Outcome']
model = RandomForestClassifier()
rfe = RFE(estimator=model, n_features_to_select=5)
rfe.fit(X, y)
selected_features_rfe = X.columns[rfe.support_]
print("Top 5 Features by RFE:", selected_features_rfe.tolist())
Results:
Top 5 Features by RFE: ['Glucose', 'BloodPressure', 'BMI', 'DiabetesPedigreeFunction', 'Age']RFE chose: Glucose, BMI, Age, Insulin, and Diabetes Pedigree Function.
3. Embedded Method (Lasso Regularization)
Finally, I used LassoCV to let the model decide which features to zero out automatically:
from sklearn.linear_model import LassoCV
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
lasso = LassoCV(cv=5)
lasso.fit(X_scaled, y)
coef = pd.Series(lasso.coef_, index=X.columns)
selected_features_lasso = coef[coef != 0].index.tolist()
print("Selected features by Lasso:", selected_features_lasso)
Lasso’s ruthless picks:
Selected features by Lasso: ['Pregnancies', 'Glucose', 'BloodPressure', 'Insulin', 'BMI', 'DiabetesPedigreeFunction', 'Age']Glucose
BMI
Age
(Lasso kept it tight and efficient.)
Building Two Competing Models
At this point, I built and trained two Logistic Regression models:
One using all features.
One using only the top 3 features selected by Lasso.
Here’s the training and evaluation:
from sklearn.model_selection import train_test_split # You need this!
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score, roc_auc_score
# Full feature model
modelrecord = LogisticRegression(max_iter=1000)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
modelrecord.fit(X_train, y_train)
y_pred = modelrecord.predict(X_test)
acc_base = accuracy_score(y_test, y_pred)
f1_base = f1_score(y_test, y_pred)
roc_base = roc_auc_score(y_test, y_pred)
# Lasso-selected feature model
X_small = dataFrame[['Glucose', 'BMI', 'Age']] # Make sure dataFrame is correct
X_small_train, X_test_s, y_train_s, y_test_s = train_test_split(X_small, y, test_size=0.2, random_state=4)
modelrecord.fit(X_small_train, y_train_s)
y_pred_s = modelrecord.predict(X_test_s)
acc_sel = accuracy_score(y_test_s, y_pred_s)
f1_sel = f1_score(y_test_s, y_pred_s)
roc_sel = roc_auc_score(y_test_s, y_pred_s)
Visualizing Performance Differences
import pandas as pd
import matplotlib.cm as cm
import matplotlib.pyplot as plt
import matplotlib.colors as mcolors
# Metrics dynamic grabbing
metricNames = ['Accuracy', 'F1 Score', 'ROC AUC']
base_scores = [acc_base, f1_base, roc_base]
selected_scores = [acc_sel, f1_sel, roc_sel]
improvements = [s - b for s, b in zip(selected_scores, base_scores)]
# Build DataFrame
metrics_df = pd.DataFrame({
'Metric': metricNames,
'Improvement': improvements
})
# Plot
fig, ax = plt.subplots(figsize=(10, 6))
colors = ['green' if imp >= 0 else 'tomato' for imp in metrics_df['Improvement']] #color
bars = ax.bar( metrics_df['Metric'], metrics_df['Improvement'], color=colors, edgecolor='brown') #bars
# Annotate bars
for bar in bars:
height = bar.get_height()
ax.annotate(f'{height:+.2%}',
xy=(bar.get_x() + bar.get_width()/2, height),
xytext=(0, 5 if height > 0 else -15),
textcoords='offset points',
ha='center', va='bottom' if height > 0 else 'top',
fontsize=10)
# Style
ax.set_title('Real Model Improvement (Green = Better, Red = Worse)', fontsize=17, pad=20)
ax.axhline(0, color='gray', linewidth=0.8)
ax.set_ylabel('Improvement (%)')
ax.grid(axis='y', linestyle='--', alpha=0.6)
ax.set_facecolor('whitesmoke')
fig.patch.set_facecolor('white')
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
plt.tight_layout()
plt.show()

What the chart shows:
All bars are green → All metrics improved when using the Lasso-selected features (Glucose, BMI, Age).
+5.19% in Accuracy, +4.99% in F1 Score, and +4.70% in ROC AUC — solid gains.
The new color scheme (green = better) makes your point visually clear and optimistic.
Using just 3 features selected by Lasso, the model actually performed better across the board.
Accuracy, F1 Score, and ROC AUC all improved, especially accuracy (+5.19%).
The green bars below highlight these gains — sometimes, less really is more.
Real-World Lessons from This Experiment
This taught me a key truth in real-world machine learning:
Simpler models are easier to explain, faster to train, and more practical.
Reducing features improves transparency — but you might sacrifice a tiny bit of predictive performance.
Healthcare ML is about trust — doctors prefer slightly less accurate models that are interpretable over black-box ones.
Interpretability beats tiny accuracy gains — especially when lives are involved.
Where I'd Take This Next
If I had more time, I would:
Try hybrid feature selection (combining filter + wrapper + embedded).
Explore interaction features (like Glucose/Age ratios).
Run the same experiments on larger medical datasets.
Feature selection isn't about chasing higher numbers — it's about telling a better story with your data.
Final Thoughts
Beyond the numbers, smart feature selection is about making models that people can trust.
In healthcare, where every percentage point matters but every explanation matters even more, feature selection isn't a side quest — it’s the main game.


