top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

Turning Sensor Noise into Maternal Knowledge: A Deep Dive into Health Data Engineering

Jan 13
3 min read

When it comes to maternal health, the stakes are as high as they can be. Monitoring a pregnancy today isn't just about monthly check-ups; it’s about the silent, continuous stream of data flowing from wearables 24/7.

But here is the reality no one tells you: raw sensor data is a disaster.

In this project, I integrated messy data from wearables (EDA, Heart Rate) and manual logs (Food intake) into a structured PostgreSQL database. My goal was to turn "noise" into "knowledge" to manage blood sugar and stress during pregnancy.


1. The Challenge: Why We Couldn't Just "Use" the Data

When I first opened the datasets, I found a jigsaw puzzle with missing pieces. The data came from 16 individuals, each generating roughly 4,000 rows of high-frequency measurements every few seconds.

  • The Wearables: EDA (stress) and Heart Rate sensors spat out thousands of rows with timestamps that looked like gibberish.

  • The Food Logs: Being manual, entries were inconsistent—people wrote "a bit of rice" or "half a cup."

  • The Fragmentation: The tables didn't "talk" to each other. You couldn't see how a meal at noon caused a glucose spike at 1:00 PM.


To solve this, I implemented a robust ETL (Extract, Transform, Load) process and a Star Schema architecture. By setting up Referential Integrity, I ensured that every heart rate spike or glucose reading was tied to a validated patient ID, preventing "orphan data."

2. Wrangling Time and Space

In healthcare technology, time is the most critical variable. If a mother’s heart rate jumps, we need to know the context: Was she sleeping or exercising?

Normalizing Timestamps

Raw data arrived with "Combined Timestamps"—massive strings of text. To fix this, I split them into dedicated DATE and TIME columns. This simple change transformed the performance; instead of the computer "chopping strings" for every query, users can now filter morning trends instantly.

The Food Log Challenge: Converting Speech to Math

Humans are inconsistent. I wrote extensive UPDATE statements to standardize entries (e.g., converting "1/2" to 0.5).

Solving for the "Missing Fat":

When users forgot to log fat content, I used the Atwater System to calculate it programmatically:


$$Total Fat = \frac{Calories - (Protein \times 4 + Total Carb \times 4)}{9}$$

3. Interpreting the Signal: The 5-Step Analysis

To turn this cleaned data into clinical insight, I followed a structured interpretation path:

Step

Action

Objective


1. Domain Mapping

Checked PhysioNet dictionaries.

Understand Dexcom G6 and Empatica E4 measurement logic.


2. Grain Determination

Set the "Grain of Analysis."

Decided on per-hour summaries to balance detail with performance.


3. Distribution Audit

Plotted glucose trends.

Identified outliers and differences between normal/abnormal patterns.


4. Relationship Mapping

Correlation analysis.

Linked glucose spikes to stress levels and specific food triggers.


5. Time-Series Behavior

Observed peaks and drops.

Analyzed glucose stability during the night vs. post-meal intervals.


4. Optimization: Performance at Scale

With hundreds of thousands of rows from sensors like Interbeat Intervals (IBI), a standard database would crawl. I implemented two key strategies:

  • Composite Indexes: By indexing (patient_id, date), the database acts like a textbook with a perfect index. It jumps straight to the data instead of reading the whole "book" (Full Table Scan).

  • Materialized Views: For complex summaries that don't change often, I used Materialized Views to store result sets physically, allowing dashboards to load in milliseconds.

5. The Clinical Impact: Why It Matters

Data cleaning isn't just about deleting rows; it’s about refinement. By moving from a "raw" state to a "normalized" state, we built a foundation that reveals:

  1. Direct Correlations: See how a high-carb lunch impacts glucose three hours later.

  2. Stress Triggers: Track how physical stress (EDA) fluctuates during different daily activities.

  3. Preventative Care: Identify dangerous patterns before they become medical emergencies.


Final Thoughts: We’ve turned a messy pile of sensor readings into a clinical asset. In the world of maternal health, better data isn't just a technical achievement—it's a tool that can save lives.

 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page