Turning Sensor Noise into Maternal Knowledge: A Deep Dive into Health Data Engineering

When it comes to maternal health, the stakes are as high as they can be. Monitoring a pregnancy today isn't just about monthly check-ups; it’s about the silent, continuous stream of data flowing from wearables 24/7.
But here is the reality no one tells you: raw sensor data is a disaster.
In this project, I integrated messy data from wearables (EDA, Heart Rate) and manual logs (Food intake) into a structured PostgreSQL database. My goal was to turn "noise" into "knowledge" to manage blood sugar and stress during pregnancy.
1. The Challenge: Why We Couldn't Just "Use" the Data
When I first opened the datasets, I found a jigsaw puzzle with missing pieces. The data came from 16 individuals, each generating roughly 4,000 rows of high-frequency measurements every few seconds.
The Wearables: EDA (stress) and Heart Rate sensors spat out thousands of rows with timestamps that looked like gibberish.
The Food Logs: Being manual, entries were inconsistent—people wrote "a bit of rice" or "half a cup."
The Fragmentation: The tables didn't "talk" to each other. You couldn't see how a meal at noon caused a glucose spike at 1:00 PM.
To solve this, I implemented a robust ETL (Extract, Transform, Load) process and a Star Schema architecture. By setting up Referential Integrity, I ensured that every heart rate spike or glucose reading was tied to a validated patient ID, preventing "orphan data."
2. Wrangling Time and Space
In healthcare technology, time is the most critical variable. If a mother’s heart rate jumps, we need to know the context: Was she sleeping or exercising?
Normalizing Timestamps
Raw data arrived with "Combined Timestamps"—massive strings of text. To fix this, I split them into dedicated DATE and TIME columns. This simple change transformed the performance; instead of the computer "chopping strings" for every query, users can now filter morning trends instantly.
The Food Log Challenge: Converting Speech to Math
Humans are inconsistent. I wrote extensive UPDATE statements to standardize entries (e.g., converting "1/2" to 0.5).
Solving for the "Missing Fat":
When users forgot to log fat content, I used the Atwater System to calculate it programmatically:
$$Total Fat = \frac{Calories - (Protein \times 4 + Total Carb \times 4)}{9}$$
3. Interpreting the Signal: The 5-Step Analysis
To turn this cleaned data into clinical insight, I followed a structured interpretation path:
Step | Action | Objective | |
1. Domain Mapping | Checked PhysioNet dictionaries. | Understand Dexcom G6 and Empatica E4 measurement logic. | |
2. Grain Determination | Set the "Grain of Analysis." | Decided on per-hour summaries to balance detail with performance. | |
3. Distribution Audit | Plotted glucose trends. | Identified outliers and differences between normal/abnormal patterns. | |
4. Relationship Mapping | Correlation analysis. | Linked glucose spikes to stress levels and specific food triggers. | |
5. Time-Series Behavior | Observed peaks and drops. | Analyzed glucose stability during the night vs. post-meal intervals. |
4. Optimization: Performance at Scale
With hundreds of thousands of rows from sensors like Interbeat Intervals (IBI), a standard database would crawl. I implemented two key strategies:
Composite Indexes: By indexing (patient_id, date), the database acts like a textbook with a perfect index. It jumps straight to the data instead of reading the whole "book" (Full Table Scan).
Materialized Views: For complex summaries that don't change often, I used Materialized Views to store result sets physically, allowing dashboards to load in milliseconds.
5. The Clinical Impact: Why It Matters
Data cleaning isn't just about deleting rows; it’s about refinement. By moving from a "raw" state to a "normalized" state, we built a foundation that reveals:
Direct Correlations: See how a high-carb lunch impacts glucose three hours later.
Stress Triggers: Track how physical stress (EDA) fluctuates during different daily activities.
Preventative Care: Identify dangerous patterns before they become medical emergencies.
Final Thoughts: We’ve turned a messy pile of sensor readings into a clinical asset. In the world of maternal health, better data isn't just a technical achievement—it's a tool that can save lives.


