top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

Beyond the Query: How I Learned to Speak Gherkin for Data Quality

Jan 14
4 min read

For a long time, my relationship with data was defined by a cycle of "What If" anxiety. I would build a complex model or a detailed visualization, only to have that nagging feeling in the back of my mind: What if the source data changed overnight? What if a null value snuck into a primary key? I lived in a world of reactive troubleshooting, where I was excellent at finding why things broke, but I was exhausted by the constant need to "fix" the foundation before I could actually perform the analysis.

I realized that to truly excel, I had to stop reacting to data anomalies and start defining the truth of the data before it ever reached my scripts. That shift began the moment I discovered Gherkin.


The Core Concept: Moving to Behavior-Driven Logic


Gherkin is a plain-English logic framework used in Behavior-Driven Development (BDD). It follows a simple, human-readable structure: Given, When, Then.

In the world of data, we often get lost in the "how" (the code) before we fully agree on the "what" (the behavior). Gherkin acts as a bridge. It allows you to define exactly how a dataset should behave in plain language, creating a "contract" that exists before a single line of SQL or Python is written.


The "Three Amigos" Strategy


One of the most valuable lessons I learned wasn't about the code, but about the process. In BDD, there is a concept called the "Three Amigos." It involves bringing three perspectives together to write the Gherkin scenarios:

  1. The Business Perspective: Defines what problem we are solving.

  2. The Data/Analytical Perspective: Defines how the logic should be calculated.

  3. The Engineering Perspective: Defines where the data comes from and its constraints.

By sitting down and drafting Gherkin scenarios together, we surfaced misunderstandings early. I once spent an hour refining the definition of "Active User" in Gherkin, only to realize that if we hadn't had that conversation, I would have built a dashboard that was fundamentally wrong.


My Hands-On Learning Path: From Words to Code

I decided to stop jumping straight into the IDE. I began by taking a single logic rule—for example, the way a "VIP" status is calculated—and translating it into a Gherkin scenario:

  • Given: A user has made more than 10 purchases in the last 30 days.

  • When: The daily loyalty aggregation script runs.

  • Then: The user_segment column should be updated to 'VIP'.

  • And: The discount_eligible flag should be set to 'True'.


Learning to Mock and Assert


The biggest technical hurdle I overcame was understanding how to test data without moving millions of rows. I learned to use Mock Data—small, hand-crafted CSVs or temporary SQL tables that contained exactly the "Given" scenario I wrote.

In my Python environment, I used the behave library to load these mocks. The "When" step would trigger my actual transformation script, and the "Then" step used Assertions. An assertion is basically a strict rule: If the result isn't exactly X, stop everything and raise an alarm. This was a massive upgrade from my old method of "eyeballing" the data to see if it looked okay.



Challenges on the Journey


Learning to "Speak Gherkin" wasn't an overnight success. I faced a few significant hurdles while trying to adapt my analytical brain to this new structure:


1. The Mindset Shift

My brain was wired for specific syntax—SQL joins and Python loops. Gherkin felt "too simple" at first. I struggled to see how plain English could drive actual technical tests. The Overcoming: I started by writing down my common "what if" fears as bullet points, then forced them into the Given-When-Then mold. I realized Gherkin wasn't about being verbose; it was about being precise.


2. Finding the Right Level of Detail

My early scenarios were either too vague ("Then the data is correct") or too technical ("When the API fetch function runs"). The Overcoming: I focused on the observable outcome. I asked: "What would signal a problem to the person using this final dataset?" This helped me find the right level of abstraction.


3. The Scalability Trap

As I got excited, I started writing 50 scenarios for one dashboard. It became a maintenance nightmare. The Overcoming: I learned the concept of "High-Value Scenarios." Instead of testing every possible combination, I focused on the "Happy Path" (what should happen normally) and the "Edge Cases" (where the data usually breaks, like null values or zero-dollar orders).



Why This is a Game-Changer for Data Careers

Transitioning to a Gherkin-first mindset has fundamentally changed how I approach my work as a Data Analyst:

  • Living Documentation: Traditionally, documentation is a static PDF that gets outdated the moment you write it. Gherkin files are Living Documentation. Because they are tied to the actual code tests, they are always accurate. If the logic changes, the Gherkin file must change, or the tests will fail.

  • Proactive vs. Reactive: Instead of waiting for a bug to appear in a chart, my Gherkin tests run automatically in the pipeline. If a test fails, I know exactly which business rule was violated before the data even moves downstream.

  • Higher Quality Code: Because I define my expectations first, my SQL and Python have become much cleaner. I’m no longer coding for every possibility; I’m coding to satisfy specific, documented scenarios.



Final Thoughts: The Result of My Journey

Moving from "What If?" to "Given, When, Then" was more than just learning a new syntax; it was a shift in how I view data integrity. It turned me from a technician who "fixes things" into a strategist who architects reliability. Today, my dashboards are stable, my confidence is high, and my troubleshooting time has dropped by nearly 80%. If you find yourself constantly caught in the cycle of data troubleshooting, it might be time to start speaking Gherkin. It’s the difference between hoping your data is right and knowing it is.

 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page