A Journey from Data Cleaning to Predictive Analytics Use Case.

A cleaned data gives us the best facts and reliable predictions to perform Data Analysis.
Predictive Analysis makes you guess what will happen in the future by creating a model out of the given data. We perform calculations using a Probability figure and assigning a Risk Score to some dependent variables to come up with an assessment for the future.
Ex. Uses of Predictive Analysis in Business: Businesses use predictions to maintain stock inventory or forecast future sales, so that they can create management plans.
Predictive Analytics comes up with what's likely to happen so that we can take action based on the predicted outcomes. Let's consider a scenario, we would like to know the future increase in cases based on the current positive cases by categorizing the location and travel habits of the local population.
We started by analyzing a raw data file, Ex. file: schema_ontario_final.csv.
We read the information about the purpose of the data collection, reason for the study, understood the frequency of data collection, various other fields and the date/time.
Our sample data was collected as text, numerical and date format, mentioned in week numbers. So, we had to come to a consensus about how to represent this field for future analysis purposes. We decided to convert most of the data fields into 0 or 1 values so that it can be used to create calculated fields for the visualizations.
We created clear column names using underscore'_' for ease of reading. The location field needed format check as it had mixed capitalizations. Some special fields needed to be converted to Boolean to be used during analysis.
We plan to use Python code as it is very friendly to write simple code and create visualizations. We realized that taking time to clean the raw data file helps us to write clear logic instead of using specialized functions and coding at a later stage during code development. We made sure that column names are well understood
Ex. We wanted to predict which location (FSA) will have the highest probable cases in the following month based on some dependent values like current infection contact rates and travel habits of the people.
Step 1. We start by loading the cleaned data file to read into a Python data object.
Step 2. Creating a mapping value for month name to numerical order for forecasting:
Ex. Code:
month_order = {'March': 3, 'April': 4, 'May': 5, 'June': 6, 'July': 7}
df['month_num'] = df['month'].map(month_order)
Step3. Calculating a numerical risk score for some categorical data.
Convert categorical travel habits into a numerical risk score for contacting the infection.
Ex. Code:
travel_map = {
'is_Travel_NonEssential': 10,
'is_Travel_Essential': 8,
'not_Travelling': 1,
'WorkFromHome': 0,
'didntTravelBefore': 0
}
df['travel_risk'] = df['travel_work_school'].map(travel_map).fillna(2)
Step 4. Identify the fields needed to calculate metrics such as aggregation, here we used location and month to categorize.
Ex. Code:
agg_df = df.groupby(['locationfsa', 'month_num']).agg(
prob_risk=('probable', 'mean'),
contact_risk=('contact_with_illness', 'mean'),
intl_travel_risk=('travel_outside_canada', 'mean'),
local_travel_risk=('travel_risk', 'mean')
).reset_index()
Step 5. Create a new variable to calculate 'The increased number of infected cases'.
Ex. Code:
agg_df = agg_df.sort_values(['location-fsa', 'month_num'])
agg_df['next_month_risk'] = agg_df.groupby('locationfsa')['prob_risk'].shift(-1)
Step 6. This is the most important step to apply logic in coding the prediction set, you also need to come up with a logic to exclude the unknowns.
Ex. Code:
predict_surge_df = agg_df.dropna(subset=['next_month_risk'])
features = ['prob_risk', 'contact_risk', 'intl_travel_risk', 'local_travel_risk']
Step 7. Make sure your logical Predictive Model works based on some values.
Ex. Code:
model = RandomForestRegressor(n_estimators=100, random_state=55)
model.fit(train_df[features], train_df['next_month_risk'])
Step 8. Predict if the cases increase in a location for the Upcoming Month
Ex. Code:
new_data = agg_df[agg_df['month_value'] == 7].copy()
new_data['pred_rate'] = model.predict(new_data[features])
new_data['pred_surge'] = new_data['pred_rate'] - new_data['prob_rate']
Step 9. Identify the location with the highest increase in cases.
Ex. Code:
high = new_data.sort_values(by='pred_surge', ascending=True).head(1)
print("The location with the highest predicted cases is: {top_surge['locationfsa'].values[0]}")
Predictive Analytics uses the data, statistical algorithms, and machine learning techniques to estimate for the future.
We wanted to predict which location will have the highest surge in probable cases next month based on current contact rates and travel habits.
So, first we show who is at risk,
Then, we show why they are at risk and
Finally, we used various charts to show the analysis proving the connection between travel behavior of the people and the spread of the virus.
We created a bar chart showing our "Prediction System." It ranks the locations in order that will see the biggest jump in cases next month. The tallest bar represents the location which needs attention immediately.

2: The horizontal chart shows the actual factors to the model. We noted some fields like current probable cases, contact with illness, international travel, and local travel risk.
The probable_rate shows the increase in cases and the longest bar proves that the infection is spreading based on how much people are interacting.

3: The Travel Risk and Predicted cases are shown on a scatter plot, each dot represents a neighborhood. We compare the Risk Score against the Predicted Rate about how many people will get sick.
The color of the dots darken and further showing higher travel risk.
The dots tend to move higher up showing more cases and turn dark color shows the bigger risk.

The visualization clearly proves that when travel increases, the number of cases increases.
The Predictive models highlight the risk before it becomes a problem.


