Showing posts with label predictive analytics. Show all posts
Showing posts with label predictive analytics. Show all posts

Thursday, September 12, 2013

Crossing Party Lines with Predictive Modeling

With the rise of Nate Silver and the emergence of mainstream data science, we've seen many uses for predictive analytics, including the entrance of predictive modeling into the political arena. Actually, although predicting election results is a booming business now, it has been around for quite some time. 

I recently got the chance to talk to Matt Hennessy, Managing Director at Tremont Public Advisors, about a campaign he worked on for Joe Lieberman in 2006, and how they implemented predictive modeling for a successful Senate election. For those who are interested, we'll be discussing this and other examples of predictive modeling in action in a webinar on Tuesday, September 17th. 

Can you give us some background on the 2006 Senate election?

In 2006 in Connecticut, Joe Lieberman was up for reelection to the Senate as a Democrat. He had been the Vice Presidential nominee in the 2000 election and had taken a position supporting the Iraq war which upset a lot of the Democratic base. He wound up losing the Democratic primary to Ned Lamont who won on a big anti-war push. Once Lieberman lost the primary election, he lost access to a considerable amount of infrastructure – union support, door to door field workers, and all of the other boots on the ground that he would have had were all gone. He lost most of his staff except for the people who had been there for a decade or two. He needed to figure out how to replace some of the advantages he’d had with other resources out there.

As someone advising him, I saw that we had a problem: without a field operation and all of those bodies, we didn’t know exactly who we wanted to get out the vote and who the likely voters for Lieberman were. We had a very expensive polling operation going which  was using the conventional method to reach some conclusions about which demographics were most likely to vote, but we decided that we needed something more.

How was the decision made to use predictive analytics in the campaign?

The resources that normally would be used for generating ‘get out the vote’ or direct voter contact were gone the day after the primary. Usually we’d go out and try to visit all of the potential voters, but this just wasn’t possible anymore. We needed to figure out a way to work smarter to compensate for a new lack of resources. We wondered if there was a way to determine which characteristics indicated a likelihood of voting for Lieberman so that we could figure out exactly who to pull out on Election Day. After a conversation with Mike Laracy about performing this type of analysis, we decided to give predictive modeling a try. Our goal was to score every registered voter on their likelihood of voting for Lieberman, and we used Rapid Insight to build a model to do that.

We knew what data we had, the voter file, and determined which additional information we would need to build a model, like demographic information. Then we hired a polling company to call about 10,000 named voters in a random phone pull so that we’d have a statistically significant result. The poll question was a very simple yes/no question on likelihood to vote and who each voter planned on voting for. We weren’t trying to persuade people at this point; all of this polling was meant to influence the field side, not the messaging side. This approach was different than what we’d been doing before because we were calling named voters – people who actually existed and were registered to vote and had demographic information that we could attach to them – and polling them. Using this poll, we scored each of the 1.9 million registered voters in  Connecticut on their likelihood of voting for Senator Lieberman.

How did predictive analytics help the campaign?

Predictive modeling allowed us to optimize our limited resources. As opposed to working with pure assumptions, we now had an actual score attached to each individual voter, which allowed us to spend our resources on the voters with the highest propensity to vote for Lieberman. At the time, it was quite a cutting-edge use of analytics – it was the first time anyone had ever scored an entire state’s voters for the purposes of an election.

Another thing the predictive model did was to disprove our assumptions about who was likely to vote for Lieberman. Some of the key indicators that we were getting from the traditional pollsters were proven to be incorrect by the model results. Based on this we changed some of our campaign messaging The model allowed us to re-allocate our resources more efficiently and it challenged some of the notions we held. In the end, the model did a good job of predicting who the voters would be.

Do you think predictive modeling affected the outcome of the election?

It’s difficult to say, but I can say that the resources that were deployed based on the predictive model were effective. Once we started deploying based on the modeling, the polling margins started to increase; this was toward the end of the race, which is when this model was implemented. I think it increased the margin of victory. The polling was showing a very tight race, but the predictive model was showing there was a margin of victory for Lieberman that was already there, and it was actually ahead of the polling in this case.

How do you see predictive modeling being used in future elections?


If you look at what the Obama campaign did with predictive modeling – taking different factors and a complex web of data points to pinpoint individuals who are likely to vote – it’s here. Predictive modeling is here, it’s now; that’s the future of elections. The complexity of the work they’re doing in this field is truly amazing. I don’t think it will be as focused on many of the smaller races – like those below governor, but it can be very, very effective. I think this last election confirmed that it’s a major part of any political campaign that’s being conducted on scale. This is here to stay.

This example of predictive modeling in action is one of three that we'll be co-presenting in a webinar with Tableau on Tuesday, September 17th, "Turbocharge your Predictive Models with Visualizations". For more information, or to register, click here

*
Matt Hennessy has over two decades of experience in federal, state, and city government. He has built a reputation as a trusted and effective advisor to leading elected officials on public policy, communications and campaign issues. He has served as a trusted political advisor and fundraiser for candidates and political campaigns ranging from Mayor to U.S. Senator to President. Matt is an alumnus of Harvard Business School and the Kennedy School of Government where he was a Wasserman Fellow. He also holds degrees from the Catholic University of America and trinity College in Hartford.

Friday, August 30, 2013

Here's to the Skeptics: Addressing Predictive Modeling Misconceptions

Photo credit: Jonny Goldstein
As a full-time analytics professional, I have a hard time conceiving of people who have not fully embraced the power of predictive analytics, but I know they’re out there and I think it’s important to address their concerns. In doing so, I’m not here to argue that predictive analytics is a perfect fit for every organization. Predictive analytics requires investment: in your data, in infrastructure and technology, and of your time. It’s also an investment in your company, your internal knowledge base, and your future. I’m here to argue that the investment is worth it. 

To do so, I’ve presented a few clarifications to address predictive modeling concerns that I’ve heard from skeptics. If you have anything to add, or if there are any big concerns I’ve missed, let me know in the comments.

You don’t need to be a PhD statistician to build predictive models
A working knowledge of statistics will help you to better interpret the results of predictive models, but you don’t need ten years’ experience or a doctorate degree to glean insight or utilize the output from a model. There are software packages out there with diagnostics that can help you understand which variables are important, which are not, and why. Knowing your data is equally important as statistical knowledge, and both will serve you well in the long run. 

A predictive model shouldn’t be a black box
There are plenty of companies and consultants whose predictive models could fall into the “black box” category.  The model building process, in this case, involves sending your data to an outside party who analyzes it and returns you a series of scores. On the surface, this may not seem like a bad thing, but once you’ve built your first model, you’ll understand why this is not nearly as valuable as doing it yourself. While the output scores are important, you also want to know about the variables used, how the model handled any missing or outlying variables, and glean insight beyond a single set of scores so that you can change or monitor specific behaviors going forward.

Even if you know your data, modeling can help
A finished predictive model will do one of two things: confirm what you’ve always believed, or bring new insights to light. In our office, we refer to this idea as “turn or confirm” – a model will either turn or confirm the things you’ve thought to be true. Most of the time, models will do both. This allows you to both validate any anecdotal evidence you might have (or realize that correlations might not be as strong as you thought) and take a look at new variables or connections that you may not have picked up on before. 

Predictive models can be implemented quickly
I've heard some horror stories about a model taking months, or even years, to implement. If this is the case at your institution, you're doing it wrong. At this point, predictive modeling software has become incredibly efficient - usually able to turn out models within seconds or minutes. The bulk of time spent working on a model is typically spent on the data clean-up, which will vary from company to company. In any case, this is time well spent. Clean data is just as good for reporting, dashboarding, and visualizing as it is for predictive modeling.

Predictive models enhance human judgment, not replace it
If models were meant to replace human judgment, I too would be uncomfortable and suspicious of the idea. However, 99% of the time, the aim of predictive modeling is to enhance and expand human expertise to allow us (the end users) to be better-informed and more data-driven in our decision making.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, July 30, 2013

Why Nonprofits Should Be Building Predictive Models

Last fall, the Whitney Museum of American Art decided to take a different approach when deciding which of their prospective donors to mail. They built their first in-house predictive model from the ground up, and felt ready to use it. They shifted their focus away from some of their prospects who "made sense" but had never given, and used the model to inform a large part of their mailing list. Within the first six months of modeling, they received a $10k donation from a donor they would not have mailed using their previous methodology.

...And they aren't the only ones. More and more nonprofits are turning to predictive modeling to drive their fundraising. For a more in-depth look at the 'hows' and 'whys', I sat down with a man who founded his own company to provide software so that nonprofits and for-profits alike could start building their own predictive models in-house. He also happens to be my boss and one of the smartest people I know - Mike Laracy:

Why would a nonprofit use predictive modeling? How can it drive fundraising?

The quest for any organization, whether a for-profit or non-profit, is to figure out how to achieve its goals and to do so in the most efficient and cost-effective manner possible.   Predictive modeling allows an organization to make better decisions and become more efficient with its use of what are often limited resources.  By using analytics, an organization can better determine who to contact, how often to contact, how much to ask for and how best to achieve their desired fundraising results. 

Although driven by very different motivations, the relationship between a nonprofit and its donors is very similar to the relationship between a for-profit company and its customers.  Customers choose whether to buy a product or not buy a product.  They can become loyal customers or non-loyal customers.  They can buy a lot or they can buy very little.  It is much the same story for nonprofits and their donors.  Donors can be loyal or not loyal.  A prospect can choose to be a donor or not be a donor. They can give large gifts or small gifts.  With accurate data and a modeling process that is easy to implement, a non-profit can begin to model a donor’s behavior using the exact same methodologies that are used to model a customer’s behavior.

What kinds of resources are needed to start building predictive models in-house?

Without quality data, predictive modeling isn’t possible.  So let’s start with that.  There needs to be a system in place that is capturing an organization’s historical data.  Almost every organization is already capturing their data, so that’s usually not a problem.  The data doesn’t necessarily need to be organized in a data warehouse.  In fact, the data needs to be available in its raw form, so sometimes having data pre-aggregated in a warehouse can be a disadvantage.  What’s important is that the data is accessible.

From a staffing perspective, you will need a person or people to collect information on the data, build the models, communicate the results and make sure the models are being used.  There needs to be someone who is making sure the right information is being collected and the right information is being communicated.  This can be a single person, but that person needs to make sure that others in the organization are on board with an understanding of why the models are being built and how they will be used.

What are good first steps for an institution looking to get into predictive modeling?

Like any new initiative, it’s vital to the success of your predictive modeling efforts that there is universal buy-in across the organization.  If there isn’t buy-in, the models won’t be utilized.  To get buy-in, start small.  Go for the early win by building and implementing a single model.  Make sure others in the organization have an understanding of what the model will do, how it will be utilized, and most importantly, how the model will benefit the organization.  Once you get that first win, the interest and buy-in will usually spread quickly across the organization.  As you share the results of those first few successes, begin to identify who the champions for this initiative will be.  Work with them to help them communicate the success of the project organization-wide. 

In your experience, how should an institution decide who should build the predictive models?

Ideally, you want someone who has an understanding of the data.  If you don’t already have someone with that knowledge, you want a person who is willing to learn the data.  Some understanding of statistics is a plus, but with current analytic software technology, there is no longer a need to rely on someone with programming skills or a PhD in statistics to be your data expert.  The people you want to dedicate as resources for predictive modeling should be creative problem solvers who are willing to learn.   

What modeling challenges might be unique to different types of nonprofits?

There are definitely different needs and different challenges depending on what type of fundraising entity you are.  A college advancement office, for example, has an advantage in that they have information on the students who graduated with them.  For example, age comes up as a predictor in many giving models.  Whereas an organization like a museum might not have good info on the age of all of its members and donors, a college or university will at the very least have each student’s year of graduation, which is a great proxy for age.  A college will also have great information like the major each student graduated with and whether or not the current year is a major reunion year.  While a non higher-ed entity won’t have this type of information, they will have information that a college advancement office won’t have.  A museum will have info on its members, how many times someone has visited the museum, and a lot of other great information for modeling that a college won’t have.

Another challenge that people may encounter is how spread out their data is.  Some organizations have more sophisticated computer systems with everything centralized and others may have the information spread across multiple spreadsheets, databases and even outside sources.  As you determine what your data needs to look like, keep in mind that you will need to pull it together and do cleanup before you can begin to model with it.  This was actually one of the reasons we originally created our Veera product.  People were looking for an easier way to clean up and merge their data before they created their models.

Are there any common mistakes to avoid when gearing up to build a model?

I think the biggest mistake to avoid is building a model without buy-in from the rest of the organization.  Another mistake is building a model without an implementation/utilization plan.  Building and scoring a model is great, but by itself the model doesn’t do anything for you.  Before building the model you should have a plan for how you are going to use the model.  For example, if you are a nonprofit and you build a model to predict each donor’s probability of giving to the annual fund, you need to utilize the model in your annual fund outreach.  You will need a plan to mail/call the top X% of your donors with the highest probability of giving, or you should have a plan to not mail donors that are below some probability threshold.  Or perhaps you only want to mail to donors who are likely to give at least a $500 gift.  There are many ways that these models can be used, but the key is that they have to be used.

Once you begin to use them, you can also begin the process of refining and measuring the effectiveness of your models.   Then you can refine them to make them even better.  

What kinds of resources/learning opportunities are out there for those looking to get started with predictive modeling?

In the fundraising world, APRA and the Data Analytic Symposium have a lot of extremely useful sessions.  I’d also recommend Prospect DMM, which is a listserv where a lot of really smart people discuss modeling topics.  We (Rapid Insight) put on a predictive modeling class not too long ago with Brown University and Chuck McClenon from the University of Texas – Austin.  Classes like those are a great place to get started and we’re thinking about doing one again soon.

What strategies can you recommend so that a customer gets the most mileage possible out of their predictive modeling efforts?

To borrow a phrase, I’d say reduce, reuse, recycle. 

Once you’ve set up a process for organizing, cleansing and analyzing your data for one model, you can use that same process for all of your models.  In fact, you can even use that same process for scoring and testing all of your models.  There’s no reason to reinvent the wheel each time. 

Another important strategy is to make sure you set up a system for knowledge capture.  Modeling is an iterative process; you don’t just build one and you’re done.  You can learn a tremendous amount as you’re building models.  A lot of that knowledge is actually knowledge about your data.  That knowledge will accumulate very quickly over time and will make you smarter and smarter as an organization.  This is one of the biggest advantages to bringing predictive modeling in-house:  if you are not doing predictive modeling yourself, you run the risk of that knowledge escaping from your organization.  Once it escapes, you miss out on an opportunity to grow your organization’s analytic intelligence.  

Remember the old proverb about giving a man a fish and feeding him for a day versus a lifetime?  The same thing is true with predictive modeling.  If you give an organization a model; you’ve made them smart for a day.  When you give them the tools to build their own models they become smarter and more competitive for a lifetime.

**
Besides being the Founder and CEO of Rapid Insight, Mike Laracy is a devoted Birkenstock fan, recently ran up Mount Washington, has an eclectic taste in music, loves talking about predictive modeling, is a sap for his two kids, and has pretty much always been a nerd. For those of you attending APRA, he'll be giving a presentation - "Preparing Your Data for Modeling" - on Wednesday, August 7th at 1:30 pm. 

Friday, July 26, 2013

Predicting Retention for Online Students: Where to Start

With the rise of enrollment in online programs and MOOCs, we’re seeing more and more students forego traditional classroom experiences in favor of more flexible online programs. With this shift comes a whole new set of guidelines for enrollment management, financial aid, and retention programs. Retention, in particular, has seen a significant downward trend as learning moves from in-person to online classrooms.


My interest lies in figuring out what variables might be worth including in an analysis attempting to predict online student retention. I did a bit of research and was hoping to find a list of variables online that had worked in the past but couldn’t find any comprehensive resource, so I’ve started to build my own. In the sections below, I’ve listed the type of information that I think would be worth analyzing broken out into four separate categories. Some of these are variables in and of themselves, and some can be broken down different ways; for example, “age” can be used by itself, but creating a “non-traditional age” flag is useful as well. Realistically, not all schools will have all of this information, so this list is meant to be a good starting point of what to shoot for when collecting data.


Also, if you have any variables to add (and I’m sure there are some I’ve missed), I’d love to hear about them in the comments. 

Student Demographic Information
  • Socioeconomic status / financial aid information
    • FAFSA info, Pell eligibility, any scholarship or award info
  • Ethnicity
    • Minority Status
  • Gender
  • Home state
  • Distance from physical campus (if applicable)
  • Age; traditional or non-traditional?
  • Military background?
  • Have children?
  • Currently employed full-time?
  • First generation college student?
  • Legacy student? (Did a parent/grandparent/sibling attend?)

Student Online Learning History
  • Registered for classes online or in person?
  • How many days did they register before the start of the term?
  • Ever attended a class on-campus?
  • Do they plan to attend both online and on-campus classes?
  • Did they attend any type of orientation?
  • Number of previous online courses taken
    • First-time online learner?

Student Academic History
  • GPA
  • SAT/ACT scores
  • Degree hours completed
  • Degree hours attempted
  • Taking developmental courses?
  • Transfer student?
  • Degree program / major 
  • Program level (Associate, Bachelors, Masters, etc.)
  • Number of program or major changes (if applicable)
  • Any previous degrees?

Course- and Program- Related
  • Amount of text vs. interactive content 
  • Lessons with immediate feedback?
  • Any peer-to-peer forum for interaction?
  • Lessons in real time or recorded?
  • Amount of teacher interaction with students
    • Chat, email exchange, turn-around time on assignments

Closing notes:

Getting course-related data might be difficult, but the variables I listed above are derived from studies about how to improve online courses as being areas to focus on; my thinking is that the more engaged a student is, both with peers and instructors, the better their chances of online success are. If you have the data available, it would be worth trying to incorporate it into your model dataset to see whether or not it is predictive.

Rather than using retention as a y-variable when building these models, we typically create an attrition variable (exactly the opposite of retention) and use that as our y instead. This way, we're getting more directly at the characteristics of a student who is likely to leave rather than stay.

Typically when building attrition models, I create separate models for freshmen and upperclassmen. I’d suggest doing that here as well, since previous online coursework will probably be a good indicator of future online coursework. In that case, you’d want to take out many of the variables listed above when modeling freshmen retention.

Finally, it’s important to keep in mind that student success has different meanings for different institutions. You could be basing success on # of credits completed, transitions from semester to semester, or a particular GPA cutoff, among other indicators. When building these different types of student success models, you will probably need to tailor some of these variables to fit the model you're building.

-Caitlin Garrett is a Statistical Analyst at Rapid Insight

Thursday, July 11, 2013

#RIUC13

For those of you who weren’t able to attend the 2013 Rapid Insight User Conference, we set a new record for most attendees and largest number of customer presentations. With two full days of dual track programming, the presenters covered a lot of ground. While we wait for some of the video recordings of customer presentations to be formatted, I thought it would be good to do a quick recap here. 

Mike Laracy, Data Geek (at right)
The conference opened with a keynote from our Founder and CEO, Mike Laracy, who talked a bit about the future of predictive analytics. With a mass public education on the value of analytics (from people like Nate Silver and Billy Bean, with a little help from Brad Pitt), as well as significant advances in data storage and processing power, a stronger need for predictive analytics is emerging. The market is shifting towards the view that more data access is better than restricted access, and that given the right tools along with access, smart people – data scientists – can turn raw data into actionable information. Given these changes, the data scientist – that’s you – will be in increasingly higher demand over the next decade and beyond, as will predictive analytics. 

The user presentations covered lots of different topics, and we’ve made all of their slide decks available here; I’d highly recommend checking them out. In addition to what’s there, I’d also recommend checking out some of the interviews we’ve done with customers on building campaign pyramids and using predictive modeling to drive fundraising efforts. The RI staff team also gave a few presentations,  including topics like Tips and Tricks in Veera, Techniques for Improving Your Predictive Models, and An Introduction to Reporting and Dashboarding with Veera.

Another thing worth mentioning is that we announced our partnership with Tableau to provide a complete solution for both predictive modeling and visualization. Now users can use Veera to clean up their data, Analytics to build their predictive models, and Tableau’s visualizations to turbocharge their presentations. For more information, check out our partner page.
My favorite part of the User Conference has always been talking to customers about the cool data projects that they’ve been tackling, and this year was no different. Kudos to our users for being so creative and smart with the ways they use our software. We also owe a big thanks to the folks at Yale for hosting us, and to all who were able to attend. Here’s to the best User Conference so far and to making next year’s even better!

Tuesday, March 5, 2013

Six Predictive Modeling Mistakes

As we mentioned in our post on Data Preparation Mistakes, we've built many predictive models in the Rapid Insight office. During the predictive modeling process, there are many places where it's easy to make mistakes. Luckily, we've compiled a few here so you can learn from our mistakes and avoid them in your own analyses:

Failing to consider enough variables
When deciding which variables to audition for a model, you want to include anything you have on-hand that you think could possibly be predictive. Weeding out the extra variables is something that your modeling program will do, so don’t be afraid to throw the kitchen sink at it for your first pass.

Not hand-crafting some additional variables
Any guide-list of variables should be used as just that – a guide – enriched by other variables that may be unique to your institution.  If there are few unique variables to be had, consider creating some to augment your dataset. Try adding new fields like “distance from institution” or creating riffs and derivations of variables you already have.

Selecting the wrong Y-variable
When building your dataset for a logistic regression model, you’ll want to select the response with the smaller number of data points as your y-variable. A great example of this from the higher ed world would come from building a retention model. In most cases, you’ll actually want to model attrition, identifying those students who are likely to leave (hopefully the smaller group!) rather than those who are likely to stay.

Not enough Y-variable responses
Along with making sure that your model population is large enough (1,000 records minimum) and spans enough time (3 years is good), you’ll want to make sure that there are enough Y-variable responses to model. Generally, you’ll want to shoot for at least 100 instances of the response you’d like to model.

Building a model on the wrong population
To borrow an example from the world of fundraising, a model built to predict future giving will look a lot different for someone with a giving history than someone who has never given before. Consider which population you’d eventually like to use the model to score and build the model tailored to that population, or consider building two models, one for each sub-group.

Judging the quality of a model using one measure
It’s difficult to capture the quality of a model in a single number, which is why modeling outputs provide so many model fit measures. Beyond the numbers, graphic outputs like decile analysis and lift analysis can provide visual insight into how well the model is fitting your data and what the gains from using a model are likely to be.

If you’re not sure which model measures to focus on, ask around. If you know someone building models similar to yours, see which ones they rely on and what ranges they shoot for. The take-home point is that with all of the information available on a model output, you’ll want to consider multiple gauges before deciding whether your model is worth moving forward with.  

-Caitlin Garrett, Statistical Analyst at Rapid Insight
Photo Credit: http://www.flickr.com/photos/mattimattila/


Have you made any of the above mistakes? Tell us about it (and how you found it!) in the comments. 

Wednesday, January 16, 2013

How to Interpret a Decile Analysis


After building a predictive model, there are several ways to determine how well the model is describing your data. One visual way to get an idea of how well a model is fitting your data is by taking a look at the decile analysis. Here we’ll take a look at what the decile analysis represents, how it’s created, and how to spot a good model.

What a Decile Analysis Represents

After building a statistical model, a decile analysis is created to test the model’s ability to predict the intended outcome. Each column in the decile analysis chart represents a collection of records that have been scored using the model. The height of each column represents the average of those records’ actual behavior.

How the Decile Analysis is Calculated

1. The hold-out or validation sample is scored according to the model being tested.
2. The records are sorted by their predicted scores in descending order and divided into ten equal-sized bins or deciles. The top decile contains the 10% of the population most likely to respond and the bottom decile contains the 10% of the population least likely to respond, based on the model scores.
3. The deciles and their actual response rates are graphed on the x and y axes, respectively. 

After the decile analysis is built, you’ll want to take a look at the height of the bars in relation to one another. Deciding whether a model is worth moving forward with depends on the pattern you see when viewing the decile analysis. 


Ideal Situation: The Staircase Effect

When you’re looking at a decile analysis, you want to see a staircase effect; that is, you’ll want the bars to descend in order from left to right, as shown below. 

This is telling you that the model is “binning” your constituents correctly from most likely to respond to least likely to respond. A model exhibiting a good staircase decile analysis is one you can consider moving forward with.

Not-So-Ideal Situations

In contrast, if the bars seem to be out of order (as shown below), the decile analysis is telling you that the model is not doing a very good job of predicting actual responses.


 If the bars seem to be the same height, or the decile analysis looks “flat”, the decile analysis is telling you that the model isn’t performing any better than randomly binning people into deciles would. In both cases, your model should be improved before moving forward with it.  

-Caitlin Garrett, Statistical Analyst at Rapid Insight


Thursday, January 10, 2013

Valuing Analytics & Predictive Modeling in Higher Ed

As promised, here is part two of my interview with Mike Laracy, Founder, President, and CEO  of Rapid Insight. Mike's 20+ years of data analytics & predictive modeling experience have provided him with many insights. Here's Mike on becoming more data-driven in higher education, which models produce the highest ROI, and mistakes to avoid:

Where does predictive modeling fit into the analytic ecosystem in higher education?

Within the analytic ecosystem in higher ed, there is a range of ways in which data is analyzed and looked at. On one side, you have historical reporting, which our clients do a lot of and is vital to every institution.  Somewhere in the middle is data exploration and analysis, where you’re slicing and dicing data to understand it better or make more informed decisions based on what happened in the past.  On the other side of the spectrum is predictive modeling.  Modeling requires taking a look at all of the variables in a given set of information to make informed predictions about what will happen in the future. What is each applicant’s probability of enrolling or what is each student’s attrition likelihood?  What will the incoming class look like based on the current admit pool?  These are the types of questions that are being answered in higher ed with predictive analytics.  The resulting probabilities can also be used in the aggregate. For example, enrollment models allow you to predict overall enrollment, enrollment by gender, by program, or by any other factor.  The models are also used to project financial outlay based on the financial aid promised to admitted applicants and their individual enrollment probabilities.

Higher education has come a long way in the last five to ten years in its use of predictive analytics. The entire student life cycle is now being modeled starting with prospect and inquiry modeling all the way through to alumni donor modeling.   It used to be that any institutions that were doing this kind of modeling were relying on outside consulting companies.  Today most are doing their modeling in-house.  Colleges and universities view their data as a strategic asset and they are extracting value from their data with the same tools and methodologies as the Fortune 500 companies.

What kinds of resources are needed and what is the first step for an institution who wants to become more data-driven in their decision making?

It’s important to have somebody who knows the data. As long as a user has an understanding of their data, our software makes it very easy to analyze data and build predictive models very quickly. And our support team is available to answer any analytic questions. 

Gaining access to their data is the first step. We see a lot of institutions that have some reporting tools which don’t allow them to ask new questions of the data. So, they might have a set of 50 reports that they’re able to run over and over but anytime someone has a new question, without access to the raw data there’s no way to answer the question. 

It really helps if the institution is committed to a culture of data driven decision making.  Then all the various stakeholders are more focused on ensuring data access for those doing the predictive modeling.

What do you say to those who are on “the quest for perfect data”?  Is it okay to implement predictive analytics before you have that data warehouse or those perfectly cleansed datasets?

No institution is ever going to have perfect data, so you work with what you have. We suggest seeing what you have, finding any obvious problems in the data, and then fixing those problems the best you can. We’ve designed our solutions such that a data warehouse is not required but, even with a clean data warehouse, the data is never going to be perfect.   As long as you as you have an understanding of the data, you can move forward.  

In your experience, which models in higher education produce the highest ROI?
We have a customer, Paul Smith’s College that has quantified their retention modeling efforts. Using their model results, they put programs into place to help those students that were predicted to be high-risk of attrition. They credit the modeling with helping them identify which students to focus on, saving them $3m in net tuition revenue so far.

We have other clients that are using predictive modeling on the prospect side and they’re realizing significant savings on their recruiting efforts. So instead of mailing to 200,000 high school seniors, they’re mailing to 50,000, and realizing significant savings by not mailing and not calling those students who have pretty much zero probability of applying or enrolling.

Although not as easily quantifiable, enrollment modeling has a pretty big ROI.  Not only on determining which applicants are likely to enroll, but in predicting class size.  If an institution overshoots and enrolls too many applicants, they’ll have dorm, classroom, and other resource issues.  If enroll too little, they’ll have revenue issues.  So predicting class size and determining who and how many applicants to admit is extremely important.

What are some common mistakes you see when approaching predictive modeling for your higher ed customers?

One mistake that I often see is when information is thrown out as not useful to the models.  Zip code is a good example.  Zip code looks like a five digit numeric variable, but you wouldn’t want to use it as a numeric variable in a model.  In some cases it can be used categorically to help identify applicants’ origins, but its most useful purpose is to for calculating a distance from campus variable.  This is a variable that we see showing up as a predictor in many prospect/ inquiry models, enrollment models, alumni models, and even retention models.  Another example of a variable that is often overlooked is application date.  Application date often contains a ton of useful information if looked at correctly.  It can be used to calculate the number of days between when the application was sent and the application deadline.  This piece of information can tell you a lot about an applicant’s intentions.  A student who gets their application in the day before the deadline probably has very different intentions than a student who applies nine months before the deadline.  This variable ends up participating in many models. 

To get our customers up to speed on best practices in predictive modeling we’ve created resources like lists of recommended variables for specific models and guides on how to create useful new variables from existing data.

Tuesday, December 11, 2012

Five Data Preparation Mistakes (and How to Avoid Them!)

After building many predictive models in the Rapid Insight office and helping our customer build many more models outside of the office, we have a list of data preparation mistakes that could fill a room. Here are some of the most common ones we've seen:


1. Including ID Fields as Predictors
Because most IDs look like continuous integers (and older IDs are typically smaller), it is possible that they may make their way into the model as a predictive variables. Be sure to exclude them as early on in the process as possible to avoid any confusion while building your model.

2. Using Anachronistic Variables
Make sure that no predictor variables contain information about the outcome. Because models are built using historical data, it is possible that some of the variables you have accessible when building your model were not available at the time the model is built to reflect. No predictor variables should be proxies for your dependent variable (ie: “made a gift” = donor, “deposited” = enrolled).

3. Allowing Duplicate Records
Don’t include duplicates in a model file. Including just two records per person gives that person twice as much predictive power. To make sure that each person’s influence counts equally, only one record per person or action being modeled should be included. It never hurts to dedupe your model file before you start building a predictive model. 

4. Modeling on Too Small of a Population
Double-check your population size. A good goal to shoot for in a modeling dataset is at least 1,000 records spanning three years. Including at least three years helps to account for any year-to-year fluctuations in your dataset. The larger your population size is, the most robust your model will be. 

5. Not Accounting for Outliers and/or Missing Values
Be sure to account for any outliers and/or missing values. Large rifts in individual variables can add up when you’re combining those variables to build a predictive model. Checking the minimum and maximum values for each variable can be a quick way to spot any records that are out of the usual realm. 

[photo credit]

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, February 21, 2012

On Target: Predicting Pregnancy



Call me biased, but I think creative uses for predictive analytics are pretty cool.  Target’s “pregnancy-prediction model”, explained Thursday in a The New York Times Magazine article, is a great example.  It should inspire all of us to take a fresh look at our data and consider what more we can accomplish with a powerful predictive analysis tool (like RI Analytics) and a little bit of creative thinking.



Target’s journey to predicting pregnancy started with an idea conceived by its marketing department. The department had previously conducted surveys which indicated that once a consumer’s shopping habits are ingrained, it can be hard to change them – except during certain brief periods of a person’s life, like after a marriage or the birth of a child, where shopping patterns and brand loyalties often change.  The birth of a child represents a new grocery and household goods list for new parents, as well as the opportunity for Target to sell things like cribs, rugs, furniture, car seats, and other items that a person or couple would not usually buy. Because birth records are public information it was already common practice for companies to send promotional items to new parents; so, to stay one step ahead of competitors, marketers at Target wanted to see if there was a way to predict pregnancy during the second trimester.

Target reviewed the shopping habits of women who had a baby-shower registry as they approached their due dates. Eventually they were able to identify about 25 different products that were indicators of pregnancy, including items like unscented lotion, vitamin supplements, hand sanitizers and washcloths. By treating the purchase of each item as a variable, they were able to create a model that assigned each shopper a pregnancy prediction score based on their purchases. This score was then used to send out relevant coupons and advertisements tailored to each woman at a specific point in her pregnancy – before other retailers even knew she was pregnant. Needless to say, sales in Target’s Mom and Baby department skyrocketed.

This is one way that a creative use of data, combined with some predictive analytics, yields some pretty cool results. Target had the data they needed all along –they just needed the right person to ask the right question. 

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Friday, January 27, 2012

Is this thing on?

Hi everyone, Caitlin here on the Rapid Insight blog. I've been reading some analytics articles this week, and one term I've heard thrown around a lot is "big data". For those of you who aren't familiar with this term, it is used to describe the massive amounts of data that have been created over the past few years in response to technological advances and an increased effort to link data from disparate sources. Those of you who are familiar with this term know that there are positive and negative implications of using such a large expanse of data to create predictive models. Lots of data can mean the ability to predict broader outcomes with more accuracy, but also leaves more room for overconfidence and error. Regardless of the merits and flaws in utilizing it, big data is here to stay. 

One article in particular, Fast Company's "Why Big Data Won't Make You Smart, Rich, Or Pretty" (written by Daniel Rasmus), provides a more in-depth look at some of the challenges that big data poses and is worth a read. One such problem that I'd like to focus on is what Rasmus headlines as "Complexity": 
          
   "I was sitting with the CIO of a large insurance company in Portland. We were talking about generational hand-offs when he raised the issue of an Excel spreadsheet used to evaluate commercial property underwriting. He said one of the older members of the organization owned that spreadsheet and he was the only one who knew how it worked. The hand-off issue was not one of getting the older employee to collaborate with the younger employee, but one of complexity. That spreadsheet was complex and tightly woven into the employee's worldview. Although the transfer could theoretically take place, it is unknowable how long it would take, if the new employee would stay, or how the process would change as multiple worldviews collided. Combining models full of nuance and obscurity increases complexity. Organizations that plan complex uses of Big Data and the algorithms that analyze the data need to think about continuity and succession planning in order to maintain the accuracy and relevance of their models over time, and they need to be very cautious about the time it will take to integrate, and the value of results achieved, from data and models that border on the cryptic."

The amount of time and energy that goes into creating spreadsheets like the one mentioned here is incalculable. As times change and datasets grow, these spreadsheets are tasked with leveraging the information and formulas already contained with a constant flow of new variables and considerations. As you can imagine, the rise of big data is only increasing the complexity of combining and augmenting various data and models. One thing I've learned from our customers is that the point of a hand-off is particularly tricky, especially when each spreadsheet can be full of its own nuances and quirks. The days when a spreadsheet was meant for only one person to understand, use, and manipulate seem to dimming and giving way to more collaborative and transparent efforts. This is the kind of environment in which I see Rapid Insight's data intelligence tool, Veera, being a big help. 

Veera is able to accommodate large amounts of data due to its ability to run data through any user-created jobs without actually storing any data within the program. It has the same abilities that Excel does for manipulating variables and much more, which means no need to further obscure or compromise the accuracy of data with constant formula editing. Not only is the Veera platform easier to use for data manipulation, it is also transparent: you can open any node at any time if you want to take a look under the hood to see what changes are occurring, and how. Passing projects on is easy - you can drag and drop Veera jobs in and out of emails, and attach explanatory notes to every node in the job if you want. I find myself using the notes feature all the time, if only to help jog my memory after taking some time away from a particular job or data path. The ability to save jobs allows you to keep data processes on hand, so rather than recreating a job for each dataset you'd like to use, you can use them as many times as needed, and of course edit any job whenever necessary. 

Although big data does come with some extra considerations, Veera is well equipped to deal with an increasing amount of data while accommodating needs for transparency and ease of explanation. With spreadsheets constantly being updated and sent from person to person, I feel that the less information one has to remember about the data, and the more that can be noted or observed during the analysis process, the better. Veera helps us to adapt to the influx of big data, keeping past work in-tact and relevant while allowing plenty of room for change and growth as new challenges present themselves. 

-Caitlin Garrett, Statistical Analyst at Rapid Insight