Showing posts with label predictive modeling. Show all posts
Showing posts with label predictive modeling. Show all posts

Thursday, September 12, 2013

Crossing Party Lines with Predictive Modeling

With the rise of Nate Silver and the emergence of mainstream data science, we've seen many uses for predictive analytics, including the entrance of predictive modeling into the political arena. Actually, although predicting election results is a booming business now, it has been around for quite some time. 

I recently got the chance to talk to Matt Hennessy, Managing Director at Tremont Public Advisors, about a campaign he worked on for Joe Lieberman in 2006, and how they implemented predictive modeling for a successful Senate election. For those who are interested, we'll be discussing this and other examples of predictive modeling in action in a webinar on Tuesday, September 17th. 

Can you give us some background on the 2006 Senate election?

In 2006 in Connecticut, Joe Lieberman was up for reelection to the Senate as a Democrat. He had been the Vice Presidential nominee in the 2000 election and had taken a position supporting the Iraq war which upset a lot of the Democratic base. He wound up losing the Democratic primary to Ned Lamont who won on a big anti-war push. Once Lieberman lost the primary election, he lost access to a considerable amount of infrastructure – union support, door to door field workers, and all of the other boots on the ground that he would have had were all gone. He lost most of his staff except for the people who had been there for a decade or two. He needed to figure out how to replace some of the advantages he’d had with other resources out there.

As someone advising him, I saw that we had a problem: without a field operation and all of those bodies, we didn’t know exactly who we wanted to get out the vote and who the likely voters for Lieberman were. We had a very expensive polling operation going which  was using the conventional method to reach some conclusions about which demographics were most likely to vote, but we decided that we needed something more.

How was the decision made to use predictive analytics in the campaign?

The resources that normally would be used for generating ‘get out the vote’ or direct voter contact were gone the day after the primary. Usually we’d go out and try to visit all of the potential voters, but this just wasn’t possible anymore. We needed to figure out a way to work smarter to compensate for a new lack of resources. We wondered if there was a way to determine which characteristics indicated a likelihood of voting for Lieberman so that we could figure out exactly who to pull out on Election Day. After a conversation with Mike Laracy about performing this type of analysis, we decided to give predictive modeling a try. Our goal was to score every registered voter on their likelihood of voting for Lieberman, and we used Rapid Insight to build a model to do that.

We knew what data we had, the voter file, and determined which additional information we would need to build a model, like demographic information. Then we hired a polling company to call about 10,000 named voters in a random phone pull so that we’d have a statistically significant result. The poll question was a very simple yes/no question on likelihood to vote and who each voter planned on voting for. We weren’t trying to persuade people at this point; all of this polling was meant to influence the field side, not the messaging side. This approach was different than what we’d been doing before because we were calling named voters – people who actually existed and were registered to vote and had demographic information that we could attach to them – and polling them. Using this poll, we scored each of the 1.9 million registered voters in  Connecticut on their likelihood of voting for Senator Lieberman.

How did predictive analytics help the campaign?

Predictive modeling allowed us to optimize our limited resources. As opposed to working with pure assumptions, we now had an actual score attached to each individual voter, which allowed us to spend our resources on the voters with the highest propensity to vote for Lieberman. At the time, it was quite a cutting-edge use of analytics – it was the first time anyone had ever scored an entire state’s voters for the purposes of an election.

Another thing the predictive model did was to disprove our assumptions about who was likely to vote for Lieberman. Some of the key indicators that we were getting from the traditional pollsters were proven to be incorrect by the model results. Based on this we changed some of our campaign messaging The model allowed us to re-allocate our resources more efficiently and it challenged some of the notions we held. In the end, the model did a good job of predicting who the voters would be.

Do you think predictive modeling affected the outcome of the election?

It’s difficult to say, but I can say that the resources that were deployed based on the predictive model were effective. Once we started deploying based on the modeling, the polling margins started to increase; this was toward the end of the race, which is when this model was implemented. I think it increased the margin of victory. The polling was showing a very tight race, but the predictive model was showing there was a margin of victory for Lieberman that was already there, and it was actually ahead of the polling in this case.

How do you see predictive modeling being used in future elections?


If you look at what the Obama campaign did with predictive modeling – taking different factors and a complex web of data points to pinpoint individuals who are likely to vote – it’s here. Predictive modeling is here, it’s now; that’s the future of elections. The complexity of the work they’re doing in this field is truly amazing. I don’t think it will be as focused on many of the smaller races – like those below governor, but it can be very, very effective. I think this last election confirmed that it’s a major part of any political campaign that’s being conducted on scale. This is here to stay.

This example of predictive modeling in action is one of three that we'll be co-presenting in a webinar with Tableau on Tuesday, September 17th, "Turbocharge your Predictive Models with Visualizations". For more information, or to register, click here

*
Matt Hennessy has over two decades of experience in federal, state, and city government. He has built a reputation as a trusted and effective advisor to leading elected officials on public policy, communications and campaign issues. He has served as a trusted political advisor and fundraiser for candidates and political campaigns ranging from Mayor to U.S. Senator to President. Matt is an alumnus of Harvard Business School and the Kennedy School of Government where he was a Wasserman Fellow. He also holds degrees from the Catholic University of America and trinity College in Hartford.

Friday, August 30, 2013

Here's to the Skeptics: Addressing Predictive Modeling Misconceptions

Photo credit: Jonny Goldstein
As a full-time analytics professional, I have a hard time conceiving of people who have not fully embraced the power of predictive analytics, but I know they’re out there and I think it’s important to address their concerns. In doing so, I’m not here to argue that predictive analytics is a perfect fit for every organization. Predictive analytics requires investment: in your data, in infrastructure and technology, and of your time. It’s also an investment in your company, your internal knowledge base, and your future. I’m here to argue that the investment is worth it. 

To do so, I’ve presented a few clarifications to address predictive modeling concerns that I’ve heard from skeptics. If you have anything to add, or if there are any big concerns I’ve missed, let me know in the comments.

You don’t need to be a PhD statistician to build predictive models
A working knowledge of statistics will help you to better interpret the results of predictive models, but you don’t need ten years’ experience or a doctorate degree to glean insight or utilize the output from a model. There are software packages out there with diagnostics that can help you understand which variables are important, which are not, and why. Knowing your data is equally important as statistical knowledge, and both will serve you well in the long run. 

A predictive model shouldn’t be a black box
There are plenty of companies and consultants whose predictive models could fall into the “black box” category.  The model building process, in this case, involves sending your data to an outside party who analyzes it and returns you a series of scores. On the surface, this may not seem like a bad thing, but once you’ve built your first model, you’ll understand why this is not nearly as valuable as doing it yourself. While the output scores are important, you also want to know about the variables used, how the model handled any missing or outlying variables, and glean insight beyond a single set of scores so that you can change or monitor specific behaviors going forward.

Even if you know your data, modeling can help
A finished predictive model will do one of two things: confirm what you’ve always believed, or bring new insights to light. In our office, we refer to this idea as “turn or confirm” – a model will either turn or confirm the things you’ve thought to be true. Most of the time, models will do both. This allows you to both validate any anecdotal evidence you might have (or realize that correlations might not be as strong as you thought) and take a look at new variables or connections that you may not have picked up on before. 

Predictive models can be implemented quickly
I've heard some horror stories about a model taking months, or even years, to implement. If this is the case at your institution, you're doing it wrong. At this point, predictive modeling software has become incredibly efficient - usually able to turn out models within seconds or minutes. The bulk of time spent working on a model is typically spent on the data clean-up, which will vary from company to company. In any case, this is time well spent. Clean data is just as good for reporting, dashboarding, and visualizing as it is for predictive modeling.

Predictive models enhance human judgment, not replace it
If models were meant to replace human judgment, I too would be uncomfortable and suspicious of the idea. However, 99% of the time, the aim of predictive modeling is to enhance and expand human expertise to allow us (the end users) to be better-informed and more data-driven in our decision making.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, July 30, 2013

Why Nonprofits Should Be Building Predictive Models

Last fall, the Whitney Museum of American Art decided to take a different approach when deciding which of their prospective donors to mail. They built their first in-house predictive model from the ground up, and felt ready to use it. They shifted their focus away from some of their prospects who "made sense" but had never given, and used the model to inform a large part of their mailing list. Within the first six months of modeling, they received a $10k donation from a donor they would not have mailed using their previous methodology.

...And they aren't the only ones. More and more nonprofits are turning to predictive modeling to drive their fundraising. For a more in-depth look at the 'hows' and 'whys', I sat down with a man who founded his own company to provide software so that nonprofits and for-profits alike could start building their own predictive models in-house. He also happens to be my boss and one of the smartest people I know - Mike Laracy:

Why would a nonprofit use predictive modeling? How can it drive fundraising?

The quest for any organization, whether a for-profit or non-profit, is to figure out how to achieve its goals and to do so in the most efficient and cost-effective manner possible.   Predictive modeling allows an organization to make better decisions and become more efficient with its use of what are often limited resources.  By using analytics, an organization can better determine who to contact, how often to contact, how much to ask for and how best to achieve their desired fundraising results. 

Although driven by very different motivations, the relationship between a nonprofit and its donors is very similar to the relationship between a for-profit company and its customers.  Customers choose whether to buy a product or not buy a product.  They can become loyal customers or non-loyal customers.  They can buy a lot or they can buy very little.  It is much the same story for nonprofits and their donors.  Donors can be loyal or not loyal.  A prospect can choose to be a donor or not be a donor. They can give large gifts or small gifts.  With accurate data and a modeling process that is easy to implement, a non-profit can begin to model a donor’s behavior using the exact same methodologies that are used to model a customer’s behavior.

What kinds of resources are needed to start building predictive models in-house?

Without quality data, predictive modeling isn’t possible.  So let’s start with that.  There needs to be a system in place that is capturing an organization’s historical data.  Almost every organization is already capturing their data, so that’s usually not a problem.  The data doesn’t necessarily need to be organized in a data warehouse.  In fact, the data needs to be available in its raw form, so sometimes having data pre-aggregated in a warehouse can be a disadvantage.  What’s important is that the data is accessible.

From a staffing perspective, you will need a person or people to collect information on the data, build the models, communicate the results and make sure the models are being used.  There needs to be someone who is making sure the right information is being collected and the right information is being communicated.  This can be a single person, but that person needs to make sure that others in the organization are on board with an understanding of why the models are being built and how they will be used.

What are good first steps for an institution looking to get into predictive modeling?

Like any new initiative, it’s vital to the success of your predictive modeling efforts that there is universal buy-in across the organization.  If there isn’t buy-in, the models won’t be utilized.  To get buy-in, start small.  Go for the early win by building and implementing a single model.  Make sure others in the organization have an understanding of what the model will do, how it will be utilized, and most importantly, how the model will benefit the organization.  Once you get that first win, the interest and buy-in will usually spread quickly across the organization.  As you share the results of those first few successes, begin to identify who the champions for this initiative will be.  Work with them to help them communicate the success of the project organization-wide. 

In your experience, how should an institution decide who should build the predictive models?

Ideally, you want someone who has an understanding of the data.  If you don’t already have someone with that knowledge, you want a person who is willing to learn the data.  Some understanding of statistics is a plus, but with current analytic software technology, there is no longer a need to rely on someone with programming skills or a PhD in statistics to be your data expert.  The people you want to dedicate as resources for predictive modeling should be creative problem solvers who are willing to learn.   

What modeling challenges might be unique to different types of nonprofits?

There are definitely different needs and different challenges depending on what type of fundraising entity you are.  A college advancement office, for example, has an advantage in that they have information on the students who graduated with them.  For example, age comes up as a predictor in many giving models.  Whereas an organization like a museum might not have good info on the age of all of its members and donors, a college or university will at the very least have each student’s year of graduation, which is a great proxy for age.  A college will also have great information like the major each student graduated with and whether or not the current year is a major reunion year.  While a non higher-ed entity won’t have this type of information, they will have information that a college advancement office won’t have.  A museum will have info on its members, how many times someone has visited the museum, and a lot of other great information for modeling that a college won’t have.

Another challenge that people may encounter is how spread out their data is.  Some organizations have more sophisticated computer systems with everything centralized and others may have the information spread across multiple spreadsheets, databases and even outside sources.  As you determine what your data needs to look like, keep in mind that you will need to pull it together and do cleanup before you can begin to model with it.  This was actually one of the reasons we originally created our Veera product.  People were looking for an easier way to clean up and merge their data before they created their models.

Are there any common mistakes to avoid when gearing up to build a model?

I think the biggest mistake to avoid is building a model without buy-in from the rest of the organization.  Another mistake is building a model without an implementation/utilization plan.  Building and scoring a model is great, but by itself the model doesn’t do anything for you.  Before building the model you should have a plan for how you are going to use the model.  For example, if you are a nonprofit and you build a model to predict each donor’s probability of giving to the annual fund, you need to utilize the model in your annual fund outreach.  You will need a plan to mail/call the top X% of your donors with the highest probability of giving, or you should have a plan to not mail donors that are below some probability threshold.  Or perhaps you only want to mail to donors who are likely to give at least a $500 gift.  There are many ways that these models can be used, but the key is that they have to be used.

Once you begin to use them, you can also begin the process of refining and measuring the effectiveness of your models.   Then you can refine them to make them even better.  

What kinds of resources/learning opportunities are out there for those looking to get started with predictive modeling?

In the fundraising world, APRA and the Data Analytic Symposium have a lot of extremely useful sessions.  I’d also recommend Prospect DMM, which is a listserv where a lot of really smart people discuss modeling topics.  We (Rapid Insight) put on a predictive modeling class not too long ago with Brown University and Chuck McClenon from the University of Texas – Austin.  Classes like those are a great place to get started and we’re thinking about doing one again soon.

What strategies can you recommend so that a customer gets the most mileage possible out of their predictive modeling efforts?

To borrow a phrase, I’d say reduce, reuse, recycle. 

Once you’ve set up a process for organizing, cleansing and analyzing your data for one model, you can use that same process for all of your models.  In fact, you can even use that same process for scoring and testing all of your models.  There’s no reason to reinvent the wheel each time. 

Another important strategy is to make sure you set up a system for knowledge capture.  Modeling is an iterative process; you don’t just build one and you’re done.  You can learn a tremendous amount as you’re building models.  A lot of that knowledge is actually knowledge about your data.  That knowledge will accumulate very quickly over time and will make you smarter and smarter as an organization.  This is one of the biggest advantages to bringing predictive modeling in-house:  if you are not doing predictive modeling yourself, you run the risk of that knowledge escaping from your organization.  Once it escapes, you miss out on an opportunity to grow your organization’s analytic intelligence.  

Remember the old proverb about giving a man a fish and feeding him for a day versus a lifetime?  The same thing is true with predictive modeling.  If you give an organization a model; you’ve made them smart for a day.  When you give them the tools to build their own models they become smarter and more competitive for a lifetime.

**
Besides being the Founder and CEO of Rapid Insight, Mike Laracy is a devoted Birkenstock fan, recently ran up Mount Washington, has an eclectic taste in music, loves talking about predictive modeling, is a sap for his two kids, and has pretty much always been a nerd. For those of you attending APRA, he'll be giving a presentation - "Preparing Your Data for Modeling" - on Wednesday, August 7th at 1:30 pm. 

Friday, July 26, 2013

Predicting Retention for Online Students: Where to Start

With the rise of enrollment in online programs and MOOCs, we’re seeing more and more students forego traditional classroom experiences in favor of more flexible online programs. With this shift comes a whole new set of guidelines for enrollment management, financial aid, and retention programs. Retention, in particular, has seen a significant downward trend as learning moves from in-person to online classrooms.


My interest lies in figuring out what variables might be worth including in an analysis attempting to predict online student retention. I did a bit of research and was hoping to find a list of variables online that had worked in the past but couldn’t find any comprehensive resource, so I’ve started to build my own. In the sections below, I’ve listed the type of information that I think would be worth analyzing broken out into four separate categories. Some of these are variables in and of themselves, and some can be broken down different ways; for example, “age” can be used by itself, but creating a “non-traditional age” flag is useful as well. Realistically, not all schools will have all of this information, so this list is meant to be a good starting point of what to shoot for when collecting data.


Also, if you have any variables to add (and I’m sure there are some I’ve missed), I’d love to hear about them in the comments. 

Student Demographic Information
  • Socioeconomic status / financial aid information
    • FAFSA info, Pell eligibility, any scholarship or award info
  • Ethnicity
    • Minority Status
  • Gender
  • Home state
  • Distance from physical campus (if applicable)
  • Age; traditional or non-traditional?
  • Military background?
  • Have children?
  • Currently employed full-time?
  • First generation college student?
  • Legacy student? (Did a parent/grandparent/sibling attend?)

Student Online Learning History
  • Registered for classes online or in person?
  • How many days did they register before the start of the term?
  • Ever attended a class on-campus?
  • Do they plan to attend both online and on-campus classes?
  • Did they attend any type of orientation?
  • Number of previous online courses taken
    • First-time online learner?

Student Academic History
  • GPA
  • SAT/ACT scores
  • Degree hours completed
  • Degree hours attempted
  • Taking developmental courses?
  • Transfer student?
  • Degree program / major 
  • Program level (Associate, Bachelors, Masters, etc.)
  • Number of program or major changes (if applicable)
  • Any previous degrees?

Course- and Program- Related
  • Amount of text vs. interactive content 
  • Lessons with immediate feedback?
  • Any peer-to-peer forum for interaction?
  • Lessons in real time or recorded?
  • Amount of teacher interaction with students
    • Chat, email exchange, turn-around time on assignments

Closing notes:

Getting course-related data might be difficult, but the variables I listed above are derived from studies about how to improve online courses as being areas to focus on; my thinking is that the more engaged a student is, both with peers and instructors, the better their chances of online success are. If you have the data available, it would be worth trying to incorporate it into your model dataset to see whether or not it is predictive.

Rather than using retention as a y-variable when building these models, we typically create an attrition variable (exactly the opposite of retention) and use that as our y instead. This way, we're getting more directly at the characteristics of a student who is likely to leave rather than stay.

Typically when building attrition models, I create separate models for freshmen and upperclassmen. I’d suggest doing that here as well, since previous online coursework will probably be a good indicator of future online coursework. In that case, you’d want to take out many of the variables listed above when modeling freshmen retention.

Finally, it’s important to keep in mind that student success has different meanings for different institutions. You could be basing success on # of credits completed, transitions from semester to semester, or a particular GPA cutoff, among other indicators. When building these different types of student success models, you will probably need to tailor some of these variables to fit the model you're building.

-Caitlin Garrett is a Statistical Analyst at Rapid Insight

Tuesday, May 21, 2013

Using Predictive Modeling to Drive Fundraising Efforts


In preparation for their presentation at our upcoming User Conference, "Using Predictive Modeling to Focus your Fundraising Efforts", I got the chance to chat with Bridget Mendoza and Brianna Lowndes from the Whitney Museum of American Art. Bridget, the Director of Development Records, and Bri, Director of Membership and Annual Fund, have been working together for the past year and a half to bring predictive modeling in-house for the Whitney Museum. 

Here are their thoughts on building their skillsets, modeling challenges, and how the process is going so far:

Bridget Mendoza
What triggered your interest in predictive modeling for the Whitney Museum?

BM – We started by thinking about how to enhance our prospecting as we lead up to our new building. Our research team routinely identifies ‘hidden’ people with higher capacities giving at an entry levels. Anecdotally we compared these prospects to active upper level donors and started seeing patterns in some of their giving and membership histories. We’d previously completed a modeling exercise with an outside company, but the problem with outsourcing was that once we got the model we didn’t have ownership and couldn’t adjust it. We know our data better than anyone else, and when we looked at some of the underlying information, we wanted the ability to alter and refine the model. As our goals are ambitious, we needed to continually grow our prospect base and predictive modeling helps us create a solid foundation for doing so.

How did you decide internally who would take on the predictive modeling project?

Brianna Lowndes
BL - We formed a committee of about ten who were involved in conversations on what we hoped to get out of a predictive modeling software or service and what our goals would be. As conversations progressed and we decided on the Rapid Insights tools it made sense from a resource perspective to deploy Bridget and I, who already work closely with the data and provide different perspectives.  Being close to the exports and metrics and being aware of nuances in member lifecycles has played really nicely into the work we’re doing in predictive modeling. The larger group meets quarterly and that cross-departmental approach helps keep us on track and engaged with the bigger picture.

How did you build your predictive modeling skillset?

BM – We started by attending conferences like MARC and the Rapid Insight modeling course at Brown. Once we made the decision to work with Rapid Insight we had the opportunity to work closely with their team and to become more familiar with their software and with basic modeling practices.  Bri and I also took a Business Statistics for Management class as a refresher.

What modeling challenges have you found that are unique to a museum?

BM –Museums are not as far along in leveraging predictive modeling as our Higher Education counterparts.  While attending the RI User Conference, we heard a really interesting presentation about student retention which sparked our thinking on how to apply what they’ve done to the museum setting.  Like many museum membership programs, our acquisitions in a given year are connected to the exhibition schedule. These cyclical patterns make it more complicated to isolate the data around the health of the program. We are excited to leverage predictive modeling tools to better understand those trends. 

Do you have any advice for non-profits who are thinking about predictive modeling in-house?

BM – There’s a learning curve, but don’t let that discourage you. That’s what a lot of our webinar will be about. Even though we haven’t been modeling for five or ten years, there’s a lot that that can be accomplished in that first year especially with the help from a partner like Rapid Insight.

BL – It’s important to have senior leadership support and take a cross-departmental approach. This ensures we are always thinking about the larger institutional needs. I’d also say that taking the stats class was helpful for us. The software does a lot of the heavy lifting for you so it’s important to get up to speed so that you feel like you’re engaging critically and asking good questions. 

BM – Having a vendor who had built this type of model before and had reliable expertise in the both the non-profit and for-profit fields has been really helpful for us. It was good to be able to collaborate with our software’s support team to build up our own knowledge of data prep and predictive modeling. Rapid Insight has been a real partner through this first year of modeling and we are excited to continue and expand this great work.   

**
If you're interested in hearing more about how to use predictive modeling to focus your fundraising efforts, Bridget and Bri are presenting at our upcoming User Conference. For more information, or to register, click here. Both users and non-users are welcome to attend.

If you have a tip you'd like to share on using predictive models to drive your fundraising efforts, please leave it as a comment below :)

Tuesday, April 30, 2013

Using Social Media Data

Every minute, millions of pieces of social media data are generated around the world. In any given minute*:

Instagram users share 3,600 photos
Brands and organizations on Facebook receive 34,722 “likes”
Twitter users send over 100,000 tweets

With millions of people sharing more information each day, we are witnessing a shift in how information is being produced online as it becomes more and more user-generated. The web is moving away from being a static library and becoming a more open, more connected place to share content. There are plenty of opportunities to mine new variables from the constantly increasing expanse of social media data. Here are a few:

  • For a college or non-profit organization’s advancement department, tapping into these variables can provide insight into how connected an individual is with the institution. If a constituent is following you (whether on Facebook, LinkedIn, Twitter, or another forum), they are actively choosing to be connected to you, which conceivably may make them more likely to donate.
  • In a college enrollment office, tracking which high school students have liked their Facebook page can give them an idea of students who are very interested in their institution. This information could be used to qualify prospects or to help shape the pool of students being marketed to.
  •  Brands looking to decide who to market to can analyze both the source and content of social media data to determine who might be most likely to buy a product or respond to a campaign. 

Tracking and leveraging these data points has the potential to add value to the predictive models you’re already building.  

So, are you already leveraging your social media data into your models, and if not, why not?

Tuesday, March 5, 2013

Six Predictive Modeling Mistakes

As we mentioned in our post on Data Preparation Mistakes, we've built many predictive models in the Rapid Insight office. During the predictive modeling process, there are many places where it's easy to make mistakes. Luckily, we've compiled a few here so you can learn from our mistakes and avoid them in your own analyses:

Failing to consider enough variables
When deciding which variables to audition for a model, you want to include anything you have on-hand that you think could possibly be predictive. Weeding out the extra variables is something that your modeling program will do, so don’t be afraid to throw the kitchen sink at it for your first pass.

Not hand-crafting some additional variables
Any guide-list of variables should be used as just that – a guide – enriched by other variables that may be unique to your institution.  If there are few unique variables to be had, consider creating some to augment your dataset. Try adding new fields like “distance from institution” or creating riffs and derivations of variables you already have.

Selecting the wrong Y-variable
When building your dataset for a logistic regression model, you’ll want to select the response with the smaller number of data points as your y-variable. A great example of this from the higher ed world would come from building a retention model. In most cases, you’ll actually want to model attrition, identifying those students who are likely to leave (hopefully the smaller group!) rather than those who are likely to stay.

Not enough Y-variable responses
Along with making sure that your model population is large enough (1,000 records minimum) and spans enough time (3 years is good), you’ll want to make sure that there are enough Y-variable responses to model. Generally, you’ll want to shoot for at least 100 instances of the response you’d like to model.

Building a model on the wrong population
To borrow an example from the world of fundraising, a model built to predict future giving will look a lot different for someone with a giving history than someone who has never given before. Consider which population you’d eventually like to use the model to score and build the model tailored to that population, or consider building two models, one for each sub-group.

Judging the quality of a model using one measure
It’s difficult to capture the quality of a model in a single number, which is why modeling outputs provide so many model fit measures. Beyond the numbers, graphic outputs like decile analysis and lift analysis can provide visual insight into how well the model is fitting your data and what the gains from using a model are likely to be.

If you’re not sure which model measures to focus on, ask around. If you know someone building models similar to yours, see which ones they rely on and what ranges they shoot for. The take-home point is that with all of the information available on a model output, you’ll want to consider multiple gauges before deciding whether your model is worth moving forward with.  

-Caitlin Garrett, Statistical Analyst at Rapid Insight
Photo Credit: http://www.flickr.com/photos/mattimattila/


Have you made any of the above mistakes? Tell us about it (and how you found it!) in the comments. 

Tuesday, January 29, 2013

Four Years of Predictive Modeling and Lessons Learned

I recently got the chance to talk with Dr. Michael Johnson, Director of Institutional Research at Dickinson College, about his experiences with predictive modeling over the past four years. Dr. Johnson will be presenting a free webinar, "Four Years of Predictive Modeling and Lessons Learned", on Thursday, January 31st at 2pm EST in which he'll provide a more in-depth look at his experiences. 

Can you give us an example of a lesson you’ve learned through your experiences with predictive modeling?

I’ve learned that predictive modeling is good but predictive modeling in real time is just more extremely beneficial. When we picked up Rapid Insight, we moved a five day turnaround time to an eight minute turnaround. That’s one of the biggest changes we’ve made, and the effects have been very apparent.

If there was another thing I’ve learned, it is to automate absolutely every process possible to remove the opportunity for human error.

What types of predictive models will you be discussing during your webinar?

The enrollment management model is our primary model but a close cousin to that is the one we’ve been using for retention. The dataset is basically the same only slightly enhanced. It’s good to use essentially the same dataset to solve two different problems.

How has predictive modeling changed the way you operate?

It is the primary tool that we use to make decisions on the incoming class. This last week has been an incredible example of that. We’re taking a look at our early action pool and asking questions: What does it look like? How does it compare with previous years? What if we swap out some people; how does that change our incoming class?

What do you hope attendees will take away from your webinar?

There’s really no need to reinvent the wheel, so I’ll share some ideas that I’ve picked up. I hope that others come on board and share their successes as well. We all have the same problem set, so it will be nice for others to take away a few things that I’ve seen that have given me success. It would be great if they had ideas that they wanted to share with others as well. 

To register for Dr. Johnson's free webinar, "Four Years of Predictive Modeling and Lessons Learned", or for more information, please click here

To read a case study about how Dickinson College uses predictive modeling for strategic enrollment management, please click here

Wednesday, January 16, 2013

How to Interpret a Decile Analysis


After building a predictive model, there are several ways to determine how well the model is describing your data. One visual way to get an idea of how well a model is fitting your data is by taking a look at the decile analysis. Here we’ll take a look at what the decile analysis represents, how it’s created, and how to spot a good model.

What a Decile Analysis Represents

After building a statistical model, a decile analysis is created to test the model’s ability to predict the intended outcome. Each column in the decile analysis chart represents a collection of records that have been scored using the model. The height of each column represents the average of those records’ actual behavior.

How the Decile Analysis is Calculated

1. The hold-out or validation sample is scored according to the model being tested.
2. The records are sorted by their predicted scores in descending order and divided into ten equal-sized bins or deciles. The top decile contains the 10% of the population most likely to respond and the bottom decile contains the 10% of the population least likely to respond, based on the model scores.
3. The deciles and their actual response rates are graphed on the x and y axes, respectively. 

After the decile analysis is built, you’ll want to take a look at the height of the bars in relation to one another. Deciding whether a model is worth moving forward with depends on the pattern you see when viewing the decile analysis. 


Ideal Situation: The Staircase Effect

When you’re looking at a decile analysis, you want to see a staircase effect; that is, you’ll want the bars to descend in order from left to right, as shown below. 

This is telling you that the model is “binning” your constituents correctly from most likely to respond to least likely to respond. A model exhibiting a good staircase decile analysis is one you can consider moving forward with.

Not-So-Ideal Situations

In contrast, if the bars seem to be out of order (as shown below), the decile analysis is telling you that the model is not doing a very good job of predicting actual responses.


 If the bars seem to be the same height, or the decile analysis looks “flat”, the decile analysis is telling you that the model isn’t performing any better than randomly binning people into deciles would. In both cases, your model should be improved before moving forward with it.  

-Caitlin Garrett, Statistical Analyst at Rapid Insight


Thursday, January 10, 2013

Valuing Analytics & Predictive Modeling in Higher Ed

As promised, here is part two of my interview with Mike Laracy, Founder, President, and CEO  of Rapid Insight. Mike's 20+ years of data analytics & predictive modeling experience have provided him with many insights. Here's Mike on becoming more data-driven in higher education, which models produce the highest ROI, and mistakes to avoid:

Where does predictive modeling fit into the analytic ecosystem in higher education?

Within the analytic ecosystem in higher ed, there is a range of ways in which data is analyzed and looked at. On one side, you have historical reporting, which our clients do a lot of and is vital to every institution.  Somewhere in the middle is data exploration and analysis, where you’re slicing and dicing data to understand it better or make more informed decisions based on what happened in the past.  On the other side of the spectrum is predictive modeling.  Modeling requires taking a look at all of the variables in a given set of information to make informed predictions about what will happen in the future. What is each applicant’s probability of enrolling or what is each student’s attrition likelihood?  What will the incoming class look like based on the current admit pool?  These are the types of questions that are being answered in higher ed with predictive analytics.  The resulting probabilities can also be used in the aggregate. For example, enrollment models allow you to predict overall enrollment, enrollment by gender, by program, or by any other factor.  The models are also used to project financial outlay based on the financial aid promised to admitted applicants and their individual enrollment probabilities.

Higher education has come a long way in the last five to ten years in its use of predictive analytics. The entire student life cycle is now being modeled starting with prospect and inquiry modeling all the way through to alumni donor modeling.   It used to be that any institutions that were doing this kind of modeling were relying on outside consulting companies.  Today most are doing their modeling in-house.  Colleges and universities view their data as a strategic asset and they are extracting value from their data with the same tools and methodologies as the Fortune 500 companies.

What kinds of resources are needed and what is the first step for an institution who wants to become more data-driven in their decision making?

It’s important to have somebody who knows the data. As long as a user has an understanding of their data, our software makes it very easy to analyze data and build predictive models very quickly. And our support team is available to answer any analytic questions. 

Gaining access to their data is the first step. We see a lot of institutions that have some reporting tools which don’t allow them to ask new questions of the data. So, they might have a set of 50 reports that they’re able to run over and over but anytime someone has a new question, without access to the raw data there’s no way to answer the question. 

It really helps if the institution is committed to a culture of data driven decision making.  Then all the various stakeholders are more focused on ensuring data access for those doing the predictive modeling.

What do you say to those who are on “the quest for perfect data”?  Is it okay to implement predictive analytics before you have that data warehouse or those perfectly cleansed datasets?

No institution is ever going to have perfect data, so you work with what you have. We suggest seeing what you have, finding any obvious problems in the data, and then fixing those problems the best you can. We’ve designed our solutions such that a data warehouse is not required but, even with a clean data warehouse, the data is never going to be perfect.   As long as you as you have an understanding of the data, you can move forward.  

In your experience, which models in higher education produce the highest ROI?
We have a customer, Paul Smith’s College that has quantified their retention modeling efforts. Using their model results, they put programs into place to help those students that were predicted to be high-risk of attrition. They credit the modeling with helping them identify which students to focus on, saving them $3m in net tuition revenue so far.

We have other clients that are using predictive modeling on the prospect side and they’re realizing significant savings on their recruiting efforts. So instead of mailing to 200,000 high school seniors, they’re mailing to 50,000, and realizing significant savings by not mailing and not calling those students who have pretty much zero probability of applying or enrolling.

Although not as easily quantifiable, enrollment modeling has a pretty big ROI.  Not only on determining which applicants are likely to enroll, but in predicting class size.  If an institution overshoots and enrolls too many applicants, they’ll have dorm, classroom, and other resource issues.  If enroll too little, they’ll have revenue issues.  So predicting class size and determining who and how many applicants to admit is extremely important.

What are some common mistakes you see when approaching predictive modeling for your higher ed customers?

One mistake that I often see is when information is thrown out as not useful to the models.  Zip code is a good example.  Zip code looks like a five digit numeric variable, but you wouldn’t want to use it as a numeric variable in a model.  In some cases it can be used categorically to help identify applicants’ origins, but its most useful purpose is to for calculating a distance from campus variable.  This is a variable that we see showing up as a predictor in many prospect/ inquiry models, enrollment models, alumni models, and even retention models.  Another example of a variable that is often overlooked is application date.  Application date often contains a ton of useful information if looked at correctly.  It can be used to calculate the number of days between when the application was sent and the application deadline.  This piece of information can tell you a lot about an applicant’s intentions.  A student who gets their application in the day before the deadline probably has very different intentions than a student who applies nine months before the deadline.  This variable ends up participating in many models. 

To get our customers up to speed on best practices in predictive modeling we’ve created resources like lists of recommended variables for specific models and guides on how to create useful new variables from existing data.

Tuesday, December 11, 2012

Five Data Preparation Mistakes (and How to Avoid Them!)

After building many predictive models in the Rapid Insight office and helping our customer build many more models outside of the office, we have a list of data preparation mistakes that could fill a room. Here are some of the most common ones we've seen:


1. Including ID Fields as Predictors
Because most IDs look like continuous integers (and older IDs are typically smaller), it is possible that they may make their way into the model as a predictive variables. Be sure to exclude them as early on in the process as possible to avoid any confusion while building your model.

2. Using Anachronistic Variables
Make sure that no predictor variables contain information about the outcome. Because models are built using historical data, it is possible that some of the variables you have accessible when building your model were not available at the time the model is built to reflect. No predictor variables should be proxies for your dependent variable (ie: “made a gift” = donor, “deposited” = enrolled).

3. Allowing Duplicate Records
Don’t include duplicates in a model file. Including just two records per person gives that person twice as much predictive power. To make sure that each person’s influence counts equally, only one record per person or action being modeled should be included. It never hurts to dedupe your model file before you start building a predictive model. 

4. Modeling on Too Small of a Population
Double-check your population size. A good goal to shoot for in a modeling dataset is at least 1,000 records spanning three years. Including at least three years helps to account for any year-to-year fluctuations in your dataset. The larger your population size is, the most robust your model will be. 

5. Not Accounting for Outliers and/or Missing Values
Be sure to account for any outliers and/or missing values. Large rifts in individual variables can add up when you’re combining those variables to build a predictive model. Checking the minimum and maximum values for each variable can be a quick way to spot any records that are out of the usual realm. 

[photo credit]

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, November 20, 2012

How to Score a Dataset Using Analytics Only

Since we’ve already covered how to score a dataset using Veera, it’s only fair that we show you how to score using the Analytics Scoring program. We’ll start at the point where you save your scoring model within Analytics. After memorizing your model in the Model tab, you’ll want to move down to the Compare Models tab. This tab allows you to compare any two models side-by-side. Once you’ve decided which model you like better, you’re ready to save it by selecting the model and clicking the “Save Scoring Model” as button, as shown below.


Analytics will prompt you to navigate to where you’d like the file to be saved, and will save it with a .rism (Rapid Insight Scoring Model) extension. After saving the .rism file, you’ll want to open the Analytics Scoring Module by going to your Start Menu and navigating to Rapid Insight Inc. -> Analytics -> Scoring, as shown below. 


Once inside the scoring module, you’ll need to click the “Select Dataset” button and navigate to where the dataset you’d like to score is located on your machine. After loading in your dataset, you’ll see all of the variables within it populate the ‘Dataset Variables’ window. Next, you’ll need to click the “Select Scoring Model” button and navigate to where the scoring model (.rism) file you’d like to use is located. Once you find the model, its equation will show up in the corresponding window.


Before you start the scoring process, you have a couple of options detailing how you’d like the model to be scored. The first option, shown above in the green box, allows you to validate the model by looking at the decile analysis resulting from the scoring process. The second option, shown in the blue box, allows you to output the scores as well as the corresponding deciles or percentiles. After you’ve selected the appropriate options, click on the “Start Scoring” button, decide where you’d like your scores to output, and Analytics will score your dataset in the way that you request. 

-Caitlin Garrett, Statistical Analyst at Rapid Insight