Showing posts with label analysis. Show all posts
Showing posts with label analysis. Show all posts

Thursday, July 11, 2013

#RIUC13

For those of you who weren’t able to attend the 2013 Rapid Insight User Conference, we set a new record for most attendees and largest number of customer presentations. With two full days of dual track programming, the presenters covered a lot of ground. While we wait for some of the video recordings of customer presentations to be formatted, I thought it would be good to do a quick recap here. 

Mike Laracy, Data Geek (at right)
The conference opened with a keynote from our Founder and CEO, Mike Laracy, who talked a bit about the future of predictive analytics. With a mass public education on the value of analytics (from people like Nate Silver and Billy Bean, with a little help from Brad Pitt), as well as significant advances in data storage and processing power, a stronger need for predictive analytics is emerging. The market is shifting towards the view that more data access is better than restricted access, and that given the right tools along with access, smart people – data scientists – can turn raw data into actionable information. Given these changes, the data scientist – that’s you – will be in increasingly higher demand over the next decade and beyond, as will predictive analytics. 

The user presentations covered lots of different topics, and we’ve made all of their slide decks available here; I’d highly recommend checking them out. In addition to what’s there, I’d also recommend checking out some of the interviews we’ve done with customers on building campaign pyramids and using predictive modeling to drive fundraising efforts. The RI staff team also gave a few presentations,  including topics like Tips and Tricks in Veera, Techniques for Improving Your Predictive Models, and An Introduction to Reporting and Dashboarding with Veera.

Another thing worth mentioning is that we announced our partnership with Tableau to provide a complete solution for both predictive modeling and visualization. Now users can use Veera to clean up their data, Analytics to build their predictive models, and Tableau’s visualizations to turbocharge their presentations. For more information, check out our partner page.
My favorite part of the User Conference has always been talking to customers about the cool data projects that they’ve been tackling, and this year was no different. Kudos to our users for being so creative and smart with the ways they use our software. We also owe a big thanks to the folks at Yale for hosting us, and to all who were able to attend. Here’s to the best User Conference so far and to making next year’s even better!

Tuesday, April 23, 2013

Five Steps for Data-Driven Strategic Enrollment Management


Establish your goals
This first step is crucial for mapping out a course to your end goal. Try to envision where you’d like to end up and formulate a specific goal to help get you there.

Possible goals include:
  • Reduce your prospect mailing budget
  • Increase accuracy of enrollment yield predictions
  • Meet diversity objectives
  • Increase your retention rate

Get to know your data
The first step to getting to know your data is gaining access to your data, which is trickier for some people than others. If you have to go through IT to access your data, it helps to have a clear goal in mind and a good idea of what fields or tables you’ll need.

Once you have your data, you’ll need some time to get well-acquainted. A good starting point is to make sure you understand what each field represents and how things are coded. If you have questions about how data is being recorded or stored, this is the time to ask. Once you have a handle on what your data represents, you’ll want to thoroughly review it.

A few suggestions:
  • Spot-check the accuracy of your data. Double-checking things like the mean, min, and max for each variable is a quick way to verify accuracy.  If you spot any data quality issues, do your best to resolve them sooner than later.
  • Check for missing values. If you have a variable with a high number of missings, you’ll need to decide whether or not to use that variable and if there’s a way to fill in what’s not there.
  • Brainstorm ideas for new variables. If you can’t create new variables from what you have on-hand, spend some time thinking about things that might be worth tracking going forward. 

Analyze your data
I realize that the word “analyze” represents a whole spectrum of techniques and applications – and that’s okay. In a general sense, you’ll want to see if fields in your dataset can give you some insight that you can relate back to your initial goal(s).

Some ideas:

  • Look at correlations within your dataset. Are they positive or negative? Large or small?
  • Look for the differences between your target and non-target population, variable by variable.
  • Visuals help! Graphs are a great way to get a feel for the relationships between your variables.
  • Try building a predictive model. The results you get will be more directly applicable to driving decisions. 
You may get some surprising results during the analysis phase. I’ve worked on projects where the end insight was the exact opposite of what was expected. Although sometimes the results can be surprising, it’s important to let your data tell its story. The other side of analysis is that your data can confirm what you’ve long-suspected to be the truth – whether it’s that students from Montana are more likely to enroll, or that the number of first term credits impacts a student’s likelihood of attrition – embrace these confirmations and continue to rely on that information. 

Turn analysis into insight
Keep your initial question in mind, take what you’ve learned from your analysis, and apply it going forward. The idea here is to replace outdated anecdotal evidence with insights from our data. If your goal was to save money on prospect marketing efforts, use the factors that correlate to a higher response rate to drive your decisions about who will receive the next round of direct mail. If you’re trying to improve retention rate, target those students who look most like previously dropped students and reach out to help keep them on campus.

Assess your decisions
Last, but certainly not least, don’t forget to circle back and re-assess your decisions. If you feel like you’re not making progress toward your initial goal, consider re-framing it or breaking it down into more manageable phases. If you feel good about the progress you’re making, start working on new goals. A data-driven decision should be sustainable under conditions similar to the past. Don’t be afraid to revisit past goals if you feel like you can improve or add something to your initial recommendation. 

...Did we miss anything? Have questions about becoming more data-driven? Leave them in the comments below.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Thursday, January 10, 2013

Valuing Analytics & Predictive Modeling in Higher Ed

As promised, here is part two of my interview with Mike Laracy, Founder, President, and CEO  of Rapid Insight. Mike's 20+ years of data analytics & predictive modeling experience have provided him with many insights. Here's Mike on becoming more data-driven in higher education, which models produce the highest ROI, and mistakes to avoid:

Where does predictive modeling fit into the analytic ecosystem in higher education?

Within the analytic ecosystem in higher ed, there is a range of ways in which data is analyzed and looked at. On one side, you have historical reporting, which our clients do a lot of and is vital to every institution.  Somewhere in the middle is data exploration and analysis, where you’re slicing and dicing data to understand it better or make more informed decisions based on what happened in the past.  On the other side of the spectrum is predictive modeling.  Modeling requires taking a look at all of the variables in a given set of information to make informed predictions about what will happen in the future. What is each applicant’s probability of enrolling or what is each student’s attrition likelihood?  What will the incoming class look like based on the current admit pool?  These are the types of questions that are being answered in higher ed with predictive analytics.  The resulting probabilities can also be used in the aggregate. For example, enrollment models allow you to predict overall enrollment, enrollment by gender, by program, or by any other factor.  The models are also used to project financial outlay based on the financial aid promised to admitted applicants and their individual enrollment probabilities.

Higher education has come a long way in the last five to ten years in its use of predictive analytics. The entire student life cycle is now being modeled starting with prospect and inquiry modeling all the way through to alumni donor modeling.   It used to be that any institutions that were doing this kind of modeling were relying on outside consulting companies.  Today most are doing their modeling in-house.  Colleges and universities view their data as a strategic asset and they are extracting value from their data with the same tools and methodologies as the Fortune 500 companies.

What kinds of resources are needed and what is the first step for an institution who wants to become more data-driven in their decision making?

It’s important to have somebody who knows the data. As long as a user has an understanding of their data, our software makes it very easy to analyze data and build predictive models very quickly. And our support team is available to answer any analytic questions. 

Gaining access to their data is the first step. We see a lot of institutions that have some reporting tools which don’t allow them to ask new questions of the data. So, they might have a set of 50 reports that they’re able to run over and over but anytime someone has a new question, without access to the raw data there’s no way to answer the question. 

It really helps if the institution is committed to a culture of data driven decision making.  Then all the various stakeholders are more focused on ensuring data access for those doing the predictive modeling.

What do you say to those who are on “the quest for perfect data”?  Is it okay to implement predictive analytics before you have that data warehouse or those perfectly cleansed datasets?

No institution is ever going to have perfect data, so you work with what you have. We suggest seeing what you have, finding any obvious problems in the data, and then fixing those problems the best you can. We’ve designed our solutions such that a data warehouse is not required but, even with a clean data warehouse, the data is never going to be perfect.   As long as you as you have an understanding of the data, you can move forward.  

In your experience, which models in higher education produce the highest ROI?
We have a customer, Paul Smith’s College that has quantified their retention modeling efforts. Using their model results, they put programs into place to help those students that were predicted to be high-risk of attrition. They credit the modeling with helping them identify which students to focus on, saving them $3m in net tuition revenue so far.

We have other clients that are using predictive modeling on the prospect side and they’re realizing significant savings on their recruiting efforts. So instead of mailing to 200,000 high school seniors, they’re mailing to 50,000, and realizing significant savings by not mailing and not calling those students who have pretty much zero probability of applying or enrolling.

Although not as easily quantifiable, enrollment modeling has a pretty big ROI.  Not only on determining which applicants are likely to enroll, but in predicting class size.  If an institution overshoots and enrolls too many applicants, they’ll have dorm, classroom, and other resource issues.  If enroll too little, they’ll have revenue issues.  So predicting class size and determining who and how many applicants to admit is extremely important.

What are some common mistakes you see when approaching predictive modeling for your higher ed customers?

One mistake that I often see is when information is thrown out as not useful to the models.  Zip code is a good example.  Zip code looks like a five digit numeric variable, but you wouldn’t want to use it as a numeric variable in a model.  In some cases it can be used categorically to help identify applicants’ origins, but its most useful purpose is to for calculating a distance from campus variable.  This is a variable that we see showing up as a predictor in many prospect/ inquiry models, enrollment models, alumni models, and even retention models.  Another example of a variable that is often overlooked is application date.  Application date often contains a ton of useful information if looked at correctly.  It can be used to calculate the number of days between when the application was sent and the application deadline.  This piece of information can tell you a lot about an applicant’s intentions.  A student who gets their application in the day before the deadline probably has very different intentions than a student who applies nine months before the deadline.  This variable ends up participating in many models. 

To get our customers up to speed on best practices in predictive modeling we’ve created resources like lists of recommended variables for specific models and guides on how to create useful new variables from existing data.

Friday, October 5, 2012

Fundraising: The Science


Now that we’ve discussed the art of fundraising, I think it’s only right that we focus a little bit on the science. After all, knowing which prospects are statistically most likely to give makes a gift officer’s contribution to the art of fundraising that much more successful. As I’ve mentioned before, my function in the fundraising spectrum is as an analyst, helping customers build models identifying which prospects are most likely to donate.

One of the most important things we do during the predictive modeling process is data preparation, which often means creating new variables from the data our customers have on-hand. I’d like to discuss some of these variables, as well as how and why to include them in a fundraising or advancement model. For the purposes of this blog entry, I’ll use a higher education example. Typically a higher education institution might have some extra variables, but these can be tailored to fit other institutions or excluded when not relevant.

Demographic Information          
It’s always smart to have an idea of what each donor looks like at a demographic level. Variables to include here are things like age, gender, marital status, and any occupational data you might have available to you. In a higher education context, this would also include things like the constituent’s class year, major, whether or not their spouse is an alumni, and whether they are a legacy alumni (meaning a parent or grandparent also attended the institution). Additionally, we often create a “reunion year flag” indicating if the analysis year is a reunion year for that person, as donors are often more likely to give (and give larger gifts) during a reunion year.

Location Information
General information about each donor’s location like ZIP code, city, and county can be useful as categorical variables (treating people that live in each one as a group). Once we have a ZIP code, we always calculate a “distance from institution” variable using one of Veera’s pre-programmed functions. This new variable, which is measured in miles, gives you a solid idea of the relationship between location and giving. If you have access to census data, we recommend appending variables relating to neighborhood or housing type. Creating flag variables for wealthy neighborhood ZIP codes can also be useful; constituents coming from these areas may be more likely to give. Although this can be created at a more local level, we often start with Forbes’ list of the top 500 wealthiest ZIP codes in the US, which is available online at http://www.forbes.com/lists/2011/7/zip-codes-11_rank.html.

Contact History
The ways in which a donor engages with you can tell you a lot about their likelihood of giving. For starters, include variables pertaining to their event history. How many events have they attended? Which types of events are they attending? How many days since their last event? Answers to questions like these can sometimes turn out to be predictive of giving. This is also where your social media variables come into play; create flags for whether a constituent is following you on LinkedIn, Facebook, Twitter, Pinterest, etc. A donor following you on one or more of these sites is an indication that they want to be connected, and therefore they may be more likely to give. Conversely, if a constituent has indicated that they do not want to be contacted, you’ll want to include this information as well, as it can be very predictive.

Gift History
This brings us to our last and most predictive set of variables: giving history. These variables should answer all kinds of questions about what a giver looks like historically, like:

  • How many gifts have they given in their lifetime?
  • What was their last gift?
  • How many days since their first gift? How many days since their last gift?
  • Have they given in the past 12 months? If so, how much?
  • What is the velocity of the gifts - are the increasing, decreasing, or staying the same?

One thing to note here is that gift dates themselves aren’t useful in a predictive model, but their translations – like the number of days since an event – allow us to use the insight they provide.

In building a predictive model, some of these variables may be predictive, while others might turn out not to be. It’s a good idea to include some combination of these variables, plus anything you have on-hand that you think could possibly be predictive.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Friday, July 20, 2012

The Forgotten Tabs: Means Analysis


Continuing with the Forgotten Tabs series, the next tab we’ll be focusing on is the Means Analysis tab. The Means Analysis tab provides the mean, number of observations, maximum values, and minimum values for any of the variables in your dataset. You also have the option to take “means by” and “subclass by” to view the means of multiple subcategories across variable combinations.

 In this case, we’re comparing the attrition rates of legacy and non-legacy students by whether or not they received financial aid. For each of these four possible categories, we are able to see the mean attrition rate, the number of observations, and the min and max values. Doing so allows us to see the differences in attrition rate over a couple of different characteristics.

The Means Analysis tab can be also useful in comparing data from different cohorts or years in order to spot trends. In the example below, we’re comparing attrition rates by year, which allows us to pick up on any trends or changes that are occurring from year to year. If for some reason we were noticing a year that had a much higher or lower attrition rate than the other years, this gives us the opportunity to pick up on that and investigate further as to why that might be. 


You might also note that beyond looking at ‘Attrition’, we are also looking at a variable called ‘Predicted Attrition’. This variable represents the predicted attrition probabilities that we’ve assigned to each student. In this case, we’ve grouped these values by year to get an idea of how well we’re predicting attrition for that year compared to the actual attrition rate. Comparing our predicted values to actual values gives us a sense of any weaknesses from year to year that our predictive model might have. If we do find any weakness in predictive ability, we have the opportunity to go back and further fine-tune our model in order to incorporate our findings. 

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Friday, July 6, 2012

The Forgotten Tabs: Frequency Analysis


During this year’s User Conference, I gave a presentation called “Analytics: The Forgotten Tabs”, which I’ve decided to expand into a blog series. The purpose of this series will be to explain how and why to use four of the lesser-known tabs in Analytics – Frequency Analysis, Means Analysis, Correlation Analysis, and Profiling Analysis. Each entry will focus on one of these tabs and we’ll start with the Frequency Analysis tab.

The Frequency Analysis tab’s output is actually fairly simple; it gives you the frequency of occurrence for any binary or categorical variable. For a single variable, it will output counts and percentages for each value of that variable. It is also capable of creating two-way cross frequencies, which output raw numbers, as well as row, column, and total percentages. 




While Frequency Analysis isn’t actually performing any statistical test – its functions are simple summing and percentage operations – it is providing valuable information about the number and percentage of observations that fall into each sub-category of a binary or categorical variable. Using this tab gives you a quick by-the-numbers glance at variables like “Ethnicity” or “Department”, which allows you to instantaneously compare subcategories without doing any manual addition or division. This is particularly useful when you’re working with a variable such as “Department” that may have a lot of sub-categories.




One other little-known fact about the output from Frequency Analysis (and other tabs) is that you can save it to the Report Bar the same way you would a graph or chart. To do so, click on the ‘Reports’ section of the taskbar and select ‘Launch Report Bar’. 




The report bar will float over your analysis; you can save things to it by dragging the outputs you wish to save into the bar itself. Saving things to the report bar allows you to export them from Analytics in a few different ways. If you select the ‘PPoint’ option before clicking ‘Export’, Analytics will create a PowerPoint such that each of the graphs our outputs you saved will become their own slide in the presentation. The other option you have is to save the information you’re interested in to the Reports tab in Analytics (by selecting the ‘Report’ option on the Report Bar), which allows you to create custom reports within the program and export these reports as Word Documents to be used later on. In any case, there are a number of ways to take the information that you’re getting from Analytics and use it in a presentation or report down the line.  

-Caitlin Garrett, Statistical Analyst at Rapid Insight