Showing posts with label enrollment. Show all posts
Showing posts with label enrollment. Show all posts

Friday, July 26, 2013

Predicting Retention for Online Students: Where to Start

With the rise of enrollment in online programs and MOOCs, we’re seeing more and more students forego traditional classroom experiences in favor of more flexible online programs. With this shift comes a whole new set of guidelines for enrollment management, financial aid, and retention programs. Retention, in particular, has seen a significant downward trend as learning moves from in-person to online classrooms.


My interest lies in figuring out what variables might be worth including in an analysis attempting to predict online student retention. I did a bit of research and was hoping to find a list of variables online that had worked in the past but couldn’t find any comprehensive resource, so I’ve started to build my own. In the sections below, I’ve listed the type of information that I think would be worth analyzing broken out into four separate categories. Some of these are variables in and of themselves, and some can be broken down different ways; for example, “age” can be used by itself, but creating a “non-traditional age” flag is useful as well. Realistically, not all schools will have all of this information, so this list is meant to be a good starting point of what to shoot for when collecting data.


Also, if you have any variables to add (and I’m sure there are some I’ve missed), I’d love to hear about them in the comments. 

Student Demographic Information
  • Socioeconomic status / financial aid information
    • FAFSA info, Pell eligibility, any scholarship or award info
  • Ethnicity
    • Minority Status
  • Gender
  • Home state
  • Distance from physical campus (if applicable)
  • Age; traditional or non-traditional?
  • Military background?
  • Have children?
  • Currently employed full-time?
  • First generation college student?
  • Legacy student? (Did a parent/grandparent/sibling attend?)

Student Online Learning History
  • Registered for classes online or in person?
  • How many days did they register before the start of the term?
  • Ever attended a class on-campus?
  • Do they plan to attend both online and on-campus classes?
  • Did they attend any type of orientation?
  • Number of previous online courses taken
    • First-time online learner?

Student Academic History
  • GPA
  • SAT/ACT scores
  • Degree hours completed
  • Degree hours attempted
  • Taking developmental courses?
  • Transfer student?
  • Degree program / major 
  • Program level (Associate, Bachelors, Masters, etc.)
  • Number of program or major changes (if applicable)
  • Any previous degrees?

Course- and Program- Related
  • Amount of text vs. interactive content 
  • Lessons with immediate feedback?
  • Any peer-to-peer forum for interaction?
  • Lessons in real time or recorded?
  • Amount of teacher interaction with students
    • Chat, email exchange, turn-around time on assignments

Closing notes:

Getting course-related data might be difficult, but the variables I listed above are derived from studies about how to improve online courses as being areas to focus on; my thinking is that the more engaged a student is, both with peers and instructors, the better their chances of online success are. If you have the data available, it would be worth trying to incorporate it into your model dataset to see whether or not it is predictive.

Rather than using retention as a y-variable when building these models, we typically create an attrition variable (exactly the opposite of retention) and use that as our y instead. This way, we're getting more directly at the characteristics of a student who is likely to leave rather than stay.

Typically when building attrition models, I create separate models for freshmen and upperclassmen. I’d suggest doing that here as well, since previous online coursework will probably be a good indicator of future online coursework. In that case, you’d want to take out many of the variables listed above when modeling freshmen retention.

Finally, it’s important to keep in mind that student success has different meanings for different institutions. You could be basing success on # of credits completed, transitions from semester to semester, or a particular GPA cutoff, among other indicators. When building these different types of student success models, you will probably need to tailor some of these variables to fit the model you're building.

-Caitlin Garrett is a Statistical Analyst at Rapid Insight

Tuesday, April 23, 2013

Five Steps for Data-Driven Strategic Enrollment Management


Establish your goals
This first step is crucial for mapping out a course to your end goal. Try to envision where you’d like to end up and formulate a specific goal to help get you there.

Possible goals include:
  • Reduce your prospect mailing budget
  • Increase accuracy of enrollment yield predictions
  • Meet diversity objectives
  • Increase your retention rate

Get to know your data
The first step to getting to know your data is gaining access to your data, which is trickier for some people than others. If you have to go through IT to access your data, it helps to have a clear goal in mind and a good idea of what fields or tables you’ll need.

Once you have your data, you’ll need some time to get well-acquainted. A good starting point is to make sure you understand what each field represents and how things are coded. If you have questions about how data is being recorded or stored, this is the time to ask. Once you have a handle on what your data represents, you’ll want to thoroughly review it.

A few suggestions:
  • Spot-check the accuracy of your data. Double-checking things like the mean, min, and max for each variable is a quick way to verify accuracy.  If you spot any data quality issues, do your best to resolve them sooner than later.
  • Check for missing values. If you have a variable with a high number of missings, you’ll need to decide whether or not to use that variable and if there’s a way to fill in what’s not there.
  • Brainstorm ideas for new variables. If you can’t create new variables from what you have on-hand, spend some time thinking about things that might be worth tracking going forward. 

Analyze your data
I realize that the word “analyze” represents a whole spectrum of techniques and applications – and that’s okay. In a general sense, you’ll want to see if fields in your dataset can give you some insight that you can relate back to your initial goal(s).

Some ideas:

  • Look at correlations within your dataset. Are they positive or negative? Large or small?
  • Look for the differences between your target and non-target population, variable by variable.
  • Visuals help! Graphs are a great way to get a feel for the relationships between your variables.
  • Try building a predictive model. The results you get will be more directly applicable to driving decisions. 
You may get some surprising results during the analysis phase. I’ve worked on projects where the end insight was the exact opposite of what was expected. Although sometimes the results can be surprising, it’s important to let your data tell its story. The other side of analysis is that your data can confirm what you’ve long-suspected to be the truth – whether it’s that students from Montana are more likely to enroll, or that the number of first term credits impacts a student’s likelihood of attrition – embrace these confirmations and continue to rely on that information. 

Turn analysis into insight
Keep your initial question in mind, take what you’ve learned from your analysis, and apply it going forward. The idea here is to replace outdated anecdotal evidence with insights from our data. If your goal was to save money on prospect marketing efforts, use the factors that correlate to a higher response rate to drive your decisions about who will receive the next round of direct mail. If you’re trying to improve retention rate, target those students who look most like previously dropped students and reach out to help keep them on campus.

Assess your decisions
Last, but certainly not least, don’t forget to circle back and re-assess your decisions. If you feel like you’re not making progress toward your initial goal, consider re-framing it or breaking it down into more manageable phases. If you feel good about the progress you’re making, start working on new goals. A data-driven decision should be sustainable under conditions similar to the past. Don’t be afraid to revisit past goals if you feel like you can improve or add something to your initial recommendation. 

...Did we miss anything? Have questions about becoming more data-driven? Leave them in the comments below.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Wednesday, March 20, 2013

Customer Webinar: Predictive Modeling for SEM

Our next customer webinar, "Strategic Enrollment Management: St. Michael's College and Predictive Analytics" will be given by Bill Anderson, CIO of Saint Michael's College today at 2pm EDT and will be re-broadcasted on Tuesday, March 26th, and Thursday, May 2nd

I got the chance to ask him a couple of questions about his session, which will describe the ways in which Veera and Analytics are utilized on campus to produce predictions and other analyses for the scoring team. 

What types of models have you been building?
Almost entirely enrollment management - mostly apply to enroll. We've been building them on and off for about five years now. I have someone on campus that I collaborate with and when we first started, she was using SPSS for the statistical analysis, but we've since abandoned that. 

How has model building changed your Enrollment and/or Financial Aid practices?
There have been a number of ways that we've used the models - one as a sort of verification of what our consultant has been doing, two to be able to do some sensitivity and what-if analysis (and suggest different practices or emphases on where the aid awards should go), and three to help confirm in-semester and in-process prediction on where the class is going to end up. 

In some occasions, this has impacted size of waiting list or the way we thought about awarding wait list spots, including the total number of admits. This last year, our model suggested that we could be more selective than we had been in the past. 

What do you hope attendees will learn from your presentation?
One thing is that you can do it on your own - it's not that hard. You have to have a background that supports responsible interpretation of the results, but you can sit down and do it. That's one element: just do it. I think there's another element that says once you start thinking this way, it can become infectious. In our enrollment management meetings, we have the opportunity to appeal to the data or look at a Veera job that identifies the applicants we could avoid accepting. This changes the internal conversation - from a culture of anecdote, you can change the conversation with data. The use of the products has been fabulous in terms of making the data accessible to people. 

Tuesday, January 29, 2013

Four Years of Predictive Modeling and Lessons Learned

I recently got the chance to talk with Dr. Michael Johnson, Director of Institutional Research at Dickinson College, about his experiences with predictive modeling over the past four years. Dr. Johnson will be presenting a free webinar, "Four Years of Predictive Modeling and Lessons Learned", on Thursday, January 31st at 2pm EST in which he'll provide a more in-depth look at his experiences. 

Can you give us an example of a lesson you’ve learned through your experiences with predictive modeling?

I’ve learned that predictive modeling is good but predictive modeling in real time is just more extremely beneficial. When we picked up Rapid Insight, we moved a five day turnaround time to an eight minute turnaround. That’s one of the biggest changes we’ve made, and the effects have been very apparent.

If there was another thing I’ve learned, it is to automate absolutely every process possible to remove the opportunity for human error.

What types of predictive models will you be discussing during your webinar?

The enrollment management model is our primary model but a close cousin to that is the one we’ve been using for retention. The dataset is basically the same only slightly enhanced. It’s good to use essentially the same dataset to solve two different problems.

How has predictive modeling changed the way you operate?

It is the primary tool that we use to make decisions on the incoming class. This last week has been an incredible example of that. We’re taking a look at our early action pool and asking questions: What does it look like? How does it compare with previous years? What if we swap out some people; how does that change our incoming class?

What do you hope attendees will take away from your webinar?

There’s really no need to reinvent the wheel, so I’ll share some ideas that I’ve picked up. I hope that others come on board and share their successes as well. We all have the same problem set, so it will be nice for others to take away a few things that I’ve seen that have given me success. It would be great if they had ideas that they wanted to share with others as well. 

To register for Dr. Johnson's free webinar, "Four Years of Predictive Modeling and Lessons Learned", or for more information, please click here

To read a case study about how Dickinson College uses predictive modeling for strategic enrollment management, please click here

Friday, February 10, 2012

Creating Variables: Distance From Campus


Hi folks. This is the first entry in a new series I'll call "Creating Variables". This series will explain the creation and use of helpful predictive variables that might not be present in your existing datasets.

Today we’ll talk about how to create a “distance from” variable.  A "distance from" variable can also be applied to things like retail sales, fundraising or donor models, or even hospital admissions. This variable is particularly useful for predicting enrollment at admission, which is the example we'll use. Because we don’t have a lot of information about each candidate at admission, we have to use each piece of information we’re given to the best of our ability. In this case, we use the zip code of each applicant and the zip code of our institution to determine each applicant’s distance from campus. Distance from campus is often very predictive of an applicant’s likelihood to enroll at a particular institution – usually, the closer an applicant lives to the institution, the more likely they are to enroll there.  Let’s get started.


 To begin, we’ll need to hook the applicant data into a transform node:




Opening the transform node, we’ll need to select “Distance Between” from the formula drop-down menu:








In the “Enter a Formula” window, you’ll want to enter:


Where “A” is the variable in your dataset that represents each applicant’s zip code, and ‘03818’ is replaced by your institution’s zip code. Be sure to set the result type to “Integer” and name your new variable “Distance from Campus” before saving. If you preview your data, you'll see that each student now has a value in the "Distance from Campus" column, which will be located all the way on the right as you scroll through your admission variables. 

Tada! At this point, you’re ready to output your dataset, augmented with a shiny new variable, and one step closer to predicting enrollment! 

-Caitlin Garrett, Statistical Analyst at Rapid Insight