Tuesday, February 5, 2013

Facebook's Graph Search and Prospect Research


“Facebook’s mission is to make the world more open and connected. The main way we do this is by giving people the tools to map out their relationships with the people and things they care about. We call this map the graph. It’s big and constantly expanding with new people, content, and connections. There are already more than a billion people, more than 240 billion photos, and more than a trillion connections. Today we’re announcing a new way to navigate these connections and make them more useful.”  [Facebook]

Introducing Graph Search

Last week, Facebook unveiled their new Graph Search tool, which allows users to search for Facebook users by interests, likes, relationship status, and location, among other qualifiers. Examples of searches include “Friends who like yoga who live in Chicago”, “Pictures of friends taken before 1998”, or “Friends who like Make-A-Wish”. The results of these searches can reveal full names, addresses, employers, friends and family, and photographs. Creative searching can yield some very telling results, as evidenced by a popular Tumblr site’s investigation into search possibilities.  Currently, graph search is still in beta, and you can join the waiting list here.

Specifics on Graph Search Data

In truth, all of the data gathered by Graph Search has been available for quite some time.  But the lack of an all-encompassing search feature made this data fairly obscure and hard to collect - until now.  So far, researchers aren’t sure how users will react to their personal data being more easily mined. Many users are likely to get a bit freaked out by their inclusion in these “big net” searches.  They’ll respond by making their information more private using Facebook’s existing privacy settings.  Chances are that most will passively accept this feature as an acceptable part of living in an age of social connectivity.  A few may even begin sharing more information in an effort to provide and receive more of the purported benefits.

Potential users of Graph Search need to remember the caveats.  Facebook’s information can be incomplete, deceptive, and even fictitious (“ironic likes” for example).  Then there are the obvious limitations – users need to like pages to generate searchable connections.   But the breadth and depth of data Facebook offers can’t be found anywhere else.  Leveraging the interlacing interests of individuals, businesses, and organizations into some very powerful insights is simply too valuable to ignore.

Using Graph Search for Prospect Research

So what does Graph Search mean for prospect researchers?  It means effectively mining the 8+ years of data that Facebook has been collecting just got a whole lot easier.  There are several ways that I see it helping immediately.

The ease of collecting data makes it easier to patch holes in current constituent datasets. With a little creativity, leveraging the new search options may make more imputation of variables possible, particularly by examining constituent relationships and interests. For example, age can be imputed by graduation year, which will become searchable.  

There will be better opportunities for identifying new constituents based on searches. Possible search ideas include: friends of those who are already involved with the organization, people who live nearby, people whose interests coincide with your institution’s mission, or any combination of the above. Finding friends of users who like a page is a quick search, and aggregating this list to people who live nearby will become a piece of cake.

Your Institution’s Facebook Page

On the flip side, the interest and ability of others to find you through a Graph Search should not be overlooked.  Information about fundraising organizations is about to become a whole lot more visible. The number of channels by which organizations can be searched will also greatly increase, which can mean more traffic for your page. Here are a few steps to take in preparation for the widespread release:

  • Fill out the basic information section of your page, and include as many relevant keywords as needed. This includes selecting a category and sub-categories if you haven't already. 
  • Make sure your address is up to date. Because users can search by address, you'll want this information to be as accurate as possible. 
  • Got photos? Label them with descriptive text, tag the people in them, and add a location to them. Photos are fair game for searches, and the more information you can provide at a glance, the better. 
  • If you haven't already, update your page's URL to be customized, preferably containing the name of your organization. This will also improve your SEO on Google. 
  • Check your content. Gathering and retaining followers is more important than ever. Make sure to keep things relevant and interesting to keep people engaged. 
  • Once Graph Search becomes available to the whole Facebook community, try constructing searches that you would hope your page would appear in. If it doesn't, look to those whose pages did appear and imitate what they did to list so well - the sincerest form of flattery!
For more information on Graph Search, visit https://www.facebook.com/about/graphsearch .

For follow-up questions, or help working with your data for this purpose, contact Caitlin Garrett at caitlin.garrett@rapidinsightinc.com. Our next exploration will be on using Graph Search for Enrollment and Recruiting. Please feel free to comment if you have thoughts on additional ways to use the tool. 

Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, January 29, 2013

Four Years of Predictive Modeling and Lessons Learned

I recently got the chance to talk with Dr. Michael Johnson, Director of Institutional Research at Dickinson College, about his experiences with predictive modeling over the past four years. Dr. Johnson will be presenting a free webinar, "Four Years of Predictive Modeling and Lessons Learned", on Thursday, January 31st at 2pm EST in which he'll provide a more in-depth look at his experiences. 

Can you give us an example of a lesson you’ve learned through your experiences with predictive modeling?

I’ve learned that predictive modeling is good but predictive modeling in real time is just more extremely beneficial. When we picked up Rapid Insight, we moved a five day turnaround time to an eight minute turnaround. That’s one of the biggest changes we’ve made, and the effects have been very apparent.

If there was another thing I’ve learned, it is to automate absolutely every process possible to remove the opportunity for human error.

What types of predictive models will you be discussing during your webinar?

The enrollment management model is our primary model but a close cousin to that is the one we’ve been using for retention. The dataset is basically the same only slightly enhanced. It’s good to use essentially the same dataset to solve two different problems.

How has predictive modeling changed the way you operate?

It is the primary tool that we use to make decisions on the incoming class. This last week has been an incredible example of that. We’re taking a look at our early action pool and asking questions: What does it look like? How does it compare with previous years? What if we swap out some people; how does that change our incoming class?

What do you hope attendees will take away from your webinar?

There’s really no need to reinvent the wheel, so I’ll share some ideas that I’ve picked up. I hope that others come on board and share their successes as well. We all have the same problem set, so it will be nice for others to take away a few things that I’ve seen that have given me success. It would be great if they had ideas that they wanted to share with others as well. 

To register for Dr. Johnson's free webinar, "Four Years of Predictive Modeling and Lessons Learned", or for more information, please click here. 

To read a case study about how Dickinson College uses predictive modeling for strategic enrollment management, please click here. 

Tuesday, January 22, 2013

Dealing with Nulls in Veera Transform Formulas

This next post comes from Jeff Fleischer, our Director of Client Operations, support wiz, and analyst extraordinaire: 

Working out the logic of a new variable you want to create with a TRANSFORM node can be challenging. But when missing data ("nulls") get into the mix, it can be especially confusing and frustrating. For example, if you'd written the conditional formula...

               IF ([A]='Freshman', 'UG', 'Grad')

...and some of the fields under column [A] were null, you would get nulls as an output for those rows rather than the desired 'Grad'. This is because trying to equate something with "nothing" confuses Veera as to what you would really want as a result. So here are some suggestions on how best to deal with those gaps and still get to the outcome you need...

1. 
The most obvious way to deal with gaps in data is to replace them with something. This may not always be desirable, but when it is, using a CLEANSE ahead of your TRANSFORM is your best bet. Select the "Is Missing" operator and use Alt-Left Mouse to select all the columns that need their data fields filled in with that new value, like 'unknown'. 

Of course, you could instead place a CLEANSE after your TRANSFORM, using it to fill in any missing values appearing in the new column. 

2.
If filling in those data holes using a cleanse is not preferable, maybe just a temporary patch will do. Look for the "Treat Missings in Formula as Zeros" checkbox just above the "New Variable Name" field in the TRANSFORM. Just as the name suggests, this will temporarily replace any missing data with a zero, allowing most operations to function. Be careful, though, if the column you're evaluating already contains zeros - the output may not be what you intended!

3.
If even temporarily replacing nulls with something else isn't an option, then change your TRANSFORM formula to deal with them ahead of everything else. To do this, you'll likely need to use one of two built-in Veera functions - IS NULL or IS NOT NULL. We might change our example to include another condition, such as...

IF ([A] IS NULL, 'Withdrawn', 
IF ([A]='Freshman', 'UG', 'Grad'))

The idea here is to catch any nulls before they affect the rest of the logic by putting that condition first. 

4.
Finally, another (if more specialized) option might be to use the "Missings:" TRANSFORM feature. Unlike the "Treat Missings in Formula as Zeros" checkbox, this control changes nulls that appear as the final result of a formula. The replacement options offered by this feature are limited (0 or 1), but it may be an easy way to fix a problem with absent data appearing in a new numeric field. 

-Jeff Fleischer

Wednesday, January 16, 2013

How to Interpret a Decile Analysis


After building a predictive model, there are several ways to determine how well the model is describing your data. One visual way to get an idea of how well a model is fitting your data is by taking a look at the decile analysis. Here we’ll take a look at what the decile analysis represents, how it’s created, and how to spot a good model.

What a Decile Analysis Represents

After building a statistical model, a decile analysis is created to test the model’s ability to predict the intended outcome. Each column in the decile analysis chart represents a collection of records that have been scored using the model. The height of each column represents the average of those records’ actual behavior.

How the Decile Analysis is Calculated

1. The hold-out or validation sample is scored according to the model being tested.
2. The records are sorted by their predicted scores in descending order and divided into ten equal-sized bins or deciles. The top decile contains the 10% of the population most likely to respond and the bottom decile contains the 10% of the population least likely to respond, based on the model scores.
3. The deciles and their actual response rates are graphed on the x and y axes, respectively. 

After the decile analysis is built, you’ll want to take a look at the height of the bars in relation to one another. Deciding whether a model is worth moving forward with depends on the pattern you see when viewing the decile analysis. 


Ideal Situation: The Staircase Effect

When you’re looking at a decile analysis, you want to see a staircase effect; that is, you’ll want the bars to descend in order from left to right, as shown below. 

This is telling you that the model is “binning” your constituents correctly from most likely to respond to least likely to respond. A model exhibiting a good staircase decile analysis is one you can consider moving forward with.

Not-So-Ideal Situations

In contrast, if the bars seem to be out of order (as shown below), the decile analysis is telling you that the model is not doing a very good job of predicting actual responses.


 If the bars seem to be the same height, or the decile analysis looks “flat”, the decile analysis is telling you that the model isn’t performing any better than randomly binning people into deciles would. In both cases, your model should be improved before moving forward with it.  

-Caitlin Garrett, Statistical Analyst at Rapid Insight


Thursday, January 10, 2013

Valuing Analytics & Predictive Modeling in Higher Ed

As promised, here is part two of my interview with Mike Laracy, Founder, President, and CEO  of Rapid Insight. Mike's 20+ years of data analytics & predictive modeling experience have provided him with many insights. Here's Mike on becoming more data-driven in higher education, which models produce the highest ROI, and mistakes to avoid:

Where does predictive modeling fit into the analytic ecosystem in higher education?

Within the analytic ecosystem in higher ed, there is a range of ways in which data is analyzed and looked at. On one side, you have historical reporting, which our clients do a lot of and is vital to every institution.  Somewhere in the middle is data exploration and analysis, where you’re slicing and dicing data to understand it better or make more informed decisions based on what happened in the past.  On the other side of the spectrum is predictive modeling.  Modeling requires taking a look at all of the variables in a given set of information to make informed predictions about what will happen in the future. What is each applicant’s probability of enrolling or what is each student’s attrition likelihood?  What will the incoming class look like based on the current admit pool?  These are the types of questions that are being answered in higher ed with predictive analytics.  The resulting probabilities can also be used in the aggregate. For example, enrollment models allow you to predict overall enrollment, enrollment by gender, by program, or by any other factor.  The models are also used to project financial outlay based on the financial aid promised to admitted applicants and their individual enrollment probabilities.

Higher education has come a long way in the last five to ten years in its use of predictive analytics. The entire student life cycle is now being modeled starting with prospect and inquiry modeling all the way through to alumni donor modeling.   It used to be that any institutions that were doing this kind of modeling were relying on outside consulting companies.  Today most are doing their modeling in-house.  Colleges and universities view their data as a strategic asset and they are extracting value from their data with the same tools and methodologies as the Fortune 500 companies.

What kinds of resources are needed and what is the first step for an institution who wants to become more data-driven in their decision making?

It’s important to have somebody who knows the data. As long as a user has an understanding of their data, our software makes it very easy to analyze data and build predictive models very quickly. And our support team is available to answer any analytic questions. 

Gaining access to their data is the first step. We see a lot of institutions that have some reporting tools which don’t allow them to ask new questions of the data. So, they might have a set of 50 reports that they’re able to run over and over but anytime someone has a new question, without access to the raw data there’s no way to answer the question. 

It really helps if the institution is committed to a culture of data driven decision making.  Then all the various stakeholders are more focused on ensuring data access for those doing the predictive modeling.

What do you say to those who are on “the quest for perfect data”?  Is it okay to implement predictive analytics before you have that data warehouse or those perfectly cleansed datasets?

No institution is ever going to have perfect data, so you work with what you have. We suggest seeing what you have, finding any obvious problems in the data, and then fixing those problems the best you can. We’ve designed our solutions such that a data warehouse is not required but, even with a clean data warehouse, the data is never going to be perfect.   As long as you as you have an understanding of the data, you can move forward.  

In your experience, which models in higher education produce the highest ROI?
We have a customer, Paul Smith’s College that has quantified their retention modeling efforts. Using their model results, they put programs into place to help those students that were predicted to be high-risk of attrition. They credit the modeling with helping them identify which students to focus on, saving them $3m in net tuition revenue so far.

We have other clients that are using predictive modeling on the prospect side and they’re realizing significant savings on their recruiting efforts. So instead of mailing to 200,000 high school seniors, they’re mailing to 50,000, and realizing significant savings by not mailing and not calling those students who have pretty much zero probability of applying or enrolling.

Although not as easily quantifiable, enrollment modeling has a pretty big ROI.  Not only on determining which applicants are likely to enroll, but in predicting class size.  If an institution overshoots and enrolls too many applicants, they’ll have dorm, classroom, and other resource issues.  If enroll too little, they’ll have revenue issues.  So predicting class size and determining who and how many applicants to admit is extremely important.

What are some common mistakes you see when approaching predictive modeling for your higher ed customers?

One mistake that I often see is when information is thrown out as not useful to the models.  Zip code is a good example.  Zip code looks like a five digit numeric variable, but you wouldn’t want to use it as a numeric variable in a model.  In some cases it can be used categorically to help identify applicants’ origins, but its most useful purpose is to for calculating a distance from campus variable.  This is a variable that we see showing up as a predictor in many prospect/ inquiry models, enrollment models, alumni models, and even retention models.  Another example of a variable that is often overlooked is application date.  Application date often contains a ton of useful information if looked at correctly.  It can be used to calculate the number of days between when the application was sent and the application deadline.  This piece of information can tell you a lot about an applicant’s intentions.  A student who gets their application in the day before the deadline probably has very different intentions than a student who applies nine months before the deadline.  This variable ends up participating in many models. 

To get our customers up to speed on best practices in predictive modeling we’ve created resources like lists of recommended variables for specific models and guides on how to create useful new variables from existing data.