Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Tuesday, June 25, 2013

Data Scientists: The Next Generation

As I’m sure you all have noticed, the data business is booming right now. (Are you tired of the term “big data” yet?) The fact that 90% of the data in world today has been created in the last two years is a great example of the growth trajectory of data. All of this data provides new opportunities for discovery for those who are willing to analyze it. Enter the data scientist.

 “Data Scientist” isn’t even listed as a career by the US Government’s Bureau of Labor Statistics yet, but it’s already been named the sexiest job of the 21st century by Harvard Business Review. With a growth pattern similar to that of data itself, it’s safe to say that data scientists are going to be in high demand. Among other skills, being a practitioner of data science requires analytical thinking, mathematical/statistical ability, a knack for communicating results to non-data people, and creativity. This combination of business acumen and technical skill isn’t easy to come by, and new graduate programs with an emphasis on data science seem to be cropping up daily to fill the gaps. One article from the New York Times recently asserted that the United States will need to increase the number of graduates with data science skills by as much as 60% to keep up with demand.  So, when you’re looking for new data scientists, where do you turn? To a generation who’s grown up with data science all around them – through Netflix recommendations, Google search results, and even at the movie theater à la Moneyball.

I was recently asked to participate in a “Job Hop Day” for a local elementary school. The idea was to expose 4-6 graders to different jobs that are available in the Mount Washington Valley in NH. It was a good opportunity to spend a fund day with elementary school students while exposing them to world of data science (and the idea that people actually get paid for doing it!). In preparing for our session, I realized that as thrilling as an hour-long lecture on data science might be for some, 10-year-olds probably wouldn’t be so interested. After ruling out a product demo and a slideshow, my coworkers and I thought about other ways to engage them. We decided the best approach for them to learn about being a data scientist was to do it themselves (in the guise of a game). 

When creating the game, we thought about some of the skills we wanted to reinforce, which were things like data mining, basic math, and the ability to make predictions. From there, we got creative – we wanted to pick a subject that kids would be interested in, and since vampires are on the brink of cliché, we settled on werewolves. The game we came up with was a variation of a Family Feud board that involved an initial data-mining phase to glean the characteristics of a werewolf.

To start, I gave the kids ten descriptions of people on color-coded index cards, five of which were designated as “werewolves” and five of which were “non-werewolves”. (Coming up with the descriptions was a good exercise for us as well, we tried to make  sure the clues weren’t too obvious, and had to plan them so that some characteristics were more popular than others. An example: three of the werewolves were vacationing in London this summer, but all five of them played some kind of sport). Each data scientist had a whiteboard to write down their descriptions as they went, and we stopped the “data mining” portion of the game once they all felt like they had come up with as many characteristics as they could. The Family Feud board I mentioned earlier had the ten characteristics listed in order of the number of times they came up, and the kids took turns guessing what was on the board.

Over the course of the day, three groups of students played the game, and all three groups seemed to really enjoy it. After we finished the game, we talked about the different uses of data and predictive modeling, covering examples spanning test scores to baseball. They were knee-deep in baseball season and pretty excited when I told them about a baseball scout’s presentation I saw at DRIVE, and how they used statistics to predict what might happen in each game. It was evident from our conversations that the kids had some knowledge of the amount of data around them and were interested in examining the world from a data-driven viewpoint. (I should probably mention here that the kids who chose to attend our session knew it would be math-related, so our sample was a bit biased.) Most of them had never heard of a data scientist or a statistical analyst before, but they were interested in the type of thinking we’d done. A few days later, a student’s mom told me that her son “loved the game” and “was so excited that it was an actual job that he could shoot for”.

Overall, our ad hoc approach to the data scientist experience seemed to go over well, but there’s always room for improvement. I’m interested in any ideas or experiences you guys have might regarding young data scientists, and would love to hear about them in the comments below. In the meantime, if you’ve had a sneaking suspicion about a certain neighbor around a full moon, or just want to have a little fun, I’d recommend trying out your own version of the game. 




-Caitlin Garrett is a Statistical Analyst at Rapid Insight

Tuesday, April 30, 2013

Using Social Media Data

Every minute, millions of pieces of social media data are generated around the world. In any given minute*:

Instagram users share 3,600 photos
Brands and organizations on Facebook receive 34,722 “likes”
Twitter users send over 100,000 tweets

With millions of people sharing more information each day, we are witnessing a shift in how information is being produced online as it becomes more and more user-generated. The web is moving away from being a static library and becoming a more open, more connected place to share content. There are plenty of opportunities to mine new variables from the constantly increasing expanse of social media data. Here are a few:

  • For a college or non-profit organization’s advancement department, tapping into these variables can provide insight into how connected an individual is with the institution. If a constituent is following you (whether on Facebook, LinkedIn, Twitter, or another forum), they are actively choosing to be connected to you, which conceivably may make them more likely to donate.
  • In a college enrollment office, tracking which high school students have liked their Facebook page can give them an idea of students who are very interested in their institution. This information could be used to qualify prospects or to help shape the pool of students being marketed to.
  •  Brands looking to decide who to market to can analyze both the source and content of social media data to determine who might be most likely to buy a product or respond to a campaign. 

Tracking and leveraging these data points has the potential to add value to the predictive models you’re already building.  

So, are you already leveraging your social media data into your models, and if not, why not?

Tuesday, April 23, 2013

Five Steps for Data-Driven Strategic Enrollment Management


Establish your goals
This first step is crucial for mapping out a course to your end goal. Try to envision where you’d like to end up and formulate a specific goal to help get you there.

Possible goals include:
  • Reduce your prospect mailing budget
  • Increase accuracy of enrollment yield predictions
  • Meet diversity objectives
  • Increase your retention rate

Get to know your data
The first step to getting to know your data is gaining access to your data, which is trickier for some people than others. If you have to go through IT to access your data, it helps to have a clear goal in mind and a good idea of what fields or tables you’ll need.

Once you have your data, you’ll need some time to get well-acquainted. A good starting point is to make sure you understand what each field represents and how things are coded. If you have questions about how data is being recorded or stored, this is the time to ask. Once you have a handle on what your data represents, you’ll want to thoroughly review it.

A few suggestions:
  • Spot-check the accuracy of your data. Double-checking things like the mean, min, and max for each variable is a quick way to verify accuracy.  If you spot any data quality issues, do your best to resolve them sooner than later.
  • Check for missing values. If you have a variable with a high number of missings, you’ll need to decide whether or not to use that variable and if there’s a way to fill in what’s not there.
  • Brainstorm ideas for new variables. If you can’t create new variables from what you have on-hand, spend some time thinking about things that might be worth tracking going forward. 

Analyze your data
I realize that the word “analyze” represents a whole spectrum of techniques and applications – and that’s okay. In a general sense, you’ll want to see if fields in your dataset can give you some insight that you can relate back to your initial goal(s).

Some ideas:

  • Look at correlations within your dataset. Are they positive or negative? Large or small?
  • Look for the differences between your target and non-target population, variable by variable.
  • Visuals help! Graphs are a great way to get a feel for the relationships between your variables.
  • Try building a predictive model. The results you get will be more directly applicable to driving decisions. 
You may get some surprising results during the analysis phase. I’ve worked on projects where the end insight was the exact opposite of what was expected. Although sometimes the results can be surprising, it’s important to let your data tell its story. The other side of analysis is that your data can confirm what you’ve long-suspected to be the truth – whether it’s that students from Montana are more likely to enroll, or that the number of first term credits impacts a student’s likelihood of attrition – embrace these confirmations and continue to rely on that information. 

Turn analysis into insight
Keep your initial question in mind, take what you’ve learned from your analysis, and apply it going forward. The idea here is to replace outdated anecdotal evidence with insights from our data. If your goal was to save money on prospect marketing efforts, use the factors that correlate to a higher response rate to drive your decisions about who will receive the next round of direct mail. If you’re trying to improve retention rate, target those students who look most like previously dropped students and reach out to help keep them on campus.

Assess your decisions
Last, but certainly not least, don’t forget to circle back and re-assess your decisions. If you feel like you’re not making progress toward your initial goal, consider re-framing it or breaking it down into more manageable phases. If you feel good about the progress you’re making, start working on new goals. A data-driven decision should be sustainable under conditions similar to the past. Don’t be afraid to revisit past goals if you feel like you can improve or add something to your initial recommendation. 

...Did we miss anything? Have questions about becoming more data-driven? Leave them in the comments below.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, February 5, 2013

Facebook's Graph Search and Prospect Research


“Facebook’s mission is to make the world more open and connected. The main way we do this is by giving people the tools to map out their relationships with the people and things they care about. We call this map the graph. It’s big and constantly expanding with new people, content, and connections. There are already more than a billion people, more than 240 billion photos, and more than a trillion connections. Today we’re announcing a new way to navigate these connections and make them more useful.”  [Facebook]

Introducing Graph Search

Last week, Facebook unveiled their new Graph Search tool, which allows users to search for Facebook users by interests, likes, relationship status, and location, among other qualifiers. Examples of searches include “Friends who like yoga who live in Chicago”, “Pictures of friends taken before 1998”, or “Friends who like Make-A-Wish”. The results of these searches can reveal full names, addresses, employers, friends and family, and photographs. Creative searching can yield some very telling results, as evidenced by a popular Tumblr site’s investigation into search possibilities.  Currently, graph search is still in beta, and you can join the waiting list here.

Specifics on Graph Search Data

In truth, all of the data gathered by Graph Search has been available for quite some time.  But the lack of an all-encompassing search feature made this data fairly obscure and hard to collect - until now.  So far, researchers aren’t sure how users will react to their personal data being more easily mined. Many users are likely to get a bit freaked out by their inclusion in these “big net” searches.  They’ll respond by making their information more private using Facebook’s existing privacy settings.  Chances are that most will passively accept this feature as an acceptable part of living in an age of social connectivity.  A few may even begin sharing more information in an effort to provide and receive more of the purported benefits.

Potential users of Graph Search need to remember the caveats.  Facebook’s information can be incomplete, deceptive, and even fictitious (“ironic likes” for example).  Then there are the obvious limitations – users need to like pages to generate searchable connections.   But the breadth and depth of data Facebook offers can’t be found anywhere else.  Leveraging the interlacing interests of individuals, businesses, and organizations into some very powerful insights is simply too valuable to ignore.

Using Graph Search for Prospect Research

So what does Graph Search mean for prospect researchers?  It means effectively mining the 8+ years of data that Facebook has been collecting just got a whole lot easier.  There are several ways that I see it helping immediately.

The ease of collecting data makes it easier to patch holes in current constituent datasets. With a little creativity, leveraging the new search options may make more imputation of variables possible, particularly by examining constituent relationships and interests. For example, age can be imputed by graduation year, which will become searchable.  

There will be better opportunities for identifying new constituents based on searches. Possible search ideas include: friends of those who are already involved with the organization, people who live nearby, people whose interests coincide with your institution’s mission, or any combination of the above. Finding friends of users who like a page is a quick search, and aggregating this list to people who live nearby will become a piece of cake.

Your Institution’s Facebook Page

On the flip side, the interest and ability of others to find you through a Graph Search should not be overlooked.  Information about fundraising organizations is about to become a whole lot more visible. The number of channels by which organizations can be searched will also greatly increase, which can mean more traffic for your page. Here are a few steps to take in preparation for the widespread release:

  • Fill out the basic information section of your page, and include as many relevant keywords as needed. This includes selecting a category and sub-categories if you haven't already. 
  • Make sure your address is up to date. Because users can search by address, you'll want this information to be as accurate as possible. 
  • Got photos? Label them with descriptive text, tag the people in them, and add a location to them. Photos are fair game for searches, and the more information you can provide at a glance, the better. 
  • If you haven't already, update your page's URL to be customized, preferably containing the name of your organization. This will also improve your SEO on Google. 
  • Check your content. Gathering and retaining followers is more important than ever. Make sure to keep things relevant and interesting to keep people engaged. 
  • Once Graph Search becomes available to the whole Facebook community, try constructing searches that you would hope your page would appear in. If it doesn't, look to those whose pages did appear and imitate what they did to list so well - the sincerest form of flattery!
For more information on Graph Search, visit https://www.facebook.com/about/graphsearch .

For follow-up questions, or help working with your data for this purpose, contact Caitlin Garrett at caitlin.garrett@rapidinsightinc.com. Our next exploration will be on using Graph Search for Enrollment and Recruiting. Please feel free to comment if you have thoughts on additional ways to use the tool. 

Caitlin Garrett, Statistical Analyst at Rapid Insight

Tuesday, April 17, 2012

Creating Variables: Out-of-state Flag


Sometimes it’s good to see which of your students or donors are in-state because an in-state population may be more likely to enroll or be retained or give than an out-of-state population. Creating an out-of-state flag from a “state” variable allows you to easily differentiate between your in-state and out-of-state prospects. I should also note that it is just as easy to create an in-state flag if that better suits your data. In any case, here’s how:

The first step is to hook your data source to a transform node:

Because we’ll be creating a binary (“yes or no”) variable, we’ll want to click on the “if” button (at the top of the buttons on the right side), which will automatically generate an equation that we can change to suit our data. 

In the “Enter a Formula” window, we’ll want to edit the auto-generated equation so it reads:



Where ‘[A]’ is the variable in our dataset that represents state, and the term it is set equal to (in this case, ‘NH’) is the term in our dataset that represents our institution’s state. Note that we could have set state equal to ‘New Hampshire’ or a numerical code, as long as it matches the term that represents New Hampshire in our dataset. The equation outputs a variable that is equal to ‘1’ when state is NOT New Hampshire and ‘0’ otherwise, thus flagging records which are out-of-state.


The final step before naming and saving your out-of-state flag is to select “binary” from the “Result Type” list.








And, voila, it’s easy as that! You now have a quick way of identifying in-state vs. out-of-state students in your dataset; let the reporting begin!

PS: If you guys have any specific requests for a variable to be featured in the "Creating Variables" series, please leave them in the comments or email me directly!

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Wednesday, April 4, 2012

Creating Variables: Age


Hi all! Today I’d like to a cover a pretty universally predictive variable: age. Age can be created in relation to the date of a particular event (like an application date or a mailing date), or as a reflection of age today, at this moment. Either way, age is often predictive and easy to add to your dataset by creating it in Veera from a “birth date” field.

The first step in doing so is to hook your data source to a transform node: 
After opening the transform node, we’ll want to click on the function button and select the second “YearsBetween” function.

[Note: Veera is capable of outputting the number of years between two dates in two separate ways. The first function on the list calculates the number of years between two dates, regardless of the actual day and month, while the second function calculates the number of years between two dates taking day and month into account. To illustrate this point, take the dates December 1, 1960, and April 1, 1980. Using the first “YearsBetween” function, the number of years between these dates is 20. Using the second “Years Between” function, the number of years between these dates is 19. See the difference?]

Here, we have two options. We can (a) calculate age today or (b) calculate age at a specific point in time, depending on what we type in the “Enter a Formula” window.

(a) Age today:  






Where ‘[A]’ corresponds to the variable in your dataset that represents birthdate, and “TODAY()” is the Today function from the drop-down menu on the right. 



or


(b) Age at a specific point in time:





Where ‘[A]’ corresponds to the variable in your dataset the represents birthdate, and ‘00/00/0000’ represents the specific date on which you’d like to measure age. 

Be sure to save before exiting the transform node, and there you have it, a brand-new age variable!

PS: If you guys have any specific requests for a variable to be featured in the "creating variables" series, please leave them in the comments or email me directly!

-Caitlin Garrett, Statistical Analyst at Rapid Insight


Thursday, March 8, 2012

Creating Variables: Days Between Application Date and Term Start


Hello everybody! This post will be a continuation of the Creating Variables series. Today we’ll be discussing how and why to create a “days between application date and term start” variable.

At first glance, this variable seems a little long-winded, but I can assure you, it’s worth its weight in characters. As you all know, for any institution that accepts applications on a non-rolling basis, there exists a window of time during which applications must be filed to be considered for acceptance. The amount of time between when an application is submitted and when the relevant admission term begins can be an indication of a student’s interest in a particular institution. For example, a student may turn in an application to his first-choice college during the first week that applications are accepted, but this same student might wait until the day or week before the deadline to turn in applications to his safety or back-up schools. In this way, the amount of time between the day that a student turns in an application and the term start date can be seen as an indicator of that student’s interest. Let’s go ahead and calculate this:

The first step is to hook applicant data into a transform node:


Next, after opening the transform node, we’ll need to select the “Days Between” formula from the drop-down menu:






In the “Enter a Formula” window, we’ll want to enter:




…Where ‘[A]’ corresponds to the variable in your dataset that represents the date each application was submitted, ‘09/01/2012’ represents the start date for the term you’re admitting for, and “date” is actually the date function from the formula drop-down menu:








Before naming and saving this new variable, be sure to switch the “Result Type” to “integer”:





And, voila! Now you have a “days between application date and term start variable” to add to your predictive variable arsenal.

-Caitlin Garrett, Statistical Analyst at Rapid Insight

Friday, February 10, 2012

Creating Variables: Distance From Campus


Hi folks. This is the first entry in a new series I'll call "Creating Variables". This series will explain the creation and use of helpful predictive variables that might not be present in your existing datasets.

Today we’ll talk about how to create a “distance from” variable.  A "distance from" variable can also be applied to things like retail sales, fundraising or donor models, or even hospital admissions. This variable is particularly useful for predicting enrollment at admission, which is the example we'll use. Because we don’t have a lot of information about each candidate at admission, we have to use each piece of information we’re given to the best of our ability. In this case, we use the zip code of each applicant and the zip code of our institution to determine each applicant’s distance from campus. Distance from campus is often very predictive of an applicant’s likelihood to enroll at a particular institution – usually, the closer an applicant lives to the institution, the more likely they are to enroll there.  Let’s get started.


 To begin, we’ll need to hook the applicant data into a transform node:




Opening the transform node, we’ll need to select “Distance Between” from the formula drop-down menu:








In the “Enter a Formula” window, you’ll want to enter:


Where “A” is the variable in your dataset that represents each applicant’s zip code, and ‘03818’ is replaced by your institution’s zip code. Be sure to set the result type to “Integer” and name your new variable “Distance from Campus” before saving. If you preview your data, you'll see that each student now has a value in the "Distance from Campus" column, which will be located all the way on the right as you scroll through your admission variables. 

Tada! At this point, you’re ready to output your dataset, augmented with a shiny new variable, and one step closer to predicting enrollment! 

-Caitlin Garrett, Statistical Analyst at Rapid Insight