Sunday, February 8, 2026

Umamusume and Generative AI

After being roped in by my teenaged daughter, I've been playing a lot of Umamusume lately. It's a Japanese gacha game about "horse girls" (uma musume, or "horse daughter", in Japanese), young women with horse ears and tails in an alternate universe who inherit their names, personalities, and careers from (Japanese) racehorses in our world. The game obviously appeals to people whose preferences run toward women with animal features (kemonomimi in Japanese); that's not my cup of tea, but I won't judge. However, in what follows I am going to render some judgments about generative AI, specifically the sometimes ridiculous answers that Google's AI prepends to search results. But first, more on the game....

Rather than the girls, what does appeal to me about the game is the strategy, which involves a lot of probabilistic reasoning—right up my alley, given the love of statistics that led me to become a data scientist. Like with any gacha game, there's a "meta" strategy of deciding how to parcel out scarce (for players who don't "whale" by spending hundreds or thousands of dollars) resources to "pull" for the assets needed to play the game (horse girls and "support" cards in this case): because it's a gacha game, you never know exactly what you're going to get when you pull, but you play the odds, and work with the information you have about probability distributions, while prioritizing the most useful targets.

Unlike the average gachaUmamusume also features a deep tactical dimension of decision-making in the daily "grind" required to build up your team of horse girls. Each day, you guide at least one trainee through a "career", making training decisions and picking skills that, if you're lucky, will give her the stats and abilities she'll need for a place on your team. Sometimes you're working with explicit probabilities (like the probability that a skill will activate if its activation trigger occurs), and other times you're working with much vaguer contingencies (like the chance that the trigger condition will occur in the first place during an actual race), but in either case you're using probabilistic reasoning to make decisions, and yay, that's fun for a person like me (YMMV).

And though horse girls are not my cup of tea, I do find many of the characters quite appealing (shout-out here to King Halo, Nice Nature, and Narita Taishin), and I have to say that their background stories are well-written. That brings me to the Google search that inspired this post. To wit, I was watching the video story for an uma named Mejiro Dober when I encountered this:

 If you're like me, you're wondering, "What's this 'Bell' business?" It's not an obvious nickname for a racer, like "Tiger" or "Beast" or even "Twinkle Toes", but its origin isn't explained within the story—or, for that matter, in any of the official Umamusume lore. So naturally I turn to Google and ask, "Why is Mejiro Dober called Bell?" and I get something like what you see below:

Note that this is the best of three answers Google's AI produced for me at different times: the first time, right after the character came out on the global server (which is ~3 years behind the original Japanese server), the AI flat-out insisted to me that Mejiro Dober is not in fact called Bell, while the answer above (and another from a few hours earlier) at least hedge the response by acknowledging uncertainty, though both of the latter answers are still categorically wrong, in that the nickname does come from an official source. The AI also doesn't explicitly acknowledge that "Bell" does actually show up in search results (note the two circled hits below the AI summary), focusing instead only on the origin of the name.

Now you may be thinking, "Hey, Scott, you asked the wrong question: you asked why she was called 'Bell', not whether she was called 'Bell'." Yeah, that's true, because I already knew she was called that, and just wanted to know why. In all fairness, the "why" question is tougher, and in the first paragraph of the summary the AI quite rightly notes that it can't find an explanation. My problem is more with the second paragraph. But just for jollies, here's the AI's answer when I asked whether she's called Bell:

Still wrong. 

Presumably, the AI is using retrieval-augmented generation (RAG), meaning it's not just spitting out something retrieved from the model's own training (like what would happen if you asked ChatGPT), but rather doing a web search and then summarizing the results. Kudos to Google for following RAG best practice by linking to the top sources next to or underneath the AI summary, but in both screenshots the websites the AI links are actually the two commonly consulted fan-made wikis, not official sources (and the embedded video in the lefthand one is actually an ad on the wiki site for an interview with the cast of Stranger Things, so not at all helpful). The wikis are reasonably authoritative, and reproduce art and text from the official website, which might confuse a well-meaning AI model, but they aren't official.

The obvious reason that Google's AI doesn't link to the official sources is the Umamusume webpage for Dober contains all of one short paragraph of information, and beyond that, to get official info, you'd have to look at the game's social media accounts, the background story videos (helpfully posted on YouTube), and the anime associated with the game (also on YouTube, but not entirely canonical, as the anime characters are often somewhat different from their game counterparts). I'm not sure how if it all Google's AI consumes these sources, but it's entirely doable with modern multimodal large language models (LLM's) to extract the text from videos (they're subtitled!) and index it along with everything else Google stores from the web; and yes, you can do the same with actual voice, which Google already does by generating subtitles with AI. Given the amount of information currently stored only in video format, you'd think the company would be doing that already. If the Google AI isn't looking at these more exotic sources (and maybe it's not, because "Bell" appears several times in the background story videos), then it's deceptive to state it couldn't find anything in official sources, since it's not actually looking at most of them in the first place.

As you might have guessed, the Dober/Bell mess is not the only questionable result about Umamusume that I've seen from Google's AI. Once, it even gave me advice on when to use skills during a race (not useful, because skill activation is automated and rule-based, not under the player's control), and I wish I'd screenshotted that one. That instance was uniquely bad, but here's a more typical response. I asked what the best skills are for Late Surgers (one of four racing styles, each defined by where they run relative to the rest of the pack for the majority of a race): 

A few of these recommendations are good, and a few are highly questionable (though they probably represent something some ill-informed player posted somewhere), but I'd like to focus on the ones that are flat-out wrong.

Most obviously, the skills circled in red can't actually be used by Late Surgers. Speed Star works only for Pace Chasers (another style, which runs closer to the front). "Seuin Sky (Reeling in the Big One)" isn't even a skill, but rather the original variant ("outfit") for the uma Seuin Sky. The outfit's unique skill, Angling and Scheming, can be inherited by another uma, but, while it technically can work on any runner, to trigger it you have to be ahead on a corner late in the race, and so it's usually used only by Front Runners.

The errors circled in blue are more subtle. Let's start with Uma Stan and Ramp Up: they're actually completely different skills, with entirely different trigger conditions, but neither of one of them is likely to trigger in the late race (Ramp Up must trigger mid-race, and Uma Stan, because it can trigger any time a runner is close to 3 other runners, tends to trigger in the early race). Furious Feat and Position Pilfer seem to be presented as if they're different versions of the same skill (by way of comparison, above that line, you'll see On Your Left!, which is the premium, or "gold", version of Slick Surge, and the same is true of Rising Dragon and Outer Swell), but Position Pilfer is actually the non-premium ("white") version of Fast & Furious, which sounds a lot like Furious Feat, but isn't the same thing at all. Notably, while Position Pilfer and Fast & Furious are restricted to Late Surgers, Furious Feat, though readily usable by Late Surgers (it works on anyone in the back half of the pack) is restricted to Mile-distance races. The conflation of these skills explains the weird reference to "Mile and other distances".

In short, the advice provided by the AI here is effectively useless: there are some sound suggestions, but you need to know Umamusume pretty well to pick the wheat from the chaff, and anyone who knew the game that well wouldn't be asking this question in the first place (or would be asking for far more detailed answers on each skill, considering pros and cons).

It's true that generative AI can excel at so-called "zero-shot" tasks, constructing new things (like a list of good Late Surger skills) by assembling information using well-established rules and relationships. But performing this kind of "transfer-learning" task successfully requires that the model discriminate between what things it can transfer from one domain to another, and what things it can't. That works in domains like politics and economics and even real-life horse-racing, where as a model is trained it can extract those rules and relationships from billions of words of text. However, it tends to fall apart in a highly specialized domain about which people have written comparatively little, especially if it's a general-purpose model (like Google's AI), in which case it might try to transfer rules and relationships it really shouldn't. This is how we get the answer I didn't think to screenshot, treating Umamusume as if it were a game that allows players to make decisions during a race (which is how most racing games work). And it's how we get a response like the one in the screenshot above, where the AI can't figure out the rules well enough to plug the nuggets of info it's pulled from the web into the right places.

The common thread between both this and the Bell problem is that the AI model just doesn't have enough information on Umamusume to work with. For a human, there's more than enough information available to figure out the game, but training a generative AI uses a brute-force approach that requires lots and lots and lots of info to learn patterns that humans can pick up with a few minutes of light reading.

Oh, in case you're still wondering why Mejiro Dober is called "Bell", I did finally figure that out. I had a hunch it was one of those things that make more sense in the original language than in translation, so I consulted Google Translate, and sure enough, turns out the Japanese word for bell is "beru", while Dober's name in Japanese is actually (ignoring the subtleties of proper transliteration) "Mejiro Doberu", because "Doberu" is short for "Doberuman", the Japanese version of "Doberman"—all the Mejiro Farm foals that year were named after breeds of dog. The pun would be so obvious to a Japanese-speaker that it likely would rarely be commented on, making it hard for a RAG AI to find references to it even if the AI were pulling from Japanese as well as English sources. An AI language model might be able to recognize and reproduce word play, but figuring out that an English nickname derives from word play in another language appears to be beyond this AI's capabilities.

 

Monday, July 10, 2017

The Relationship between Machine Learning and Statistics

UPDATE from original 7/10/17 version to 9/14/17 version: I erroneously maligned confidence intervals for models of big datasets, conflating them with statistical significance; I've fixed that mistake below.


Like anyone who practices data science, I often get asked, by relatives and acquaintances, what "data science" is. Like any such question, it's not very hard to answer this one to the satisfaction of someone who knows little about the topic: in my case, I tend to describe the discipline as applying the principles of traditional statistics to large amounts of data, and I throw in mentions of the importance of writing code and manipulating databases. Nowadays, you can also mention machine learning, and many people will have at least a vague idea of what you're talking about.

However, even if the answer satisfies most listeners, it bothers me—because I've always wondered exactly where "traditional statistics" ends and "machine learning" begins. Defining that boundary turns out to be surprisingly difficult, but also pretty useful: it's one of those cases where the journey is more important than the destination. It's not really important exactly where we draw that line, but thinking about how machine learning differs from traditional statistics leads to further questions about whether (or rather, when) we can apply the accumulated wisdom of decades of statistical practice and quantitative research to the newer domain, and the answers to those further questions prove to be quite valuable.

The Easy Answer

The easy answer to the question is that machine learning is what statistics becomes when there's too much data to manipulate using traditional statistical algorithms. The most obvious illustration of this transformation is linear regression, where gradient descent replaces direct solution. The transformation has other implications as well: while direct solution (and other traditional algorithms) have long been built into statistical packages like SPSS, until recently, someone who wanted to use gradient descent would have to know at least enough code to install the right package and call the right function. (Mind you, we're seeing more and more machine learning algorithms packaged into easy-to-ease GUI's nowadays, which will leave the coding for those who want to tweak algorithms, create new ones, or build them into applications--much as most users of SPSS never learn scripting, but experts in statistical methods can use it to create powerful extensions to the original package.) Likewise, storing and processing large datasets lends itself to database applications, which can serve up all that data much more efficiently than the traditional method of reading in a CSV.

But the easy answer, while coherent, isn't entirely right. Data scientists actually use a number of techniques that we think of as "machine learning" even when the amounts of data involved are relatively small—indeed, no bigger than what a quantitative researcher in the 1980's would have dealt with. In 2015, my first year as someone with the job title "data scientist", my team worked on a number of demonstration projects for recommender systems. Because we hadn't deployed those systems yet, we usually didn't have real user data, let alone petabytes of it, and even where we did have all of the real data, it wasn't necessarily very big: for example, we used natural language processing (NLP) to measure the similarity between different pages on a website, and that website only had about 1200 pages, each of which included actual content of about two paragraphs. Nonetheless, we never doubted that our applications of collaborative filtering and NLP were "machine learning".

Why? Well, I've never been entirely sure, but I think the answer is that machine learning includes all of those algorithms whose development was prompted by increasing amounts of data and increasing amounts of computing power. The Doc2Vec we used to analyze those 1200 web pages could probably have run on my TRS-80 Color Computer back in the 1980's (it might have been an all-night job), but no one had invented it yet. The same applies to collaborative filters and any number of other recently developed methods that produce useful results even with smallish datasets. All of these algorithms get labeled "machine learning" because they were invented by people who did "machine learning", and, just like the methods used on truly big data, they're usually applied through code rather a traditional statistical package.

However, that's a pretty messy answer, and it really begs the question of the extent to which the difference between traditional statistics and machine learning is a matter of style (or, to put it more nicely, work methods and habits of thought) than of substance.

Interesting Discussion, Scott, But Why Does That Matter?

Yes, there's a point to all this. To wit, the important thing to understand here is that, because there's no bright line between traditional statistics and machine learning, the laws of statistics weren't abolished the first time someone programmed a gradient descent algorithm onto a computer. To me, as a former quantitative researcher in the social sciences, that point has always been blindingly obvious—but in all the machine learning classes I've taken over the years, I've seen only occasional mentions of the relationship between older and newer methods, and I've almost never seen a discussion of the implications of the laws of statistics for machine learning. I've always been struck by this, because really, it's pretty easy to figure out some of those implications.

For example, when your data really is big, you don't have to worry about certain things: the variance due to random sampling is infinitesimal, which means that any differences you find are statistically significant (i.e., if your sample is unbiased, etc., you can be sure the differences are real, though that doesn't in itself imply that they're meaningful). But, as I pointed out above, the data handled by machine learning algorithms isn't always big, and how many data scientists bother to think about exactly how big a dataset has to get before you can stop thinking about significance tests? Confidence intervals present a somewhat more complex problem: with enough data to eliminate error due to random sampling, confidence intervals will be smaller, but when you've got randomness in the model (that is, your model doesn't account for 100% of the variance in outcomes), you still need confidence intervals (or something equivalent) to express the variability of possible outcomes. I've met data scientists who worry about these problems, but not many of them. Heck, for some of the new techniques, like neural nets, I'm not even sure how you'd go about computing a confidence interval. Feel free to Google it: yes, it can be done, but it's not something that even crosses the mind of the average data scientist, and I've never seen the topic so much as mentioned in a machine learning class I've taken.

The implication of statistics that causes me personally the most grief is regularization: regularization is really, really useful because it allows us to solve a linear regression equation even when the number of independent variables (er, sorry, "features") is greater than the number of cases—for someone trained in traditional statistics, it's nothing short of glorious magic, allowing you to do what should be impossible. So why my grief? Well, there are often cases (remember, data today can get very, very big) when the number of lines of data far exceeds the number of features in the model.

Having put much thought into the problem, I cannot figure out a very good reason why you actually need regularization in such a case, and I can see some real downsides to it: it requires more processing, and it will likely produce a less accurate result. And yet, in all of the machine learning classes I've taken, I've never seen a discussion of this issue, and I rarely see a machine learning package whose functions allow the programmer to decide not to use regularization—you can accomplish the same effect by putting in a tiny number (yes, the model still converges without any meaningful regularization, provided you have enough degrees of freedom), but of course, in doing so you can't get the computational advantages of leaving out regularization completely. There's an analogous argument for validation to avoid overfitting: if your dataset is huge, and your training sample is randomly selected, you really shouldn't have overfitting.

I may be utterly wrong on both of these points, but the larger concern is that none of the classes I've taken on machine learning has even raised these issues. The silence is so deafening that, in executing the coding exercises that are often required for job applications, I've submitted regularized models when I knew (or at least suspected) that regularization was pointless (I did, though, note that in my response, and in one case, I submitted an unregularized model alongside the regularized one--I sometimes wonder if that might have kept me from getting the job.) Even if I'm wrong, and the people teaching classes and coding machine learning packages have thought carefully about whether regularization and validation are actually needed in all cases, it would be useful to learn about the reasons for their decisions; after all, there are always situations in which a given method doesn't apply very well, and if you don't understand the assumptions behind a method, you won't be able to identify those situations.

And don't even get me started about the importance of training in statistical research for distinguishing causation from spurious correlationn, as well as avoiding a variety of other analytical pitfalls.

So...when do we start giving every aspiring data scientist real training in statistics?

Tuesday, August 4, 2015

Online Course Review: Coursera and Stanford University's Mining Massive Datasets

It's been quite a while since I last posted—and eight months since I finished the class I'm reviewing. As it happens, I finally got a job as a data scientist in late January, and work has kept me busy. That job will be the subject of my next post, but right now, we're talking about Mining Massive Datasets, offered by Coursera and three professors from Stanford University, Jure Leskovec, Anand Rajaraman, and Jeff Ullman.

The seven-week course covers the same ground as the trio's book Mining of Massive Datasets, though in much less detail. (That link, incidentally, is to the e-book; if you really want the hardcover, you're welcome to follow this sponsored link to Amazon.) It's now been offered twice on Coursera, with a third iteration set to start on September 15th. Oddly, I never intended to take this class: a friend of mine was interested in it, and I signed up so that we could take it together, but initially I had passed on it. I had taken a look at the schedule, seen a few topics that I had covered before in other courses, and decided that I wouldn't get a lot out of it.

What I missed, in my hasty scan of the course description, was that Mining Massive Datasets is not the typical data science course that shows students how to put useful algorithms into practice through code. Instead, this course is about the algorithms themselves: how they work, why they work at scale, and how they've been modified to improve performance or cover different situations. There is a lot of math: this is not a course for someone without a solid background in calculus and linear algebra (it's not like you need to remember how to integrate esoteric functions—but you do need to understand the basics). There are not, on the other hand, any programs to write: many of the exercises absolutely can't be done without writing short scripts or using a statistical language from the command line, but the professors don't require any specific language, and the code isn't graded, only the answers. The point is not to create functioning implementations of the algorithms in question, but rather to understand the nuts and bolts of how they work. The exercises are challenging, and sometimes require consulting the e-book, especially when the complexity of a topic makes it hard to grasp in a short lecture.

Although the bulk of the course is devoted to algorithms, Week 1 provides an excellent description of how HDFS and MapReduce work—without ever giving details on Hadoop or the various languages used to write mappers and reducers. In fact, I came away from Mining Massive Datasets with a far better conceptual grasp of distributed file systems than I got from the Udacity course devoted entirely to the subject. (See my review of that course here.)

It should be said that Jeff Ullman is an excruciatingly monotonic lecturer, who sounds like he's reading everything directly from notes—and you likely wouldn't be at all surprised if he suddenly called out, "Bueller...Bueller...." In addition, his explanations are not as clear as those of Leskovec and Rajaraman. In fairness, though, Ullman tends to cover the most complex topics in the course, and I was always able to figure things out by consulting the book (which covers the material in greater depth, anyway).

In summary, I strongly recommend Mining Massive Datasets for anyone who wants to understand the nitty-gritty of algorithm design for big data. The course is not, however, for the faint of heart. You could make a very successful career in data science without ever taking or it taking anything like it—but taking it will certainly make you better at the profession.

Saturday, September 13, 2014

The Ethical Challenge of "Passive Predation" in Data Science: Can Data Science Provide the Solution, and Not Just the Problem?

I recently ran across an intriguing blog post from Michael Malek, on "Predatory Data Science". Malek notes that data science methods, especially "black box" machine learning, can unintentionally create what he calls "passive predation"—that is, taking advantage of some vulnerable group despite having no intention to do so. He uses the example of a machine learning model, created for a gun manufacturer, that ends up targeting marketing efforts at the suicidal, by identifying keywords associated with depression. The data scientist using the tool in question wouldn't have intended that result, and probably would never even be aware of it, because the group of suicidal depressives would be buried amidst thousands of other micro-segments identified by the same application.

Malek perhaps overdraws his point in the middle part of the post—a historical account of the dehumanizing effects of technology that's reminiscent of Marx's condemnation of working for money in "The Alienation of Labor"—but his main argument is quite sound, and not a little scary.

I wonder, though, if data science itself could provide a solution to this problem. I hereby announce a very unofficial contest, with prizes that will prove trivial at best (I might take a winner out to lunch, or talk about his or her idea at a Data Comunity DC meetup). Pretty much any method of accomplishing this goal, technical or non-technical, is fair game. Any takers?

Thursday, September 11, 2014

Online Course Review: Udacity's Intro to Hadoop and MapReduce

For my first course on Udacity, I decided to take Intro to Hadoop and MapReduce, a course created in conjunction with Cloudera, a company whose business model is based on the open-source Apache Hadoop. To sum up my asseessment, the course was useful, but could have been done much better.

The four-lesson course (short by Udacity standards) is supposed to take about a month to complete—like all Udacity course, and unlike those of Coursera, this is not a true MOOC, taken alongside other students in real time, but rather an interactive tutorial. However, Udacity's model does feature student discussion forums; customers who pay (at the rate of $150/month) also get help from live coaches, feedback on their final projects, and the opportunity to earn a "verified certificate", similar to Coursera's Signature Track, with the difference that Udacity, unlike Coursera, no longer offers certificates for non-paying students. (As I've mentioned before, a verified certificate and two dollars may buy you a cup of coffee, but I wouldn't count on its having any greater worth.)

Before I delve into the specifics of this course, let me say that I'm not a real fan of the Udacity interface. While both providers break each lesson up into a series of short videos, Coursera labels each of those videos with a topic, making it relatively easy to go back and find the material you need; by contrast, Udacity strings all the videos for a particular lesson together under a single heading, and so you have to hunt through all of them to find something (you can click on individual videos, and each one has its own label, but you have to click on or hover over a video to see the label). In addition, whenever the video stops for a quiz, it drops out of fullscreen (assuming you're in fullscreen, of course). Moreover, Udacity's discussion forum (note the singular there) has no organization whatsoever, aside from keyword tags—making a search for specific information rather laborious.

Thr first three lessons of this particular course, which features two instructors from Cloudera, are structured in a manner that the director of a music video would appreciate: many of the videos are very short, and switch jarringly from one instructor to the other. Nonetheless, the instructors are engaging, and there's a nice interview with Doug Cutting about how he helped to create Hadoop, and named it after his toddler son's stuffed elephant. The first two lessons, which explain the basics of how Hadoop and HDFS work, can best be described as "lite"—unchallenging nearly to the point of tedium.

Lesson 3 marks an abrupt change: this is where the programming exercises began. The class requires previous experience with Python, which I lacked, and so the exercises took more time for me than they should have, but I managed. One student in the forum questioned whether this was a course on Hadoop, or a course on Python regular expressions, but doing the exercises helped me learn some Python, and, much as I hate the language, it does have a very powerful vocabulary of regular expressions. Unfortunately, the instructor blew by the concept of Hadoop streaming so fast (in Lesson 2) that I wasn't entirely sure for a while what exactly I was doing, though I was managing to get it to work—and once I looked up Hadoop streaming on my own (it is, for the record, an API that allows Hadoop mappers and reducers to be written to be written in any language), I realized that the interface would work just as well for R.

Although the simpler exercises use an online Python compiler, for the exercises that require large datasets, the course's creators deserve kudos for having students install a virtual UNIX box on which a virtual two-machine Hadoop cluster has already been set up, and then manipulate data and write code in this realistic environment. Unfortunately, the exercises that require this virtual machine seem half-baked.

First off, the instructors haven't actually detailed how to write and execute Python scripts on the UNIX machine (the class discussion forum was very helpful here). Second, the syntax needed to make the scripts work is different from the syntax presented in the video lectures (though, fortunately, there are working sample scripts saved on the virtual machine). Third, and most seriously, one particularly tricky exercise requires knowledge that students could not possibly get from the instructions, or, in all probability, the data itself, but could only get from the hints that emerge from a trial-and-error process of submitting answers to the automated grader—it was an interesting little mystery to solve, but there are no automated graders in real life, and so I'm not sure what I gained from the effort.

Yes, figuring out ambiguous instructions does have some pedagogical value, and in the end, completing the exercises was very satisfying, but, especially in the case of the problem that was insoluble without the automated grader, I got the feeling that the difficultes I faced were the result, not of a pedagological choice, but of a simple lack of effort on the part of the instructors—and I felt like I had wasted part of my time.

According to posts in the forum, Lesson 4 was not part of the original class, though I'm not sure if it was planned all along, or tacked on later. To paraphrase Monty Python and the Holy Grail, the course was completed in an entirely different style at great expense and at the last minute. The lectures feature a different intstructor, a Udacity employee, in place of the Cloudera instructors. This lesson covers design patterns, specifically filtering patterns (more regular expressions), summarization patterns (minimums, maximums, and means, for example), and structural patterns (combining data sets); one lecture also deals with combiners, scripts inserted between mappers and reducers to make things more efficient by doing some of the reduction locally on each machine in the cluster.

I found these lectures better than the previous ones, and the exercises better prepared. I will say, though, that I eventually got bored with writing new and different regular expressions in Python, and didn't finish the last few exercises (or the final project, which isn't graded for non-paying students in any case), though I did watch all of the lectures.

In the end, this half-baked pastiche of a course at least gave me a decent idea of how Hadoop works, and removed the mystique of manipulating data stored on a Hadoop cluster. I wouldn't know how to set up a cluster myself (that wasn't the intent of the class, though I don't think it would be all that hard to do), but I do know how to use Hadoop streaming—and I've realized it's not exactly rocket science.

Monday, August 4, 2014

Online Course Review: Exploratory Data Analysis, from Coursera's Data Science Specialization

Back in May, I reviewed two of the short courses that make up Coursera's Data Science specialization. Although the four-week format greatly limits the content of any one course, I was generally impressed by the scientific approach of the specialization (something all too often lacking in data "science"), and, in the case of Getting and Cleaning Data, by the many pointers provided to R packages and sources of information for further study: the course may not have gone into a lot of depth, but it provided a good overview of what you can do with R.

I recently completed a third course in the specialization, Exploratory Data Analysis, taught by Roger D. Peng (the previous courses I took were taught by Jeff Leek). While I enjoy Peng's lecture style (unlike Leek, he engages the audience by showing his face at the beginnings of lectures), and I learned a lot, the course suffers greatly from the short format.

I initially overlooked this class: from the name (more on this in a minute), I never would have guessed that 3/4 of the lectures would cover graphics in R. Peng teaches the basics of the language's three major graphics packages, the base graphics, lattice, and ggplot2. As is the case for Getting and Cleaning Data, the lectures manage only to skim the surface, particularly for ggplot2, but they do give the student a decent idea of what's possible in R. I do though think that for ggplot2 Peng could do a better job of outling the advanced features than simply pointing students to the book written by the package's author, Hadley Wickham (thankfully, it's possible to find free PDF's of the book online, but I'm not sure it's where I'd want to start for solving a discrete problem, rather than studying ggplot2 in a methodical way).

So what's with the name of the course? Peng presents visualization in R as a way of conducting initial exploration of data, but it's obviously useful for more than that, since R can create decent visualizations of the results of analysis. I suspect that the course name was chosen so that one week of lectures on clustering and dimensionality reduction could be shoehorned into the syllabus. This material probably belongs instead in the Pratical Machine Learning course, but something had to be cut to limit that course to four weeks (cf. the nine-week Machine Learning, also offered by Coursera, and which I've reviewed previouslytwice, actually). The fact that clustering and dimensionality reduction can be used for exploratory analysis and visualization is the only thing that ties the entire course together.

What's particular disturbing is the way that all of this combines with the specialization's unique approach to exercises and evaluation. Each course includes a hands-on project, and, because open-ended projects in a MOOC must, for logistical reasons, be graded using a peer grading system, the final project for Exploratory Data Analysis only ends up covering material from the first two weeks of the course, since students need the third week to work on the project, followed by the fourth week to grade it; therefore, half the content of the class doesn't play any role in the project. On top of this—I suppose to avoid overloading students—there's no quiz, homework, or any other form of practice or evaluation covering the material on clustering and dimensionality reduction, which makes it hard for a student to know if he or she really understands those topics.

To sum up, I did find the information on data visualization in R useful, but I would have appreciated a full four weeks on the subject. The coverage of clustering and dimensionality reduction was out of place in the course; nonetheless, many will find it valuable (I had already seen most if not all of it in Machine Learning and another Coursera course, Social Network Analysis, which I've also reviewed).

I do have one more comment, though this applies to the Data Science specialization in general, and to Coursera, rather than solely to this course. Normally, after completing a Coursera course, a student can go back and look at the course archives at any later time; I've found this valuable when I suddenly find myself needing to refresh my memory or find out where I can learn more about a topic. Coursera has apparently disabled this feature for the Data Science courses: their archives are no longer accessible after the grading period is over (about a week after the finish of a course). I say "apparently" because, when I contacted Coursera a few months ago to ask why I could no longer access the archives of Getting and Cleaning Data, I never got a response—this is becoming something of a theme with Coursera, which, as I noted in my second review of Machine Learning, ignores most bug reports for that class. I suppose that paying customers might get better service, but I'm not going to pay just to find out if that's true.

Of course, you can always sign up for the current iteration of a class, since they're offered continuously, but it's annoying to have to do that each month. Fortunately, all of the class materials are also available in a GitHub repository, but it's not as easy to display documents on GitHub as in Coursera's web interface. For a set of courses that only skim the surface, and whose major value is in providing links to deeper information, this is a major failing.

Programming Languages for Big Data, Part 3

And now, one more word on the subject of R's speed. At my prodding, Tommy Jones contacted the authors of the study on programming language speed, and a productive discussion ensued. It turns out that the task in question was one that couldn't be vectorized, which means that R's main strength couldn't be applied in this case. However, it was possible to speed it up by writing C++ functions in R using Rcpp. The authors tried this, and revised their paper, reporting that, using Rcpp, R performed the task only 4-5 times slower than C++. For details, see Tommy's blog post, and the revised paper.

Friday, July 11, 2014

Programming Languages for Big Data, Part 2

I mentioned the recent study on the relative speeds of programming languages to Tommy Jones, a specialist in natural language processing and fellow member of the Data Community DC, and he, being more industrious than I, dove into the code used by the authors of the paper in question. In their R code, he found gems such as a triple-nested "for" loop inside a "while" loop (instead of the much faster "apply" functions), which made the comparisons pretty useless, at least in the case of R. See Tommy's blog, Biased Estimates, for more details.

Nonetheless, it's a pretty interesting question, and I'd love to see someone who's proficient in all of the languages involved try this test again, using better code. I'm still intrigued by the very high speed of MATLAB/Octave—something that leads Andrew Ng to recommend those languages over R for prototyping—though Tommy pointed out to me that, since R is closer to being a full-featured language, it's more flexible than the former languages.

Sunday, July 6, 2014

Programming Languages for Big Data

I'm a big fan of R: it just seems intuitive to me, and there's a package available for practically any type of analysis you might want to do. I have some experience with Octave (essentially the open-source version of proprietary MATLAB) and Python (which I detest, especially with its confusing statements-masquarading-as-function-calls), but I find R easiest to work with.

Therefore, I find a new study comparing the speeds of various languages for a statistical problem pretty depressing. When looking at this kind of a study, it's important to keep one big thing in mind: the authors tested the various languages on only a single task (albeit a common task, at least in economic modeling), and different languages will have different strengths and weaknesses at different tasks.

Nonetheless, the differences in run time are so large that it's not unreasonable to draw some conclusions. Even when compiled, R takes 240 to 340 times as long to run as C++. How about Python and MATLAB? Python with the default CPython compiler is nearly as slow as compiled R (155 to 269 times), but with Pypy it reaches 1/44 of the speed of C++. MATLAB takes only about 10 times as long to run as C++, or only about 50% longer when using Mex files (C, C++, or Fortran subroutines called by MATLAB). (Octave has Oct files written in C++, which serve a similar purpose; Octave can use Mex files, but not as well as MATLAB. See the GNU documentation on the subject for details.)

Wow. The botton line is that R might not be the best choice for time-consuming applications—in other words, those that have to crunch through a lot of data, especially if the calcuations involved are complex. I had read that it's slower than the alternatives, but I had no idea that the differences were so dramatic. I really should polish my Octave skills, and, judging by many of the job ads I see, knowing some C++ would not only open up possibilities for faster-running code, but would also make me more employable.

Thursday, May 29, 2014

Online Course Review: Coursera's Machine Learning, Part 2

Back in October, I reviewed Coursera's Machine Learning course, taught by Stanford professor and Coursera co-founder Andrew Ng. As I mentioned when I first reviewed the course, I wasn't able to finish it, because I was starting a new job and moving halfway across the country. I've just been able to complete the most recent iteration of Machine Learning (which was nearly identical to earlier versions), and I'd like to add a few more thoughts to my original review.

My last time through the course (the session that began on April 22nd, 2013), I completed almost all of the lessons on supervised machine learning methods, such as regression, logistic regression, and neural networks. In this session (which began on March 3rd, 2014), I repeated those lessons, and also finished the rest, most of which covered unsupervised techniques, such as clustering and recommender systems. I won't repeat the contents of my earlier review, except to note that what I said then remains true: Andrew Ng is a clear and charismatic lecturer, he covers advanced techniques, and he provides a number of practical tips, but the programming exercises are a bit canned, and may not fully prepare students to write their own scripts in Octave.

My new comments mostly reflect comparisons to other MOOC's, particularly the two courses from Coursera's Data Science specialization that I took recently. First of all, I think that Machine Learning could do more with the online format. In fact, most MOOC's consist largely of video-recorded lectures, with the addition of a sprinkling of interactive content, but Machine Learning falls short even by comparison with other online courses. The class does feature a very effective automatic grader, but it lacks any links to additional resources, or, very importantly, notes or slides from the lectures. While the latter omission may seem trivial (I didn't notice it the first time I took the course), a lack of lecture notes makes it difficult to go back later and review material from a lecture, except by watching the whole thing again. It's true that the programming exercises include detailed instructions, but not all of the course's topics are covered by these exercises, and at any rate the organization of the instructions can make it difficult to locate information on a specific subject.

I might also amplify my comment from the earlier review that the programming exercises involve mostly copying and pasting, rather than writing entire scripts. There's a reason for this: the focus of the course is on algorithms, not on other parts of solving machine learning problems. Nonetheless, my experiences taking other courses, especially those from the Data Science specialization, have demonstrated the practical value of forcing students to think about the nuts and bolts of a research project. Machine Learning's lack of a big final project also arguably deprives students of valuable practical experience, especially since these projects usually require students to explore the course material in greater depth than do short exercises; on the other hand, the fact that a final project can only cover a single topic from the course—or at most a handful of them—calls the value of such projects into question.

My final concern is that Machine Learning seems to have gone on autopilot at this point, with little or no attention from Ng or anyone else who helped him prepare the course materials. Questions in the discussion forum are answered instead by "Community TA's", that is, volunteers who took earlier sessions of the course. Most disturbingly, the majority of reports of errors in the course materials go unanswered, and those that are answered are answered by Community TA's, who lack the ability to fix the errors. For example, a month ago I discovered that the automatic grader accepted one version of my code and rejected another, even though the two versions were algebraically equivalent. My report of this apparent bug still hasn't been answered.

Despite these concerns, I still heartily recommend Machine Learning as a valuable starting point for anyone interested in data science. While the course was offered twice in 2013, the start date of the next iteration, on June 16th, 2014, suggests that Coursera may be planning to offer sessions of the 10-week course almost back-to-back, meaning several sessions each year.

What's next for me? I'll soon be posting a review of Udacity's short Intro to Hadoop and MapReduce. After that, I'm considering taking two more courses from the Data Science specialization, first Exploratory Data Analysis, which will give me some practical experience with graphics programing in R, and then Practical Machine Learning, which will provide experience using R for machine learning, as well as a basis for comparing the machine learning course reviewed above (though the course for the Data Science specialization, at four weeks, is much shorter, and can't possibly cover the same ground).

In the meantime, while I'm still looking for work as a data scientist, I've had a number of interviews, and some of the potential employers have read and commented positively on this blog. I hope that provides an example for other social scientists out there that, yes, you can become a data scientist.

Thursday, May 8, 2014

Online Course Reviews: The Data Scientist's Toolbox, and Getting and Cleaning Data, from Coursera's Data Science Specialization

I recently completed Coursera's The Data Scientist's Toolbox and Getting and Cleaning Data, two courses that form part of the online learning provider's new Data Science specialization, taught by Brian Caffo, Jeffrey Leek, and Roger D. Peng, biostatistics professors at Johns Hopkins University, and, in the cases of Leek and Peng, authors of the Simply Statistics blog. Both of the courses I took were taught by Jeff Leek (referred to in my earlier post today). I found Getting and Cleaning Data to be an especially useful course, teaching some practical skills that are quite essential to the real-world practice of data science. However, I probably wouldn't recommend the entire specialization to anyone coming from the world of quantitative research in academia, since a big focus of the program is teaching the scientific method and the logic of statistical inference—that is, things a quantitative social scientist should know already. First, however, a little background on Coursera's specializations....

Coursera has recently introduced a handful of "specializations", each consisting of a series of short courses followed by a capstone project. The specializations continue Coursera's effort to monetize its offerings through the Signature Track, which offers a "Verified Certificate" for those who pay a fee (typically about $50-$100) to take the course.

The Signature Track in itself has dubious value. Allegedly, its purpose is to provide a more useful credential than the certificates Coursera has traditionally offered for its free classes. To make the Verified Certificate more useful (that is, more impressive to potential employers), Coursera takes measures to guarantee you did the work yourself, but these measures seem fairly easy to circumvent. Specializations add an additional sweetener: if you take every class in the specialization on the Signature Track, you can then take the capstone project (offered as an additional class), which is not available to students who take the courses for free (or even, for that matter, to students who pay for only some of the courses). Students completing the specialization also receive a specialization certificate.

The Data Science specialization includes 10 short (four-week) classes, including the capstone, each priced at $49 for the Signature Track. If you stump up the whole $490 at once, you can take any of the courses as many times as you like over the next two years (in case you don't pass the first time); if you pay for the courses one at a time, you can only retake each one once (which is probably enough—honestly, if you can't pass one of these classes, you probably don't belong in the profession, but sometimes life gets busy, and you can't finish the work for a class). Each of the first nine courses will be offered once a month; the first six are available already, and the remaining three will be offered for the first time in June. For a couple of the classes, there's also an option to substitute an alternate course on Coursera. The capstone has yet to be schedule (word in the forums has it that it'll be offered in fall), and I'm not sure how often Coursera plans to offer it.

The course that really interested me was Getting and Cleaning Data, but I signed up for The Data Scientist's Toolbox because it's required for the rest of the specialization; R Programming is also required, but I already had some experience with R, and I had no intention of completing the entire specialization, and so I skipped this one. Taking one of the later courses at the same time as the required intro course didn't pose any difficulties for me, but I think that someone who has no experience with R would probably want to complete that class before tackling any of the others.

Much of The Data Scientist's Toolbox is devoted to introducing the topics of the specialization's other eight courses; frankly, you can skip this if you don't intend to take those courses (or possibly, even if you do intend to take them—you will, after all, cover that information later, though if you're taking the whole specialization, you may need to watch the video lectures in question in order to complete the quiz for Week 1). For me, the most useful content of this class was its introductions to Git, GitHub, and RStudio (I had been using the plain old R Console, and RStudio makes things considerably easier). RStudio is required for the programming necessary to complete the assignments in the later courses, and Git and GitHub are necessary to complete the projects at the end of each course (you have to upload your work to GitHub so that other students can peform peer assessments on it). For the sake of full disclosure, let me say that I skipped the introductory lectures in Week 1 of this course (though I did pass all the quizzes), and did not complete the course project, which consisted of taking screenshots to prove that you'd installed Git, GitHub, and RStudio (I installed all three, but I wasn't really concerned with getting the course certificate).

I found Getting and Cleaning Data invaluable. I took the course because I wanted to learn how to get data off the web. For example, in the project I did for Coursera's Social Network Analysis last year, I ended up saving data from several hundred web pages by hand, which is not a particularly efficient way of doing things. Getting and Cleaning Data promises to teach students how to extract data from common data storage formats (including databases, specifically SQL, XML, JSON, and HDF5), and from the web using API's and web scraping. The syllabus also includes tips on using R to clean and recode data, and, in the last lecture, a long list of links to sources of data. It's also worth noting that the style of the video lectures is a bit different from those of other classses I've taken: there's never any video of the instructor, just the instructor's voice over the lecture notes.

Initially, I was skeptical, because most of the lectures amount to little more than a list of R packages, functions (with a few short examples), and links for further information. The information blows past you so fast that there's no hope of remembering much of it. However, the lecture notes (in both HTML5 and PDF—the HMTL5 is a little awkward to navigate, but the links work, unlike in the PDF) provide a wonderful resource that you'll find yourself referring to again and again. I've often found that the hardest part of a project is knowing where to start, and the lectures in Getting and Cleaning Data point you in the right direction; in fact, I'm using information from the lectures on web-scraping and JSON right now to do an updated version of my project for Social Network Analysis, a statistically informed visualization of which cards in the game Android: Netrunner appear together in the decks designed by players. Look for that to be posted here soon!

Among the data science courses that I've taken online, Getting and Cleaning Data is the first one that taught me how to go out and get data and then put it in a form that's usable for analysis. By contrast, Coursera's Machine Learning, taught by Stanford's Andrew Ng, provides highly practical advice on selecting and using algorithms, but does so uses very much canned programming exercises, in which the data has already been collected and processed. In fact, the two course are highly complementary, at least inasmuch as they give you ideas about how to handle different stages of a data science project. It should though be noted that Machine Learning uses Octave (essentially the open-source version of MATLAB) rather than R; the Data Science specialization includes its own (much shorter) Practical Machine Learning course, as well as an earlier course on Regression Models that delves far more deeply into that topic than does Machine Learning.

I should add that, for this class too, I never completed the final project: it looks like a highly practical exercise, but I was short on time, and more interested in my own project; again, I didn't care much about earning a certificate, with my main concern being to learn the nuts and bolts of getting data from the web.

Finally, let me offer a few comments on the Data Science specialization as a whole. I would not recommend completing the entire specialization for anyone who's well-versed in statistics and the scientfic method: if you're a competent social scientist (as opposed to someone who took one stats course as an undergraduate), you already understand important issues like sampling, causal inference, and reproducibility (though, admittedly, I've read more than a few articles by social scientists who evidently had shaky grasps on these concepts). For a specialization that labels itself as "Data Science", there's also scant coverage of databases. That being said, anyone interested in data science might find Getting and Cleaning Data, R Programming, and Practical Machine Learning useful, and for someone who doesn't have a background as a quantitative researcher, I can't recommend this specialization's focus on the scientific method and applied statistics highly enough.

Why Data Science Needs Statistics

If you've read my earlier posts about why a scientific approach is important to data science, you won't find it surprising that I recommend Jeff Leek's recent post on the Simply Statistics blog, "Why Big Data Is in Trouble: They Forgot about Applied Statistics". Leek, a biostatistics professor at Johns Hopkins, and one of the instructors in Coursera's Data Science specialization, argues that a number of recent big data failures, including that of Google Flu Trends, can be chalked up to a lack of statistical knowledge among the researchers in question. Leek cites sampling, data collection, causal logic, model specification, and sensitivity analysis as areas where a solid knowledge of applied statistics could have prevented serious errors. It's a short but cogent read.

Friday, March 14, 2014

Why Scientists Make Better Data Scientists

Have a look at this blog post by Mike Walker on why it's useful for data scientists to have a scientific background. The link came to me in a list of "featured articles" I receive weekly from Data Science Central. The tl;dr is that analysts without scientific training (the author singles out those with undegraduate business degrees) lack the tools for distinguishing correlation from causation. This leads to a range of maladies, including spurious correlations, cherry-picking data, and stringing "disconnected facts" together to construct a fallacious narrative. Walker acknowledges that not all successful analysis requires starting out with a hypothesis, but stresses that there are scientifically rigorous ways to explore data for unexpected relationships, such as A/B testing.

I find this refreshing, after spending a great deal of time lately looking at job ads for data scientists: most ads focus on experience with specific software packages, rather than experience conducting rigorous research. I suppose the former is more of an objective measure than the latter, but I'm not sure how useful it is to hire based on what applications a person has used before, especially in a profession where the start of the art changes rapidly. Another problem is that people have started slapping the word "data scientist" on a wide variety of jobs: I've seen it applied frequently to database architect positions, or even to positions that have more to do with software development than data analysis.

At the moment, all of this matters to me because the contract on which I was working ended last December, and I'm now looking for a job again. I've had two good interviews, but I'm finding it very hard to break into a profession with a background different from traditional data analysts and business analaysts. One thing I have learned is the power of networking: one of my interviews came from a contact my wife made while carpooling, and the other resulted from my submitting a resume to a small-business group recommended by a former co-worker. (Oh, and if anyone has any good job leads, I'm happy to network here, too. :)


Those of you who frequently visit my links page might have noticed that I've updated it quite a bit over the past few weeks, particularly in the sections covering online courses ("Self-teaching Resources" and "Formal Learning Resources"). Coursera and Udacity have some interesting new offerings that you might want to check out. I'm also planning to add a section listing portals and other commercial websites, and I need to go through all the links to make sure the information on them is up to date. As ever, if you have any suggestions for additional resources, please let me know!

Wednesday, October 9, 2013

Online Course Reviews: Coursera's Machine Learning and Probabilistic Graphical Models

Whoops, I haven't posted in a while.

In May, I started a new job. It has nothing to do with data science, but it has given me experience in supervising other writers, and it's also kept me quite busy. The fact that work kept me busy explains why I haven't posted recently. It also explains the one caveat I have to add to the reviews I'm about to give you: I was never able to finish all the material for either course. I got busy with the new job and moving my family into a new house, and by the time I came up for air, it was too late to finish.

As I mentioned in my April post, I signed up for Coursera's Probabilistic Graphical Models, Machine Learning, and An Introduction to Interactive Programming in Python. I dropped the An Introduction to Interactive Programming in Python almost immediately, after realizing that the course's focus on programming video games made it not as useful for my purposes as I had hoped.

Probabilistic Graphical Models was taught by Stanford Professor and Coursera co-founder Daphne Koller. Coursera hasn't yet listed a new iteration of it, but if the previous pattern holds up, it should be offered again next year. As I mentioned before, I took this course because it includes Bayesian and Markov models, both of which show up in many job ads for data scientists. I decided not to take the optional programming track, figuring that it probably wasn't a good idea to be writing programs for two different courses in a language I was just learning (both this course and Machine Learning use MATLAB and/or the very similar Octave).

Machine Learning was taught by Andrew Ng, also a Stanford professor and Coursera co-founder, and is one of Coursera's best-known and most popular courses. It's also been taught by the University of Washington's Pedro Domingos, but Ng's version will be offered again starting October 14th. I signed up for the course because machine learning is one of the basic skills of data science, but I also wanted the chance to learn one of the most commonly used statistical programming languages, MATLAB/Octave.

As I said in my last post, Daphne Koller is not the most charismatic lecturer, and her explanations can be confusing. What I didn't say last time is that I don't think Koller entirely understands the medium in which she's working. In the classroom, asking questions of the professor can make up for a confusing lecture; Koller seems to be giving the same lecture she would give in the classroom, but without the opportunity to stop her and ask questions about each topic before moving on to the next, that same lecture doesn't work very well.

While the lectures are less than ideal, the quizzes are particularly troubling: rather than presenting a simple test of the material covered in the lecture, the quizzes ask students to move beyond the lecture material, drawing out implications on their own. Asking students to do this is a great pedagogical technique, especially in a graduate-level class. However, it works a lot better when the students have discussed the material in class, giving them the opportunity to start down that path together, with the professor's guidance. None of this is possible in an online class, and, even with discussion forums, rules that prevent students from providing answers to one another prevent full exploration of the quiz topics; part of the problem is that students can see the quiz questions before beginning their discussion, rather than receiving a quiz or homework assignment only after the classroom discussion is over.

It might help to begin each quiz with more straightforward questions, giving students a little practice, before moving on the ones that require additional thinking. Far from adopting this model, Koller actually exacerbates the problem by adopting an unusually strict rule (by MOOC standards) for retaking quizzes: any attempt after the second is penalized. Because of this, I found myself taking quizzes I had no way to prepare for, because they introduced concepts for the first time, and I had no way to practice applying those concepts beforehand.

I want to stress here that I'm not simply some idiot who was in over my head. I'm trained in statistics, and I have experience using structural equation and time series models, both of which share similarities with probabilistic graphical models—and I was really interested in the course material. Koller acknowledges in her lectures that the course is challenging, and even seems to take pride in that fact. However, while the material is indeed challenging, the course is hard partly because it's badly taught. It's also possible that Koller is trying to cover too much material for the online format—the lack of classroom discussion not only makes individual topics more difficult, but increases the time required to cover each topic, since the teacher has to provide a much more detailed lecture, rather than relying on student questions to fill in holes.

While I didn't pursue the programming track, other reviewers have complained that they spent more time trying to figure out how to read in the data than they did conducting the analysis. Mind you, this is a problem that data scientists face in the real world, and so the criticism might not be completely fair.

Andrew Ng's Machine Learning is another beast altogether. Ng is in fact a charismatic, and very clear, lecturer; indeed, Koller uses a couple of his lectures in areas where the material in the two courses overlaps. Not only does Ng convey his topics clearly, but he stresses the practical aspects of the methods he's teaching, and provides useful tips about how to apply them in the real world. While Ng pulls students along at pace much gentler than Koller's, he's still able to teach methods that, he insists, are advanced enough to be unfamiliar to many practicing data scientists. I should add that the automated system used to grade programming assignments works quite well. If I do have one criticism, it's that the programming assignments probably involve a little more copying and pasting than might be ideal for learning Octave, but then, copying and pasting isn't uncommon in real programming.

In short, this is a very good course, and I strongly recommend signing up for the session that starts October 14th. Now that things have calmed down a bit for me, I might even sign up for it myself.

Wednesday, April 24, 2013

Online Course Reviews: Coursera's Social Network Analysis and Foundations of Business Strategy—Plus New Courses to Check Out

I've recently completed Foundations of Business Strategy, taught by the University of Virginia's Michael J. Lenox, and I've submitted the final project for Social Network Analysis, taught by the University of Michigan's Lada Adamic, and I'd like to share some comments on both of these oferrings from Coursera, as well as give readers a heads-up to other courses that have just started.

Social Network Analysis provided a good survey of the methods and applications in the field, covering random networks, measures of centrality, small world networks (and other topics related to the question of optimization), and the dynamic aspects of networks, such as contagion and opinion formation. Adamic's explanations were usually clear, and even a student with little knowledge of probability could have gotten the gist of most of the course material (and made use of Gephi to perform basic analysis), but equations were presented for those who wanted them, and the readings gave further detail. In fact, this is the only course I've had so far that made extensive use of academic journal articles (and a few written for a wider audience), some of them required and some recommended—they gave a much better impression of the history of social network analysis and the current state of the art than Professor Adamic could have given by herself. From a personal perspective, this topic particularly interests me because I can see how social network analysis might be applied to the study of ethnic politics, my previous area of research.

The course's only weakness lay in the (optional) programming track: the first three programming assignments, two in R and one in NetLogo, were largely exercises in copy-and-paste, rather than posing full-fledged coding tasks; they were, however, enough to give students basic familiarity with the two programming languages, and with the igraph package for R. In contrast to these "canned" assignments, the final project was almost completely unstructured, and while this provided welcome freedom to explore whatever topic a student wished, it also meant a steep learning curve for someone whose experience with R or NetLogo was limited to the earlier exercises.

Compared to the other courses I've taken, Foundations of Business Strategy proved much less time-consuming, with required readings limited to very short chapters from a forthcoming book by Lenox (and when I say "short", I mean it's more a pamphlet than a book, with chapters only a few pages long), and a business case each week. The professor encouraged discussion of each case, both in small groups and in the discussion forum, and each week recorded a debriefing that made reference to students' comments in the forum; however, the only assignments that needed to ber turned in before the final project were (relatively easy) weekly quizzes that covered the lecture topics. Quizzes, hence the lectures the quizzes covered, could be completed at any time during the course, though completing the lectures late meant that a student had no chance to participate in the discussios of the associated cases. The final project was a 1500-word "executive summary" of a strategic analysis of an organization of the each student's choice.

Despite the sparseness of the course material, the class provided a useful framework for business strategy, a framework—and this is the part that surprised and impressed me, after all the scurrilous rumors I've heard about business schools, and the weak business students I've taught in my own classes—that was solidly grounded in microeconomics, with no mention at all made of the latest management fads. No, someone who took this course isn't guaranteed to become a strategic genius, or even, necessarily, an effective strategic thinker, but that's because strategy requires making decisions in an environment that's inherently complex and ambiguous, the upshot of which is that giving students a good framework for organizing thought—and a chance to practice strategic thinking on real cases—is about the best that a teacher can do.

My one serious concern with the course was the rubric used for peer review of the final project: the assignment presented a set of criteria for grading that focused largely on the quality of analysis and writing, and specifically warned against trying to use all the methods of analysis that had been covered in the course; by contrast, the rubric that was actually used for peer grading of the assignment amounted to a checklist of topics in the course, and penalized students for not covering each one, while leaving little room for judging the quality of the report. As a former professor, I certainly understand that detailed rubrics, while beloved by students, tend to push attention in grading towards mechanical aspects of the assignment, and this probably goes double for peer assessment, but it is possible to create rubrics that give more or less clear guidance for making qualitative judgments. More importantly, whatever rubric is used needs to match the criteria that are spelled out in the original assignment.

New Courses to Check Out

I've signed up for three courses that have just started, though odds are I'll need to withdraw from one of them, due to time constraints. I stumbled upon Probabilistic Graphical Models—the term "graphical" didn't suggest anything I was interested in, but a look at the course description revealed that probabalistic graphical models (PGM's) includes Bayesian and Markov networks, both of which feature in decision-making and machine learning, and both of which show up repatedly in job ads for data scientists. Coursera co-founder Daphne Koller, of Stanford University, lacks the charisma and clear explanations of the three MOOC professors I've previously learned from; she also comes off (if I may be subjective here) as a bit pretentious, an impression that makes her inclusion of the Simpsons' family tree as an example of a genealogical network more cringeworthy than cool. I've found most of the material so far to be readily understandable, but I'm guessing it wouldn't be for someone without my background in statistics (especially since I've used by structural equation and time series models in my research, and these feature many of the same concepts found in PGM's). This is a graduate-level course, and Koller herself describes it as challenging even by those standards. Needless to say, the other side of that coin is that anyone who gets through the course will have a solid foundation in PGM's, and also, for those taking the programming track, knowledge of Octave and/or MATLAB (the two are close relatives), especially given weekly programming assignments in a 11-week course.

I've also just started An Introduction to Interactive Programming in Python, taught by multiple instructors from Rice University, and Machine Learning, this iteration taught by Coursera co-founder Andrew Ng of Stanford, one of two professors on Coursera to teach this course; like Probabilistic Graphical Models, Machine Learning makes use of Octave. I doubt however that I'll have time for all three courses, meaning that I'll likely have to withdraw from one of them, which will probably be Probabilistic Graphical Models or Machine Learning, given their overlap in programming language, and, to a lesser extent, subject matter.

Thursday, March 14, 2013

Online Courses to Check Out

Having nearly completed Stanford Professor Jennifer Widom's Introduction to Databases, I've recently begun two more massive open online courses (MOOC's) that you, the reader, might want to take a look at, both of them offered by Coursera.

The first course is Social Network Analysis, taught by Lada Adamic of the University of Michigan. This methodology, which can be applied to topics as divergent as infrastructure and epedemiology (as well as the more obvious targets such as Facebook), obviously plays a prominent role in data science, which is one reason to take the course. A second reason is that the course features an optional programming track with four assignments (including a peer-graded final project), some using NetLogo and some using R, and in my case I'm taking the course in part as a way to learn R. The course also makes use of Gephi for basic network analysis. In the second week, there are two versions of the lectures, with an advanced version for students with a background in probability distributions and differential equations; it's not clear if this will be the case in later weeks. This is a nine-week class, and if you're reading this soon after I've posted it, you can still sign up and get full credit, since the first assignment isn't due until Friday night (March 15th).

Taking the advice of one of my contacts to learn something about business, I've also signed up for a non-technical course, Foundations of Business Strategy, taught by Michael J. Lenox of the Unviersity of Virginia. This six-week class features a textbook that Lenox is currently developing, as well as the case method typical of business-school education (Lenox recommends small-group discussion to get the full impact of this method). The most interesting feature of the course is a peer-graded final assignment in which each student writes a short but well-researched strategy memo for a the CEO of a company of his or her choice; more interesting still, Lenox has invited organizations that would like their strategy assessed to join the course and offer themselves as cases for the students' final projecdts. Though we're already about 25% of the way through the course, the assignments all have the same deadline of April 14th, and so it's easy to catch up.

You might also keep a lookout for two courses starting the latter half of April, An Introduction to Interactive Programming in Python, taught by a team from Rice University, and the perennially popular Machine Learning, taught by Coursera co-founder Andrew Ng of Stanford (this one uses Octave, a close relative of MATLAB, for those keeping track of programming languages).

Tuesday, February 26, 2013

Is There an Academic Role for the Social Sciences in Data Science?

I've spent most of the space in this blog exploring routes for a social scientist to become a data scientist. However, Justin Kern's recent article "The State of Business Intelligence in Academics" has turned my thoughts briefly back to academia. Kern reports on a survey by the organizers of BI Congress that found that, although business intelligence courses are taught primariliy by information technology and management information systems departments, they're increasingly being offered outside those discplines, in particular in finance, marketing, and accounting.

I wonder what role social scientists should be playing here? Should we merely be taking classes from other departments, or should economics, poiltical science, and sociology departments have their own offerings in data science? (Since this survey reported specifically on business intelligence courses, it's possible that similar courses in the broader field of data science were missed, but I would suspect that such offerings are pretty rare in the social sciences.) I'd love to get some comments on this subject.

Tuesday, February 19, 2013

Political Ideology and Consumer Brand Choice: Applying Social Science to a Marketing Problem

This morning Shep Parke sent me a copy of Harvard Business School newsletter The Daily Stat, which links to "Ideology and Brand Consumption", a 2011 paper from marketing professors Romana Khan, Kanishka Misra, and Vishal Singh. The jist is that buyers of consumer packaged goods (CPG's) in counties that have voted more Republican or where more people go to church buy more goods from established brands and fewer generic or store-branded goods, and also fewer goods from newer brands.

The authors' explanation of these findings is that people who have conservative ideologies are more fond of tradition and the status quo, and wary of change. The paper obviously caught my eye because of the political component, especially as it touches on voting behavior, which is one of my specialties. I thought that this article might provide me a chance to give an example of how a social scientist can be valuable in analyzing business data.

The value of a social scientist here might not be obvious: after all, three social scientists already did the hard work, and now, don't we have an actionable insight for producers of consumer products? Not quite: the paper's findings are interesting, and I'd have no qualms publishing them in an academic journal as a starting point for discussion, but I wouldn't risk money and brand loyalty on something this vague and uncertain. As a social scientist, not only can I identify the questions that this research doesn't answer, but I can also suggest some practical ways to answer those questions. The short version is that we need more detailed data, preferably at the individual level, and that we might even want to run a few experiments to test our hypotheses.

There are two basic issues here. The first is whether we can make the jump from county-level data to individual consumers. The second is whether conservative personality traits actually influence shoppers' buying choices, or whether there's something else going on that just happens to be related to both factors. Let's address the problem of county-level data first. The problem we have is that we know more products of certain types are leaving the shelves in conservative counties, but we don't know exactly who's buying them. For example, it's logically possible (if unlikely) that the liberals in conversative counties buy more goods from established brands than the liberals in other places.

A more realistic concern is that, as any student of market segmentation can tell you, human psychology doesn't divide us neatly into two big groups, "conservatives" and "liberals". The fact that a candidate has to win a majority of the electorate leads voters naturally to bunch up into two competing groups , but there's a lot of diversity within each of those groups, as all of us who went to college have probably seen in the two-dimensional political graph that college Libertarian clubs like to trot out. But though we only have two choices as voters, as consumers, we have many more, and people who vote together might not shop together.

In the nineteenth century, when the big issues were things like freeing the slaves or allowing men without property to vote, we could safely say that conservative people favored the status quo, and no doubt many "conservatives" still do, but are those people who seek the safety of the well-known really the same as Tea Party members who want a revolution to roll back decades of big government, or the libertarians who favor gay marriage as ardently as they do low taxes? If it's actually just one group of Republican voters that favors the tried and true, we'd get a lot more bang for our marketing buck by focusing directly on them, or, at least, on places where they make up the biggest part of the population. There are other questions we could ask here, but you get the idea.

Even if we can identify the relevant segment of conservative voters who buy established brands, how can we be sure their conservative personality traits are what lead them to make those buying decisions? This question of causality is the one that, more than any other, keeps social scientists up at night. Sure, the psychological explanation offered by the paper's authors is a plausible story, but you can create lots of different plausable stories to explain any given set of facts (anyone who doesn't believe that should consider how quickly the latest management and marketing advice changes).

Here's one plausible story: the authors looked at sales from the same chain of stores in different counties, but the same chain typically offers a different mix of products at different stores. Places that are less densely populated typically have smaller stores, which offer a narrower range of goods, and I'd be willing to bet that where stores offer a narrower range, those offerings are dominated by established brands. And do you know what else is true of less populated places (that is, rural areas vs. big cities)? They tend to be more conservative. In other words, it's entirely possible that people in conservative counties buy more established brands because they don't have much of a choice.

How does a social scientist address these issues to pull out some information that we can act on? First of all, I'd try to find some individual-level data. We may already have data on individual brand choices from store loyalty cards, but to make that useful, we need individual-level pyschological data—that is, we need to know that specific individuals with personality trait X buy brand Y. The closest thing to that we're likely to have is demographic data on the holders of loyalty cards (both the data we gather from the loyalty card program, and data we can obtain from other sources and join with the loyalty-card data), but that may actually be counterproductive: sure, people with high incomes are more likely to vote Republican, but what actually interests us is people who vote Republican because they favor tradition, and their demographic data doesn't tell us a whole lot about that or any other personality trait. We're not, after all, actually interested in whether or not people vote Republican, but in the personality traits that make them both vote Republican and buy one brand rather than another.

In the end, assuming I work for a retailer with a loyalty-card program, I would try to survey a sample of card-holders (perhaps we could offer them coupons or some other incentives to participate). Actually, I probably wouldn't even ask political questions in the survey, because personality traits are what we're actually after (even if the observation about politics was what inspired us in the first place), and political questions might well offend our shoppers.

Getting individual-level data would get us closer to showing that personality traits cause consumers to make particular brand choices, both by looking directly at the traits and choices, and by allowing us to rule out other possible causes—for example, we could look at demographic data and psychological data at the same time in order to see which are related more closely to brand choices. And frankly, in the social sciences, that's about the best we can usually do. The "gold standard", though, is randomized experiments, because, if you divide people into two, randomly-chosen and essentially identical, groups, and then do X to one and Y to the other, you can be pretty sure that any differences you see after that point are due to the difference between X and Y. We rarely do experiments that look at behavior in the real world (as opposed to a psychology lab), because they're expensive and they pose ethical questions, but they're pretty viable for a big retailier with outlets all over the place.

With personality traits, a true experiment is never possible, because you can't force people to have certain traits (how many parents, teachers, and managers have wished it were otherwise?), but we could, for example, make sure that two (or more) stores in places with different sorts of shoppers carried exactly the same selections in one or a few product groups, thus ruling out different lineups as a cause of different brand choices. This might cost us money (lost sales or extra inventory), but unlike an academic researcher, we can recoup that cost in higher sales that result from the new information.

It's interesting to note that this experimental approach can yield results even without individual-level data—that's the logic behind introducing products in test markets, after all—but if we can combine the experiment with a survey of the shoppers at the stores taking part in the experiment (even though using both the survey and the experiment is our most expensive option), we can leverage the data from the experiment and the survey to get more benefit out of both.


I'm an Author a Musician? Who Knew?

I've noticed that many of the books advertised by Amazon on the sidebar of this blog are by one "Scott Orr". Just in case it isn't already clear, that Scott Orr is not the same Scott Orr writing this blog. Heck, I don't even know who that guy is, though I'm sure he's perfectly nice.

UPDATE:. Actually, yes, I am an author. I wrote this, after all. I've written other things, too, for that matter, many of them published. More to the point, my namesake appears, on closer examination, to be selling music.