Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Sunday, September 25, 2011

It's not "the job market"; it's the profession (and it's your problem too)

I enjoyed Kathleen Fitzpatrick's recent piece in the Chronicle* on risk-taking and the responsibility of mentors to back up those junior scholars who are doing nontraditional work. The piece's key insight is that it's one thing to urge people to "innovate" and quite another to create the institutional frameworks that make innovation not only possible but consequential.**

Kathleen's observation comports with some ideas that have been floating around in my head lately, especially around "digital humanities." I and my Fox Center colleague Bart Brinkman were recently called upon to define digital humanities for the other fellows in residence, and in the process of talking it over with Bart, and during the discussion at the CHI, I've come to realize that I have some real pet peeves around the notion of the "job market" that come into relief specifically around the field of digital humanities.

It boils down to this: peeps, we're all connected.

The recent rise to prominence of digital humanities is indistinguishable from its new importance in "the job market" (I insist on those scare quotes); after all, digital humanities and its predecessor, humanities computing, have been active fields for decades. What's happening now is that they are institutionalizing in new ways. So when we talk about "digital humanities and the 'job market,'" we are not just talking about a young scholar's problem (or opportunity, depending on how you see it). We are talking about a shift in the institutional structures of the profession. And, senior scholars, this is not something that is happening to you. You are, after all, the ones on the hiring and t&p committees. It is a thing you are making—through choices that you make, and through choices that you decline to make.

There's something a little strange about the way that digital humanities gets promoted from the top down; it gets a lot of buzz in the New York Times; it's well known as dean-candy and so gets tacked onto requests for hires; digital humanities grant money seems to pour in (thanks, NEH!) even as philosophy departments across the country are getting shut down; university libraries start up initiatives to promote digital humanities among their faculty. I am waiting for the day when administrators and librarians descend upon the natural sciences faculty to promote history of science. No, I really am.

So it seems quite natural that there should be wariness and resistance to the growing presence of digital humanities. Perhaps there is some bitterness that you might get your new Americanist only on condition that her work involves a Google Maps mashup, because it was easy to persuade people that your department needed a new "digital humanist," whatever the hell that is, and it was not easy to persuade people that you needed somebody to teach Faulkner.

The situation is not improved by the confrontational attitudes of certain factions of the digital humanities establishment (such as it is), which are occasionally prone to snotty comments about how innovative DH is and how tired and intellectually bankrupt everybody else's work is. (Not so often, I find—but even a little is enough to be a problem.) Under those circumstances, DH seems clubby and not liberating; not a way of advocating the humanities but an attack on it, and specifically on the worth of that Faulkner seminar that you teach, and that non-digital research that you do. Why, an established scholar might reasonably ask, should I even deal with this "digital humanities" nonsense? Shouldn't I just keep teaching my Faulkner seminar, because somebody ought to do it, for Christ's sake?

Well, whatever else DH is, it is highly political, and it has political consequences. So, in short, no.

I'm persuaded that the widespread appeal of DH has much to do with the leveling fantasy it offers, a fantasy of meritocracy that is increasingly belied elsewhere in the professional humanities. As Tom Scheinfeldt points out in his useful "Stuff Digital Humanists Like,"
Innovation in digital humanities frequently comes from the edges of the scholarly community rather than from its center—small institutions and even individual actors with few resources are able to make important innovations. Institutions like George Mason, the University of Mary Washington, and CUNY and their staff members play totally out-sized roles in digital humanities when compared to their roles in higher ed more generally, and the community of digital humanities makes room for and values these contributions from the nodes.
This is true. Those involved in digital humanities have also seen the ways that THATCamps, blogs, and Twitter allow junior scholars and scholars at non-R1 institutions to cut geodesics across the profession, allowing them to spread their ideas, collaborate, and achieve a certain prominence that would have been impossible through traditional channels. I'm convinced that real possibilities lie here.

And as traditional scholarly publishing becomes more and more constricted and humanities department budgets are slashed, the fiction of academic meritocracy becomes harder and harder to sustain. Perhaps on the web, we think, through lean DIY publishing and postprint review, meritocracy (or its semblance) can return to the academy. It seems at once a way forward and a way to return to a (fabled) time when people cared about scholarship for the sake of scholarship—not because they needed X number of well-placed articles or a line on the cv or a connection at Y institution without which their careers would disappear. Perhaps DH offers us a way out of the increasingly rationalized death-spiral of "impact scores" and credential inflation. Perhaps it will let us out-quantify the quantifiers, or sidestep them altogether.

Of course, the web always comes with liberatory rhetoric that usually turns out to mean little more than "what the market will bear," and the ostensible meritocracy of digital humanities in the present moment is really no more than a misalignment between its alternative (and potentially even more aggressively capitalistic) value systems and those of the institutionalized humanities more generally. It can be disturbingly easy for the genuinely progressive intentions of digital humanists to become assimilated to the vague libertarianisms of "information wants to be free" and "DIY U," and from there to Google Books and charter schools and the privatization of knowledge—an enclosure of the digital commons ironically in the name of openness. At the same time, the naming of the "alt-ac" "track" (it is generally not a track, of course, by definition) seems to provide new opportunities for young scholars even as it raises research expectations for staff and requires those on the "track" to subordinate their research interests to those of the institutional structure that employs them. Digital forms are exceptionally good at occluding labor. How to navigate those waters thoughtfully—to realize the real promise of DH—is a question to which we must all apply ourselves.

So you see what I mean when I say that "digital humanities and 'the job market'" as it now manifests isn't a narrow, merely administrative sliver of life of interest solely to junior academics who are still gravely listening to advice about how to "tailor" the teaching paragraphs in their cover letters. Digital humanities has become important to "the job market" exactly insofar as it is causing major shifts in the institutions of the profession. These shifts are political. And if you are in my profession, then they are your concern.

*I know, "enjoyed" and "Chronicle" in one sentence... mirabile dictu.

**As we all know, I have a complex relationship with the word "innovation" and do not consider it an unqualified good, nor a transhistorical value. For today, however, we will leave that particular word a black box.

Thanks to Bart and Colleen for sitting through a less-worked-out live version of this rant last week.

Friday, August 5, 2011

Via Rebekah Higgitt, a Tumblr on "The Art of Google Books." They're images that break the illusion that the books have been spiritually whooshed whole and entire into the ether.

[Related.]

Thursday, August 4, 2011

New stuff

I sure have been neglecting this blog lately, and all I have of late is a little linkspam.

First of all, Ryan Cordell writes on the overly-Facebooky-but-okay-I-guess Google+ about a cool project for people who want to get started in digital humanities but aren't sure where to begin:
Know someone who wants to get started in the digital humanities but doesn't know how to do so? They should apply for DHCommons' preconvention workshop at MLA 2012, cosponsored by NITLE and the Texas A&M Initiative for Digital Humanities, Media, and Culture. Representatives from a range of prominent DH projects and centers will be on hand for training and consultation.

Apply here.

And second, we've finally launched Colloquies at Arcade! Here's my Ed Blog post introducing Colloquies, the Colloquies landing page, which will soon have more than one Colloquy on it (specifically, on September 1, when I release the next Colloquy), and our very first Colloquy, "The Contemporary Novel," introduced by the great Lee Konstantinou.

And, shhh, Arcade may also be seeing a long-awaited up-hay-ade-gray.

Oh, and I switched my California driver's license over to a Georgia one. Guess I'm a Georgian now.

Image: Peaches. Bryan Costin, 2005. CC NC-BY-SA 2.0.

Thursday, July 21, 2011

Another question about G+ integration with other Google products: when will I be able to share a GDoc with a circle?

And: when that happens, will Google's total infiltration of the universe be scary/ier?

And: have cats weighed in on the issue?

(Duh.)

Sunday, July 10, 2011

G+, briefly

I'm trying out Google Plus, because I'm a sucker like that, and also because Google already owns most of my life, so what's the loss? (Copies of my diss on GDocs, etc.)

Apart from a brief exploration of Facebook (which is ridiculous), I have hitherto confined my internet activities to the public: a blog with my name on it, a Twitter account with my name on it. This seems to me to be right and just. (Well, I am pseudonymous on one other social network that will remain nameless but which is by far the best designed and most useful social network I have ever seen.)

G+ offers the same temptations as Facebook: the walled garden, the ability to form little clubs. That's both the good news and the bad news, I guess.

I've heard it said (a lot) that big search is dying (because spammers and similar are so good at SEO) and social search is the future. That strikes me as likely. This changes the nature of the "publicness" that I've tried to maintain in my web presence, but I'm not sure how just yet.

It occurs to me that there may be a day when G+ has nicer integration with Blogger blogs than just the ugly "+1" button, since both are Google properties. (The social network that shall go nameless has fairly nontrivial integration with blogs.) Blogs are said to be an old new medium, but I still like them. They're a damn good place to put text.

A note on commentary through taxonomy:

One thing I love about Twitter that G+ doesn't have is hashtags. This is a feature of its nonpublicness. Tagging is one of the best things about the web; commentary through taxonomy has become standard, and this is a curious and lovely thing. So far you can't really do it with G+. But this is the internet, so I'm sure people will eventually find a way.

Wednesday, May 4, 2011

Rough draft

It strikes me that Tuesday's post is actually just an expansion of a series of tweets and retweets. This blog is called Works Cited, so in the interests of citation, here is, as it were, the rough draft:



Semirelatedly, apparently Google has just renamed its search group the "knowledge" group.

This is completely hilarious.

Tuesday, May 3, 2011

The visible hand

Things on the internet are not made by magic; they're created by human labor. Who pays for that labor, and to what ends? Often, private corporations like Google pay for it. Wherefore?

Siva Vaidhyanathan's recent piece in the Chronicle argues that "Our uncritical dependence on Google is the result of an elaborate political fraud. Google has deftly capitalized on a decades-long tradition of creating 'public failure,'" which is to say, setting public projects up to fail so that private interests can swoop in and save us from our "broken" public sector:
Public failure may occur when the public sector has been intentionally dismantled, degraded, or underfinanced while expectations for its performance remain high. [...]

In such circumstances, the failure of public institutions gives rise to the circular logic that dominates political debate. Public institutions can fail; public institutions need tax revenue; therefore we must reduce the support for public institutions. The resulting failures then supply more anecdotes supporting the view that public institutions fail by design rather than by political choice.

[...]

Google officials, promoting their effort to scan millions of books purchased with public money [e.g. University of California, University of Michigan --N.C.] and donated by shortsighted universities, claimed they were trying to preserve libraries and perform an essential public service—just the sort of service that our great university libraries could have been working toward had they been allowed to succeed. Publicly supported institutions fail, so we leap into the arms of the private actor, ready to believe its sweet nothings.

Google Books is certainly read by most as a sort of public service that happens to be provided by a private corporation. Remember when the Bibliothèque Nationale de France resisted Google's digitization offers, only to later concede that they lacked funds to carry out their digitization project (the excellent Gallica) on their own? "La BNF se laisse séduire par Google," Le Figaro reported, using the language of sexual danger that Vaidhyanathan picks up in his Chronicle piece.

I'm largely persuaded by Vaidhyanathan's argument, although the persistence of this language of seduction (all literary critics know what comes next: betrayal) probably warrants further cogitation.

That Lovelace Google has practically unlimited funds to pour into whatever it wants is widely taken for granted, and it's well known as a place that is generous with said funds, especially with its workers. But despite its much-touted mission of non-evil (evil is such a strong word, isn't it?) its practices seem increasingly disturbing, including when it comes to digitization. Via Ryan Shaw, I recently came across the bizarre narrative of Andrew Norman Wilson, who says he was fired from a Google contract after inquiring into, and trying to document, the working conditions of "ScanOps":

They scan books, page by page, for Google Book Search. The workers wearing yellow badges are not allowed any of the privileges that I was allowed – ride the Google bikes, take the Google luxury limo shuttles home, eat free gourmet Google meals, attend Authors@Google talks and receive free, signed copies of the author’s books, or set foot anywhere else on campus except for the building they work in. They also are not given backpacks, mobile devices, thumb drives, or any chance for social interaction with any other Google employees. Most Google employees don’t know about the yellow badge class. Their building, 3.14159~, was next to mine, and I used to see them leave everyday at precisely 2:15 PM, like a bell just rang, telling the workers to leave the factory. Their shift starts at 4 am.

They are not elves; I repeat, not elves. Today Glenn Fleishman tweeted the picture below (via @GreatDismal):

hand spotted in a Google Book by Glenn Fleishman

Whose hand is this?

The image reminded me of Caleb Crain's post on encountering the finger of a Google technician in a translation of a Kant essay. As he wrote in a review of Adrian Johns's recent book on copyright and piracy,
... Kant didn't think that an author could mount a strong legal case against piracy based on property rights in words. After all, even after pirates copied an author's words, the author himself still had them. It was better for an author to argue that his book was not an object but an exercise of his powers which "he can concede, it is true, to others, but never alienate". In other words, Kant explained - in a passage partly obscured by the fingers of the Google technician who turned the pages in the scanner - a pirated book was not to be understood as property that had been stolen; it was rather a speech act that had been compromised. The business arrangement that an author made with an editor might make it look as if words could be traded like watches or pork bellies, but it just wasn't so. Could there be a fitter representation of copyright's contemporary plight than the fingers of a Google technician obscuring Kant's defence of writer's rights? An author's consent, Kant cautions in a footnote, "can by no means be presumed because he has already given it exclusively to another", yet Google is struggling to effect exactly this sort of transfer of consent today, as it attempts to win approval for a legal settlement in the United States that will allow it to republish works whose copyright owners have not come forward. I couldn't have read Kant's essay so easily without the Google technician's labour - in fact, without Google, I might not have got around to reading it at all - but her fingers were nonetheless in the way. The internet's attitude toward Kant's words is ambiguous, combining respect, appropriation, liberation and accidental vandalism.
hand spotted by Caleb Crain

This scan is particularly ghostly, the hand covered over with a second hand reasserting the text of the Kant translation.

The hand--always the synecdoche for the worker (the mediator between the head and the hand, we learn in Metropolis, must be the heart)--is inserted literally into our view of the text, disrupting for a moment our sense that Google Books are, quite simply, books that have been "put online," as if books themselves could simply leap media and enter a disembodied realm. The intrusion of the hand shows us that these are photographs (of a sort) and that someone must have made them.

In an inversion of our usual intuition that images are less mediated than text, these hands make us realize that Google Books made us feel as though digital texts were unmediated--were the books themselves. In contrast, the awareness that the digital object is an OCRed image of text--a photograph of its own scene of production, complete with visual evidence of the hand that wrought it--forces us to acknowledge the strange backwards ekphrasis (text to image to fallen, "corrupted" text--OCR is a silent diplomatic edition) in a Google Book, the labor by which it was created and uploaded, and the person who labored, now knowable only through the operative, synecdochal appendages that both create and corrupt the digital object.

This is not to argue for some kind of metaphysics of book presence wherein only a paper book is a real book, not haunted by ghostly disembodied hands. Our tendency to efface the digital laborer, as well as the work of editors, designers, etc., is precisely what enables the widespread belief that e-books are necessarily cheaper to produce than paper books, as if the cost of the book lay in the printing. At least with the heft of paper one is reminded that there was, somewhere, a scene of labor. A Google Book effaces the medium of the medium, until a latex-draped finger appears before us, as if to reassert the tactile element always running beneath the digital.

Obviously, this is a Blogger blog, i.e. run on Google resources.
More on GB hand scans:

Sunday, December 19, 2010

A Supposedly Fun Thing: Text-Mining and the Amusement/Knowledge System; or, the Epistemological Sentimentalists

If we could text-mine the internets of the last few days for the correlation between the words "n-gram" and "fun," I'm sure we'd get a nontrivial number. One of the most striking things about the reception of the Google Books Ngrams, largely in the form of the web tool, is the giddy delight with which people have announced how much fun it is. Exhibit A is the bit I quoted yesterday from Patricia Cohen at the New York Times:
The intended audience is scholarly, but a simple online tool allows anyone with a computer to plug in a string of up to five words and see a graph that charts the phrase’s use over time — a diversion that can quickly become as addictive as the habit-forming game Angry Birds.
But that's just one example--the fun of the Google Books Ngrams tool is almost universally noted. See, for instance, "Fun With Google's Ngram Viewer" (Mother Jones), "Fun with Google NGram Viewer" (WSJ), and "BRB, Can't Stop NGraming" (The Awl). And Dorothea Salo tweets,
What I like about the GBooks n-grams is seeing all kinds of people playing with it. Just playing. THAT, friends, is how one learns.

The prevalence of this language of play raises two questions.

1. What rhetorical work is this move (calling Google Books Ngrams a fun toy) doing?

2. What experiential dimension of Google Books Ngrams does this rhetorical move describe, and what does it tell us about the tool's epistemic significance?

Answering the first question feeds into answering the second. To call the Google Books Ngrams web tool (henceforth "GBN") a fun toy is to hedge one's bets, to express approval without necessarily venturing into the higher-stakes terrain of approving it as a research method. Any assessment of the tool's epistemic value is channeled through an expression of pleasure (or, as Patricia Cohen and The Awl's Choire Sicha rather interestingly suggest, compulsion). Play can of course be a form of learning, and very important--that's what Dorothea Salo's tweet indicates. But play is a good learning environment precisely because the stakes are low and mistakes can be made safely, as a comment by Bill Flesch suggests: "I played around with it for about half an hour. Now I'm bored." New toy, please! With respect to knowledge, the language of play is deeply ambivalent.

As I read it, the universal declaration of fun that has surrounded the release of GBN is as much about guilt as about pleasure. Those who are compulsively "ngraming," as Sicha so amusingly puts it, are often all too aware of GBN's limitations, which have been blogged extensively, all the way down to what Natalie Binder points out, in her much-retweeted post, has to underlie the whole operation: inevitably imperfect OCR.*

Why does the GBN web tool even exist? Not to advance knowledge, I don't think, or at least not directly, but rather because it's fun. Because it directs interest toward the more substantive element of the project, the downloadable data set that relatively few people are actually going to download.

There are huge problems with using GBN (and throughout I'm alluding to the web tool/toy that everybody is saying is so much fun) as any sort of meaningful index of culture, and everyone knows it. And yet.

I would argue that the universal declaration of fun is a form of confession: I am deriving epistemological satisfaction from this unsound tool, with its built-in Words for Snowism. It's a guilty pleasure, epistemic candy: the sensation of knowledge, lacking in any nutritional value.

But the guilt goes rather deeper than the simple tension between GBN's unreliability for actual research and the "gee whiz!" quality of the graphs: GBN is fun because it is so limited.

That great scholar of nineteenth-century culture, Walter Benjamin, described a mode of writing that he called "information."
Villemessant, the founder of Le Figaro, characterized the nature of information in a famous formulation. 'To my readers,' he used to say, 'an attic fire in the Latin Quarter [Paris] is more important than a revolution in Madrid.' This makes strikingly clear that what gets the readiest hearing is no longer intelligence coming from afar, but the information which supplies a handle for what is nearest. Intelligence that came from afar--whether over spatial distance (from foreign countries) or temporal (from tradition)--possessed an authority which gave it validity, even when it was not subject to verification. Information, however, lays claim to prompt verifiability. The prime requirement is that it appear 'understandable in itself.' (147, emphasis added)
What GBN delivers is information in this sense. It is near at hand, easy to use, and puts out a nice visualization that appears "understandable in itself." It's easy to deliver, in that way, not unlike a pizza. It's no good to point out, as Mark Davies does, that the Corpus of Historical American English (COHA) allows one to look at specific syntactic forms, or include related words, or track usages by the genre of the source. Such capacities only raise anxieties. (For example, what gets tagged as "nonfiction"? Where, for instance, do autobiographies go? I once, to my astonishment, saw The Autobiography of Alice B. Toklas in the nonfiction section of a book store--along with Three Lives! But I digress.)

As soon as we raise such questions, the graph stops being "understandable in itself," stops being information. Conversely, when you aren't given the choice to sort by genre, how genres are defined necessarily stops being a question. It's the very fact that the toy is a black box and a blunt instrument that makes it feel immediate and incontrovertible and, in that very satisfying way, obvious. We get the epistemic satisfaction of information, and the thing that gives it to us is precisely that information's lack of nuance.

Yesterday I used the word "cheap" to describe the kind of historical narratives GBN suggests. There is indeed a kind of economic dimension to the satisfaction that GBN delivers. Of Oscar Wilde's many quotable lines, I am reminded of this one:
The fact is that you were, and are I suppose still, a typical sentimentalist. For a sentimentalist is simply one who desires to have the luxury of an emotion without paying for it. (768)
Feeling, Wilde suggests, has to be earned.** Bracketing the question of whether this is a good description of sentimentalism, it's a good analogue for the epistemic candy of GBN. One receives the apparent solidity of research--the nice graph that summarizes and visualizes what might otherwise be years of labor in the making--without having to have actually done any research. This is only a cheap thrill, "fun," when it is actually cheap--that is, when we don't inquire into how the corpus was prepared, or what effects GBN's case-sensitivity is having on our results.

The analogy to sentimentalism is useful not only because it gives us a model for understanding the economy of feeling here, but also because it allows us to recognize that there is an element of feeling in the way that we encounter information. We are likely to find it ethically reprehensible when our emotions or what we believe we know are manipulated. And yet there are times when we want the cheap thrill. Most people I know will freely cop to liking a good emotionally manipulative movie or novel, whether a thriller or a romance or one of those movies where the dog dies. As the fun of ngrams demonstrates, we like a little intellectual manipulation too.

(I know, I know, it doesn't tell you anything conclusively, but...try Foucault versus Habermas!)

What does it mean, this liking it?

I mentioned Bill Brown's term, the "amusement/knowledge system," in my title above because it's another, perhaps more explicit way of describing the close interweaving of knowledge and fun at the end of the nineteenth century that so fascinated Benjamin (208). In my own work I have tried to make a case for taking seriously both the knowledge and the amusement in that system, notably in naturalist fiction, because it's often in such liminal places that the terms of what counts as knowledge are most at stake. Part of the reason experimental literature seems to be here to stay is that the amusement/knowledge system is, too.

The point is not to condemn fun as something that has no place in knowledge--far from it. Fun is central to how we vet knowledge--just think of how important it is that research be "interesting"! It is our highest (and also most common) praise.*** Indeed, play lies at the heart of our most cherished models of intellectual inquiry--a nonutilitarian curiosity to "see what happens." As I quoted Dorothea Salo at the beginning of this post: "THAT, friends, is how one learns."

So condemning fun is not at all on my agenda. Rather, I want to draw attention to the emotional content of the way we talk about knowledge, and to the ambivalence that intellectual "fun" signifies. Ours is an age of "news junkies" (again with the pleasure bordering on unpleasurable compulsion, à la the "addictive" ngrams) and "armchair policy wonks" and people who read voraciously, but only in the proverbial dubiously defined "nonfiction" category. Nate Silver and the Freakonomics dudes are minor celebrities. Lies, damned lies, and statistics are our idea of fun, as powerfully as a Victorian melodrama was ever considered fun. Which means we need to think much more about how fun operates, and why, and what that means for knowledge. And just as crucially: what knowledge means for pleasure.


*In fairness, Ben Schmidt argues that GBN's OCR is pretty accurate, given the state of the field, and also that "No one is in a position to be holier-than-thou about metadata. We all live in a sub-development of glass houses." But there's a big difference between "this is really good, for OCR" and "this degree of accuracy is good enough for supplying evidence for X kinds of claims."

**Taken out of context, Wilde appears here to be describing sentimentalism through an economic metaphor. In fact, it's rather the reverse, or at the very least something more confused than that: most of the surrounding text is taken up with Wilde chastising Douglas for his financial mooching.

***As Sianne Ngai points out, the "interesting," like the language of play, has a hedging quality, bridging epistemological and aesthetic domains.

Benjamin, Walter. "The Storyteller: Observations on the Works of Nikolai Leskov." Trans. Harry Zohn. Selected Writings: Volume 3, 1935-1938. Ed. Howard Eiland and Michael W. Jennings. Cambridge, Mass.: Belknap-Harvard UP, 2002. Print.

Brown, Bill. The Material Unconscious: American Amusement, Stephen Crane, and the Economies of Play. Cambridge, Mass.: Harvard UP, 1996. Print.

Ngai, Sianne. "Merely Interesting." Critical Inquiry 34.4 (Summer 2008): 777-817. Print.

Wilde, Oscar. "To Alfred Douglas." Jan.-Mar. 1897. The Complete Letters of Oscar Wilde. Eds. Merlin Holland and Rupert Hart-David. New York: Henry Holt, 2000. Print.

Previously on text-mining:
Google Books Ngrams and the number of words for "snow"
Dec. 16, 2010
Dec. 14, 2010
Google's automatic writing and the gendering of birds

Friday, December 17, 2010

Google Books Ngrams and the number of words for "snow"

As I mentioned yesterday, Google has put out a big data set (downloadable) and a handy interface for tracking the incidence of words and phrases. As many have pointed out, one can do a lot more with the raw data set than with the handy, handy online tool, but it's that latter that the New York Times called
a diversion that can quickly become as addictive as the habit-forming game Angry Birds.
(I've never heard of Angry Birds, but that's the kind of thing I'm likely to be out of the loop on, so okay.)

I said yesterday that Google Books Ngrams was a lot more sophisticated than Googlefight, and it is. But I'm troubled by the model of cheap history that's presented in the NYT article--as if to suggest that if you want to do cultural studies now, all you need to do is Google (Books Ngram) it:
With a click you can see that “women,” in comparison with “men,” is rarely mentioned until the early 1970s, when feminism gained a foothold. The lines eventually cross paths about 1986.

You can also learn that Mickey Mouse and Marilyn Monroe don’t get nearly as much attention in print as Jimmy Carter; compare the many more references in English than in Chinese to “Tiananmen Square” after 1989; or follow the ascent of “grilling” from the late 1990s until it outpaced “roasting” and “frying” in 2004.

“The goal is to give an 8-year-old the ability to browse cultural trends throughout history, as recorded in books,” said Erez Lieberman Aiden, a junior fellow at the Society of Fellows at Harvard.
I will concede that newspaper articles are necessarily glib, but it's easy to see how the fallacy that this article promotes would be broadly accepted. The first quoted paragraph above correlates the incidence of words with known historical events; the second moves on to suggest the ngrams' predictive capacity. There's a narrative implicit in each statement of "just the facts," only the assumptions that go into them are effaced.

Let's look at the first of these reports: "With a click you can see that “women,” in comparison with “men,” is rarely mentioned until the early 1970s, when feminism gained a foothold."

The implicit narrative is that nobody even bothered to talk about women until second-wave feminism came along. In fact, if you go by the incidence of the words "men" and "women" in the Google Books Ngrams data set, sure, you might be tempted to really believe that the 1970s was the time "when feminism gained a foothold." I can imagine the suffragists who fought for and won the franchise that I as a woman can enjoy annually asking, "what are we, chopped liver?"

What distinguishes the feminist movements of the 1970s, for the purposes of this data set, is its renewed attention to language. The suffragists wanted a policy change: they wanted the vote (and the freedoms that the vote could give them). The second-wave feminists wanted policy changes too (still working on that wage gap, people!) but they also wanted a deeper change: they wanted to change the way we thought about women and--here's the kicker--spoke about women. The 1970s is when it became broadly recognized as problematic to treat "man" as a synonym for "person," and I suspect that a significant percentage of the uses of "men" were and remain the "universal" usage. That's a nuance that the online Ngrams tool can't give you ("with a click").

Likewise, if you got your understanding of history through Google Books Ngrams, you wouldn't expect to hear this from 1929:
Have you any notion of how many books are written about women in the course of one year? Have you any notion how many are written by men? Are you aware that you are, perhaps, the most discussed animal in the universe? Here had I come with a notebook and a pencil proposing to spend a morning reading, supposing that at the end of the morning I should have transferred the truth to my notebook. But I should need to be a herd of elephants, I thought, and a wilderness of spiders, desperately referring to the animals that are reputed longest lived and most multitudinously eyed, to cope with all this. I should need claws of steel and beak of brass even to penetrate the husk. How shall I ever find the grains of truth embedded in all this mass of paper, I asked myself, and in despair began running my eye up and down the long list of titles. Even the names of the books gave me food for thought. Sex and its nature might well attract doctors and biologists; but what was surprising and difficult of explanation was the fact that sex--woman, that is to say--also attracts agreeable essayists, light-fingered novelists, young men who have taken the M.A. degree; men who have taken no degree; men who have no apparent qualification save that they are not women. (27)
That's Virginia Woolf, of course, giving a fictionalized, subjective encounter with the British Library. Yes, it's a bit longer than a sentence, and you have to read it; you can't just click! But it gives you much more women's history than does the Google Books Ngrams example cited by the NYT.

Google Books Ngrams is a fun tool (as everyone keeps pointing out) and, if you download the data set, even a useful one. But it can only get you so far, and uncontextualized, it encourages assumptions that it does not announce. I mention the number of words for "snow" in my title above because it's a famous fallacy--the notion that Inuit has [insert high number here] words for snow, always with the implicit suggestion that having a lot of words for something means that something is extremely important to the culture. Language Log uses this as their go-to example of stupid assertions about language widely believed by the public; it's a cheap Whorfism, claiming broad cultural significance for something incidental. We have a widely accepted term for a magical being that flies by night and runs a clandestine cash-for-baby-teeth operation. That doesn't make it central to American culture. ("Mom, is the Tooth Fairy real?" "Yes! Check Google Books Ngrams if you don't believe me!")

There's a certain Words For Snowism in the online Google Books Ngrams tool, the suggestion that the more frequently a word is used, the more important it is in a collective unconscious of which the Google Books data set serves as a convenient index. This importance is not the same thing as significance, in the sense of significant digits or statistical significance; it's not the difference that makes a difference, but rather a psychologized importance--attachment, cathexis. Which is really kind of garbage.

The web interface is, as my friend Will says, a toy. For the serious scholar, there's much more to be done with ngrams, and one can be careful as well as lazy with the conclusions one draws. But the toy has a "boom! proven with statistics!" quality, a reality-effect that's enormously pleasurable, even, as Patricia Cohen writes for the NYT, "addictive." (That's the point of toys, isn't it?) That's why I'm inclined to agree with Jen Howard, who writes that her "skepticism is mostly directed at how people will use it and what kinds of conclusions they will jump to on dubious evidence." That sort of jumping is practically built into the ngrams tool.


Woolf, Virginia. A Room of One's Own. Annot. and introd. Susan Gubar. 1929; Orlando: Harcourt, 2005. Print.

Previously on text-mining:
Dec. 16, 2010
Dec. 14, 2010
Google's automatic writing and the gendering of birds

Friday, December 3, 2010

Google's automatic writing and the gendering of birds

The almost meaningless faux-text-mining of a Google search on "birdlike woman" and "birdlike man" turns up the following results:

Vanilla Google:
"woman""man"ratio "woman"/"man"
"birdlike"16, 1002, 9905.38
"bird-like"74,400272,0000.27


Google Books:
"woman""man"ratio "woman"/"man"
"birdlike"1, 5206062.5
"bird-like"6334901.29

This probably tells us more about Google than about the correlation of gender and the term "birdlike." The hyphen makes a big difference in the search. This particular search also doesn't catch instances like "her movements were quick and birdlike."

I often think it would be interesting to do some small bit of real text-mining, just to have a global look at a corpus, but it's always incidental to the argument, so I never follow up.

The appeal of text-mining, which I think is actually magnified in the Google search, is that it's a kind of automatic writing, in which the body of the text (corpus) is made to give up its latent spirit. That the Google algorithm is unknown except insofar as it is known to maximize ad revenue does not diminish this appeal, the temptation to present Google hits as data. Since so much of our daily information is filtered through the Google algorithm anyway, it serves as a sort of corporate unconscious, whose essence is perhaps more compelling than truth.

The appeal of the Google search in lieu of text-mining is formalized in toys like Googlefight, which simply runs two Google searches at once and visualizes the results:

(Source.)

The bar graph calls on a visual form designed to represent meaningful data; although of course such forms are routinely abused (I particularly enjoy April Winchell's pie charts), the form still invites one to seriously compare the numbers. Yet the tongue-in-cheek cheesy stick-figure animation acknowledges the unseriousness of the Google fight. A Google fight is only good for settling a certain kind of argument, the confrontational flame-war variety that isn't particularly invested in actually solving a problem, not a debate but a "FIGHT." (I tried to get a screen shot of the "FIGHT" title, but I'm just not that quick on the draw, apparently.)


Yet for all that, toys like Google Fight are amusing (try Foucault versus Habermas!) and a little beguiling. I don't have time to prepare a corpus and an algorithm, but I do have three seconds to do a Google search, or make a Wordle.

Word cloud for Beatrice Forbes-Robertson Hale's The Nest-Builder (1916).

Such tools get you somewhere; they just don't get you far. It's interesting ("merely" interesting?) that the above word cloud says nothing about birds or nests, and that some of the most prominent words are "know" and "time." But of course not all words are weighted equally in a novel, and it matters that the chapters are titled "Mate-Song," "Mated," "The Nestling," "Wings," etc.--that indeed the whole marriage plot is structured around a bird allegory that disappears in the word cloud. And this may be another reason it's so appealing to let a simple Google search stand in for data, even when its unreliability is universally acknowledged. It gets you somewhere but it doesn't get you far, and in the end this is true of most text-mining, too. In the end we're fascinated by automatic writing, the possibility of forcing the body to secrete a hidden spirit, but we're also agnostic about spirit tout court. A highly sophisticated search with a known margin of error probes an ontological terrain that's suspiciously similar to the corporate unconscious, which we're tempted to say is all phony advertising anyway--or it isn't--one or the other.