Literary Criticism for an AI Age
In my humanities-dominated circles, I risk being labeled a heretic by admitting that I think AI prose can be quite skillful, often better than what most college students are capable of.
he New York Times ran an experiment back in March that asked readers to blindly choose which they preferred between a series of juxtaposed passages, one written by an artificial intelligence model and the other written by a Homo sapien. I am disobedient, and so I instead attempted to guess which one was written by AI — so as to avoid giving the machines my precious vote.
My performance left much to be desired. It was genuinely very hard to tell which passage was synthetic and which anthropogenic.
Our heuristics for detecting AI tend to focus on the uncannily “fluid” quality of most AI prose, as the Times notes. Real authors, with their distinguishing idiosyncrasies, tend not to be so smooth. Cormac McCarthy doesn’t much like punctuation, for instance, and Sally Rooney takes a defiant stand against the presumed tyranny of quotation marks.
We may loosely think of any given AI-generated text as a probabilistically determined response, essentially simply predicting the next word in the sentence, producing a sort of weighted-average response. This is why AI prose can read as so uninspired; it is definitionally average.
Yet models can also be prompted to produce less probable outcomes, or to optimize for a different norm — say, a bizarrely capitalized Trumpian norm, or a Dostoyevskyian norm of long, roundabout tortured monologues. They can overcome their voice problem by emulating particular stylistic patterns.
The Times links a study that found “readers prefer outputs of AI trained on copyrighted books over expert human writers.” The study asked readers to indicate their preference between two attempts at emulating an array of eminent authors, one written by a human, the other by a large language model “fine-tuned” to the given author.
Even the MFAs preferred the AI attempt. We may glean from this that it is not necessarily because someone is poorly read or uncultured that they might prefer an AI-generated text over a human one.
In my humanities-dominated circles, I risk being labeled a heretic by admitting that I think AI prose can be quite skillful, often better than what most college students are capable of. The position of some proponents of the humanities, who utterly reject the notion that much of human writing can be imitated and mechanically reproduced by machines, is not a sustainable one.
If there is continued merit in reading human-generated text — which I think there is — it cannot stem solely from the aesthetic quality of the prose.
An age of AI generated texts that are indistinguishable from, and indeed frequently preferred to, certain works penned by humans, demands of us a more expansive sort of criticism that puts more weight on authorship and the humanistic virtues of literature, not merely the empirical quality of prose.
***
Critics have long debated the extent to which we may consider the author when interpreting a work of literature.
W. K. Wimsatt Jr. and M.C. Beardsley argued in their famous (or infamous, to my high school English teacher) 1946 essay “The Intentional Fallacy” against trying to discern the author’s intentions when interpreting the meaning of a work of literature. Only public evidence ought to be considered (i.e. the work), not the private journals or letters of the author. What mattered was not the meaning the writer intended to invoke through his words, but the meaning he actually achieved.
A later critical counterrevolution notwithstanding, Wimsatt and Beardsley helped redefine the relationship between the author and the text within literary criticism, loosening the iron control the author once held over his own text and rendering the text a more independent, public artifact. If we wish to bring the author into greater focus, the natural question arises about whether a machine can have intentions — and how we should consider their relationship with the text.
In some sense, LLMs “intend” to choose the next best word in the sentence; they “intend” to produce the most likely desired output. How does this compare to human intentions? I might think I intend, by composing a sonnet, to woo the recipient of my tender affection. Or I might intend to express a deep emotion, or a beautiful truth about nature. But is my true intention that expression, or is there something more reptilian motivating me?
One can, in principle, trace back our every intention to evolution. I intend to woo my lover because, biologically speaking, such relations tend to result in reproduction, something selected for by evolution. An innocent expression of beauty might link to a baser intention to better connect with one’s community, which improves survival and is thus also selected for by evolution.
If we are to distinguish anthropogenic texts from AI generated texts, we cannot do so on the basis that our intention in creating that text was superior, or indeed that any intention we have actually renders prose of a higher quality.
Two decades after “The Intentional Fallacy,” a Frenchman called Roland Barthes went even farther than Wimsatt and Beardsley, arguing in “The Death of the Author” that there are in fact no authors at all, only “scriptors.” Barthes pedalled in semantic autonomy, believing in the supremacy of the text and that meaning is created by the reader. The scriptor’s interpretation of what she has written is as legitimate as anyone else’s.
If, to Barthes, the identity of the person that produces a text is not terribly relevant, would he care whether the scriptor was human at all? There is a particular line from Barthes that, 60 years of technological advance later, takes on a quite striking meaning:
“We know now that a text is not a line of words releasing a single ‘theological’ meaning,” he writes in Death of an Author, “but a multi-dimensional space in which a variety of writings, none of them original, blend and clash. The text is a tissue of quotations drawn from the innumerable centers of culture.”
Even the reader who is most remote from the mechanistic details of LLMs must find this strikingly similar to the way we describe an AI model’s workings. According to Barthes, we might not have any right at all to have a discerning literary distaste for AI-generated works, because humans are no more authors than models. Creativity, to Barthes, is just a reworking of previous forms; there is nothing really original.
Perhaps, then, literary critics ought to listen to Barthes and be satisfied by a world replete with thousands of AI written books that are all very “good” but not written by humans. I do not think such a world is desirable. While it might be valuable to consider the aesthetic qualities of a work over the qualities of its author, not all of the value we take from prose is an appreciation for technical prowess or its literary merit. We also appreciate fresh ideas. We have a conception of whether art is “interesting” that is not entirely a function of its impressiveness.
In this vein, some argue that though AI may write excellent (or at least not bland) prose when it is emulating existing patterns that humans have thought up, it cannot be creative itself and therefore produces nothing of interest.
Such is the view of the science fiction writer Ted Chiang, who describes in The New Yorker a theory of creativity that centers around choices. When John H. Updike ’54 wrote “Rabbit, Run,” he made many hundreds of thousands of both conscious and unconscious choices about his prose; the totality of these choices can be said to constitute the creativity of Updike.
When one writes a prompt for an AI model, by contrast, they make only a few hundred words worth of choices, deferring to the model’s judgement for the vast majority of the choices manifested in the final product. As Chiang writes, plenty of people have an idea for a novel, but ideas alone do not constitute a meritable or worthwhile creative work. The little details are hardly a “nuisance,” as idea-men might think, but are actually what make the artwork substantive.
When we cede so many choices to an LLM, when we abandon the choice of each word and the grammar of every sentence that coalesce to mean something, do we not lose the very artistry of the work? The prompter is in no real sense determining any of those small decisions; there is none of the special care that renders art meaningful.
I am not entirely convinced of the notion that there is some capital-C “Creativity” that is fundamentally impossible to replicate with a model. Creativity must not become an unfalsifiable ideal of the human; there are empirical ways of assessing creativity (such as the Torrance Tests of Creative Thinking), and that models are usually plagiarists does not imply that they are never creative.
Randomness — trying many thousands of patterns of choices — can surely yield at least a few meritable novel results. Barthes has a point when he says that authors are more influenced by what has already been written than they know. (But then arises the matter of taste: it matters what we think is good, not what the AI thinks is good.)
Yet at the same time, I think that we must be interested in the human behind the work. When we read descriptive writing, do we not value that rich imagery was envisioned by a human? Do we not value viewing the world through the discrete lens of a peer, rather than through the homogenized view of the model, defined by probabilistic emulation of training data? Do we not value that an author is documenting the world around her, filtered through her own subjective lens?
Literature is also moral, acting as a tool of empathy and communion; creative expression is a virtuous and meaningful endeavor. It makes little sense to mechanize virtue.
Tobias Wolff’s “Old School” is a highly autobiographical work that offers a very particular window into the literary culture of certain New England boarding schools, filtered through the subjective lens of a pupil who was expelled for academic fraudulence. Not only do I fail to see how an AI model could tell such a unique story except by plagiarizing other pre-existing stories (or entirely manufacturing a non-existent perspective), but I fail to see what interest or moral value such works would have.
Certainly, I think an LLM could replicate the qualities of Wolff’s prose that so please my eyes, but not the aura, not the meaning, not the subjective truth in Wolff’s work. Subjectivity is at the heart of good literature, and it is both discrete and highly personal. Wolff’s writerly choices are one with the meaning he imbues and such a voice can only be superficially emulated, can only be faked.
—Magazine writer Andrew W. Shlomchik can be reached at [email protected]. His column, “The Humanist,” explores what it means to be human in an era of unprecedented technological turmoil.
