On Apples and Fixing AI Writing: The Spatial and Sensory Qualities of Text

Famed AI researcher Yann LeCun, one of the visionaries of artificial intelligence and a winner of the Turing Award who was formerly the head of Meta’s artificial intelligence programs, has argued that LLMs will never be able to replicate human thought because they do not have models of how the world works. As Newsweek summarized, “For LeCun, today’s AI models—even those bearing his intellectual imprint—are relatively specialized tools operating in a simple, discrete space—language—while lacking any meaningful understanding of the physical world that humans and animals navigate with ease.” LeCun is correct about the limitations of LLMs, but since he is a computer scientist not a linguist or a humanist, the problem is worse than he knows.

The first erroneous assumption is that language is a simple, discrete space and not also an artifact of the physical world—it is. First, the very process of reading itself is spatial. Scientific studies of the reading process have established that when we read, we ingest much more than the words themselves. The simplest way we do this is to encode where we read something on a page. Anyone who has written a term paper and tried to find something they had read will remember this feeling—“It was somewhere on the bottom right of a page, about a quarter of the way through the book.” Interestingly, this experience vanishes when reading an endlessly-scrolling tablet text, which is one reason why readers retain less from a digital reading experience.

But our brains also record a lot of information that might seem entirely extraneous, for example, where we were when we read; the way the light from a window slanted across the page and then moved, ever so slowly, as the window panes formed a sundial across the surface; whether we were satiated as we read, and the growing feeling of hunger as hours passed by; and the network of visual and other sensory associations that each sentence and thought created. All of these other sensory and spatial tags create an intricate lattice that then allows us to recall the information in the text, even though it happens subconsciously.

In addition, the reader pulls each new thought into a pre-existing reservoir of other thoughts gained from other experiences—not only other books and articles, but music, political actions and commitments, movies, conversations, arguments, and training from educators and other experts. In other words, our brains house intricate, four dimensional, ever branching synaptic models in which language is just one form of encoding but in which every word, phrase, and idea interacts with literally trillions of other pieces of information. To think, therefore, that one can approximate human writing by only studying human text is a vast underestimation of how human brains work.

 My writing is deeply visual and spatial. When I think of  the words “New York City,” my mind immediately takes me to a place that I have not only read about but lived in and visited for years, and unspools millions of images, sounds, smells, perceptions of humidity, heat, cold, shapes, and human relationships, over the course of four decades of my life. For example, I first see the tall buildings of Midtown—for some reason, my brain zooms down like a diving bird on the Lower West side of Manhattan, over the 30 and 50 story towers of what used to be called San Juan Hill, not on the more iconic buildings to the East and South. I then sense the bustle of the human traffic and the loud yellow taxis, and smell and see a hot dog stand. I have distinct memories for each borough and neighborhood at various times of day and night. I remember buying and biting into a bagel with a giant plasticky slab of cream cheese and a coffee when I was 23 years old and rushing to work, the breakfast of so many commuting champions. I feel the sensation of humidity as I descend into a subway. I remember looking up one of the canyons of the avenues in midtown, and remember crossing the anodyne avenues of modest apartments along Lower Lexington Avenue.

I remember the intersection of 135th and Lenox outside the Schomburg Library where I spent so much time researching and writing African American history, and I remember walking hand-in-hand with my fiancé on a snowy January first through a picture-postcard Central Park, the brisk air hitting our faces. All of these memories and more flood me as soon as I read the words “New York City,” a complex, four-dimensional matrix of information that is not spatial, it is emotional and multi-sensory.

And those are a tiny fraction of my trillions of datapoints of memory acquired through lived experience; I also can recall everything I remember about the city from reading about it in novels and history books; I could tell you about the gang turf wars on “San Juan Hill,” and the battle in Cuba for which it was named, and the racial politics of each, as well as the way that urban renewal programs wiped it out and planted the Lincoln Center for the Performing Arts on top of it; and I could give you a brief explanation, from memory of all of these concepts, eras, movements, and the economic, racial, and even sexuality dimensions of each. As a historian of New York City, I can also remember which authors made which arguments having something to do with the city, and I can bundle those thinkers into various tranches and know, at least on a surface level, how each tranche grew from and differed from other ones.

But what I’ve described doesn’t make me special. Every human does something similar, even if they are creating and working with text and image in a culturally-devalued and widely-disparaged medium like Tik-Tok.

In other words, if the model designers and engineers who build LLMs think that their models understand what the words “New York City” means because they are comparing it to trillions of other textual mentions, they are totally missing the point of how brains manipulate, store, and recall language. And here is where LeCun is correct: simply throwing more computing power at the problem won’t work. Even if we keep the LLM approach, we need to move outside the merely textual if we want to approximate how humans manipulate text. But what I am describing is a neural network so detailed that it may simply be outside the constraints of silicon, today or tomorrow.

The Big Apple example suggests that today’s LLMs can precisely model the surfaces of an apple, but they lack the flesh itself and all of its internal connections: the flesh, the seeds, the stem, and the juice.

Next
Next

Doing More with Less: an LLM Response to Less is More.