Why AI Can’t Be Creative (Yet)
What happened when I asked Claude to write the most creative sentence in human history.
Last night I suddenly got an idea. I opened up Claude, the overly-friendly neighborhood AI bot, and asked it: “Hey buddy, write the most creative sentence you possibly could.” Claude almost immediately responded:
“She kept her childhood in a jar of fireflies and opened it only on the nights the dark felt honest.”
Cringe!
Really, Claude? That’s literally the best you can do? So I decided to go even bigger and give it another chance. The biggest, in fact: “You know what, write the most creative sentence ever written in human history!” A few seconds later, Claude responded:
“The universe wrote itself into being one lonely word at a time, and named the ache between the words you.”
Bigger cringe!
How can this be? The second prompt, which was about ALL OF HUMAN HISTORY was even worse than the first! So this led me down a whole rabbit hole linking this to my lifelong research on the science of creativity and what you get is this 100% human-written post. Let’s go.
✦✦✦
Sixty years ago, the creativity researcher E. Paul Torrance was frustrated that the field of psychology was neglecting the study of creativity. Psychology was so focused on IQ (hello Lewis Terman and his obsession with gifted children with high IQs) that barely any researchers found creativity important enough to study. Enter Torrance.
He created his legendary “Torrance Tests of Creative Thinking”, which measures creativity along four major dimensions:
Fluency: The sheer quantity of meaningful responses. For example, if the question is “list all the uses for a paperclip”, 30 responses in the span of a few minutes would demonstrate high fluency.
Flexibility: The ability to shift categories. For example, a high flexibility score would include uses that cross categories such as office supplies (holding documents together), tools (picking a lock), hygiene (toothpick), and arts/crafts (making a paperclip sculpture).
Originality: The statistical rarity of the ideas. Most people would immediately think of “hold paper together” but an original response would be “bend it into a miniature replica of the Empire State Building”.
Elaboration: The depth and detail of the ideas. An example of high elaboration would be something like “Use it as a plant support system to train and support young seedlings. You can stretch out the paperclip and bend it into an elongated “S” hook, sticking one end into the soil next to the stem and wrapping the upper loop gently around the plant. This provides rigid vertical stability to prevent the stem from snapping while allowing the plant to grow naturally.”
Torrance used his tests to follow children from elementary school all the way into adulthood to see how well they predicted creative life outcomes. One follow-up study by Jonathan Plucker found that just under half of the variance in adult creative achievement is explained by his creativity test scores, with the contribution of Torrance’s tests being more than three times that of IQ scores.
Clearly there are deep implications here for human intelligence. But in this article I want to explore something that isn’t explored nearly enough, and that’s the deep implications for artificial intelligence.
✦✦✦
Let’s zoom in on two of the major scoring dimensions: fluency and originality. Torrance’s brilliant observation, which has been empirically supported by many years of creativity research, is that fluency and originality are partially separable aspects of creative thinking. Yes, they are positively correlated with each other (people who can come up with a lot of ideas tend to also come up with more original ideas overall), but this isn’t necessarily the case. They are worth measuring separately.
Now let’s return to Claude’s responses and how it actually got more predictable when I asked it to be more creative. When Claude raised its weights to the maximum of creativity— the most creative sentence in all of human history— Claude went right to the dead center of the archetypical “profound sentence” cluster of its training: the cosmos, loneliness, love, the intimate second-person perspective. To Claude, the deepest thing it could come up with is an assemblage of a million things that have already been labeled deep during the course of human history. So the bigger the demand for originality, the harder the machine regressed to the average of what it thought “genius” is supposed to sound like. In a nutshell: It literally gave me the most average version of profundity imaginable with only a few seconds of processing.
Here’s the incredible irony as I see it: AI just isn’t built for rarity, which is literally the definition of “genius”!! Turn the creativity dial to maximum and you can pulled as close as possible to the average of “sounds like genius”, but you get something that clearly is not genius.
Now let’s analyze for a moment why neither sentence it gave me is actually creative (and why that’s actually built into the very nature of AI). According to Torrance and modern-day creativity researchers, creativity involves both novelty and meaningfulness. Just originality and you get a schizophrenic word salad. Just meaningfulness with no originality a you get an everyday thought. A sentence is only creative if it says something the way almost nobody else would say it in a way that people (even if it’s only domain experts) can comprehend. But current AI models are built entirely from phrases and images that lots of people already reach for when they’re trying to sound deep or creative.
Take: “She kept her childhood in a jar of fireflies and opened it only on the nights the dark felt honest.”
The idea of putting a memory in a jar is a common convention for creative writers. This is Hallmark Greeting Card common. “Fireflies” are the go-to image for a wistful childhood. We’ve all seen it a thousand times. “The dark felt honest” may on the surface sound meaningful but take a moment and ask yourself: What the hell does that actually mean??! It’s literally a mood pretending to be a thought. The key here is that nothing in this creative attempt could actually be true for any one real human living person. There is no specific memory here, and no real detail. Anyone could have written it.
I only discovered Claude recently, and I’ll be honest: At first, I did think some of its editing of my drafts were profound. Wow, that’s so clever to add the word “quiet” there, wish I thought of that myself! But then it started replacing my writing with the same words over and over again to the point where I would keep banning it from using certain words because they became so grating to me.
Also, I would tell Claude to try some actual variety, but it just couldn’t. It’s almost as though it couldn’t process what “variety of word choice” even means. I’ve become so frustrated by just how unoriginal and repetitive AI really is that I decided to break up with Claude. I told it that. I will have it let me know if I have any spelling mistakes or hugely egregious errors in reasoning but I don’t want it touching any of my human-made sentences ever again!
✦✦✦
I want to end this article on a note that illustrates why current AI models will not be taking over human writing anytime soon. Here’s the thing: What I am writing right now at this very moment comes from MY heart. Not the average heart of human history, but my fucking heart (and yes, Claude would never spontaneously come up with a curse word in that context. Most people don’t put the word “fucking” before the word “heart”).
The thing that makes human writing so special is that we can write from a real lived human experience and we can choose to deliberately create sentences and use word choices that are infinitely original and interesting. AI models in their current LLM instantiation are built to NOT be capable of doing that because statistical rarity is the opposite of what it’s grasping for when it composes sentences.***
Interestingly enough, the way AI detectors work is by analyzing how much a certain chunk of text is predictable. Because that’s all AI knows how to write. Just for kicks, I went onto Pangram and put this Claude-generated text in there:
“She kept her childhood in a jar of fireflies and opened it only on the nights the dark felt honest. The universe wrote itself into being one lonely word at a time, and named the ache between the words you. Grief moved into the guest room and paid rent in small forgettings.”
It came back: “100% of this text is AI.” I then replaced it with a paragraph from this very article, and it came back: “100% of this text is Human Written”.
Isn’t that interesting? I don’t think the value of human writing is going away anytime soon. I, for one, am going to stick to my own soul revealing itself on the paper. Not only does that feel more fun and meaningful, but I also bet that sentence right there is probably actually more creative than any other sentence ever written in human history (according to Claude).
*** Just as this article was going to press, Cameron Berg (an AI researcher I’m interviewing) told me that Claude is the best AI writer he’s used. I don’t doubt that. But “best writer” is meant here as fluency, not originality. Philosopher Henry Shevlin frames the distinction nicely using Margaret Boden’s three types of creativity . LLMs handle combinatorial and exploratory creativity, but transformational creativity (radically new ideas that redefine an entire domain) is exactly what they can’t do (yet).



I love this Scott - you inspired me to try the remote associates test (RAT) with AI and it crushed it. Then I asked it to create the most creative and challenging RAT item and this was it:
Sight • Wind • Nature
It took me a couple minutes to solve it. Then I asked it for the most creative sentence ever and it was pure cringe: "Consciousness is the universe remembering that every boundary is a useful fiction, so that through countless finite lives it can discover, by freely choosing truth over comfort, that the deepest form of power is to become the kind of being capable of loving reality exactly as it is while never ceasing to transform it." This is so interesting. Compared to Dylan (not cringe): "Ain't it just like the night to play tricks when you're trying to be so quiet"
I don't use AI for specific language or editing (I created a set of principles for myself to govern my use of it. One of them is "AI should never speak for you"), but one thing I have found it occasionally useful for is to put an entire essay into Claude and ask "Where does this drag? Where does it get boring?" Sometimes I take its answer into account, other times I ignore it, but it can help offer me a new lens to reconsider something I've written and if there's a better way to say it.