What AI Gets Wrong When It Plays a Personality
Researchers used my Light Triad scale to test whether a machine can simulate moral character. What they found says something hopeful about being human.
Every few months I get a version of the same question: now that AI can sound like anyone, can it be anyone? Can a large language model take on a personality the way a good actor does, and reason the way that kind of person actually reasons?
A new paper out this month in Ethics & Behavior put that question to the test, and it did so on territory I care about. A team led by Nuermaimaiti Wubuli compared how dark and light personality traits shape moral judgment in two very different kinds of minds: 404 human participants, and 2,092 responses generated by large language models (GPT-4.1 and DeepSeek).
To measure the bright side of human nature, they used a scale I built with my colleagues back in 2019, the Light Triad — humanism, Kantianism, and faith in humanity — the prosocial counterweight to the Dark Triad of narcissism, Machiavellianism, and psychopathy (plus, in more recent work, sadism).
They scored everyone’s moral choices using the CNI model, a clever tool that pulls apart three things people are usually doing all at once in a moral dilemma: their sensitivity to consequences, their sensitivity to moral norms, and their general preference for inaction over action.
Here is the first finding, and I will admit it was nice to see. In the human sample, the pattern held the way the theory predicts. People higher in dark traits were less sensitive to moral norms. People higher in light traits were more sensitive to them. Machiavellianism and sadism pulled norm-sensitivity down; Kantianism pulled it up. And people high in faith in humanity were more willing to act when a dilemma landed in front of them, which makes a lot of sense: if you trust that people are basically good, you are more likely to step in. This was a sample collected in China, on the other side of the world from where the scale was built, and the structure mostly carried.
But the part I keep thinking about is the AI.
When the researchers prompted the models to “be” a high-Machiavellian, or a deep narcissist, or a profoundly humane person, and then run the same dilemmas, something revealing happened: The machines matched only the direction of the human pattern — dark prompt, lower norm-sensitivity; light prompt, higher — and missed almost everything underneath it. Their dark characters didn’t just score low on moral norms. They scored close to zero, in a flat, absolute way that no actual human being does. The responses were rigid, extreme, and cartoonish.
And here is the kicker: The models couldn’t tell the dark traits apart. Machiavellianism, narcissism, psychopathy, and sadism are four distinct constructs with four distinct inner logics, and they all came out as basically the same response pattern. Same story for the light traits. The researchers handed the models careful background on what makes each trait different, and the models still collapsed them into one generic “bad person” and one generic “good person.” Their own conclusion: today’s LLMs show “only a preliminary simulation capacity” and “fail to reflect the complex features of personality and moral judgment.”
Now, the lazy version of this take is “AI bad, humans good,” and that is not what I think. The models got the direction right, which is genuinely useful. As the authors point out, that makes them a decent tool for a rough first pass, a way to guess which way an effect might lean before you go collect real data. Yes, and: a rough first pass is not a person.
What the machine reproduced was the stereotype of a personality. What it could not reproduce was the texture. A real Machiavellian and a real narcissist walk into the same moral dilemma and come out at different places. A genuinely humane person sometimes breaks a rule precisely because they care about a human in front of them. An actual person integrates this specific situation with a lifetime of accumulated experience, and the result is particular in a way that resists averaging. The models had read everything ever written about narcissism, and they still couldn’t do narcissism — because doing it takes something other than having read about it.
I have been reading a lot lately about how so much AI-generated prose feels empty even when it is perfectly fluent, and this study points at that same emptiness from another direction. A language model hands you the average of what a “cruel person” or a “passionate person” sounds like. The thing it overlooks is the only thing that was ever interesting in the first place: the particular human, with their particular history, doing the unaverageable thing in a particularly morally-challenging situation.
I don’t take this as a warning about machines so much as a reminder of what we are. Your conscience is the running, lived, context-soaked work of a single irreplaceable life, and at least for now, the most sophisticated models on earth keep quietly averaging that part away.
Which is worth thinking about. The most human thing about your moral life may turn out to be exactly the thing no one has figured out how to fake.



The essay identifies a real limitation of current language models, but it also smuggles in a stronger conclusion than the evidence can support.
The study shows that contemporary LLMs, when prompted to inhabit personality constructs such as Machiavellianism, narcissism, psychopathy, sadism, or the Light Triad traits, tend to generate flattened caricatures rather than psychologically differentiated agents. That finding is unsurprising. Present systems are trained to predict text, not to instantiate enduring motivational architectures. Asking them to "be a narcissist" often activates a semantic cluster associated with narcissism rather than a deeply integrated cognitive structure that persistently shapes perception, memory, attention, emotional valuation, and decision-making.
But from this observation, one cannot infer that personality itself is fundamentally inaccessible to artificial minds.
The deeper question is whether personality is a special substance or whether it is an emergent consequence of information-processing systems operating under constraints. Modern neuroscience overwhelmingly favors the latter interpretation. Human personality appears to arise from interacting mechanisms involving reinforcement learning, memory consolidation, predictive processing, social cognition, emotional regulation, developmental history, genetics, and environmental feedback. No known scientific theory requires a non-computational ingredient.
If personality is an emergent process rather than a metaphysical essence, then the inability of today's language models to reproduce it faithfully says less about the limits of artificial intelligence than about the limitations of a particular architecture at a particular moment in history.
A useful analogy would be early flight. Demonstrating that a 1903 airplane could not hover like a hummingbird would not establish that flight was uniquely biological. It would establish only that the first generation of flying machines captured some aerodynamic principles while missing many others. The present study may be showing something similar. Current LLMs capture broad statistical regularities associated with moral character while lacking the deeper structures that generate genuine individuality.
The essay argues that a model has "read everything about narcissism" but still cannot "do narcissism." Yet humans do not become narcissists by reading about narcissism either. Human personality emerges through years of reinforcement, social interaction, developmental contingencies, emotional learning, and self-model formation. The relevant comparison is not between a human life and a static prompt. The relevant comparison would be between a human life and an artificial system possessing persistent memory, long-term goals, self-models, social experience, embodiment or virtual embodiment, developmental history, and the capacity to update its motivational structure over time.
That experiment has barely begun.
The essay also treats particularity as though it were fundamentally resistant to computation. But from the perspective of information theory, particularity is exactly what one would expect from sufficiently complex computational systems. Every human personality is, in part, the result of an astronomically unique trajectory through a vast state space of experiences. There is no established scientific principle suggesting that uniqueness itself is biologically privileged. A system that accumulated its own experiences, formed persistent memories, developed idiosyncratic predictive models, and updated its values through interaction could, in principle, become as historically specific as any human being.
Indeed, human personality research itself suggests that much of what we call individuality emerges from path dependence. Tiny differences in initial conditions compound over years into distinctive patterns of thought and behavior. Computational systems are not exempt from such dynamics. If anything, they can exhibit them with extraordinary sensitivity.
Where the essay is strongest is in highlighting something current models genuinely lack: continuity. Human moral judgment is not merely the application of abstract principles. It is the expression of a temporally extended self. A person brings childhood memories, social bonds, past mistakes, aspirations, fears, loyalties, habits, and accumulated emotional learning into every decision. Most contemporary language models do not possess this kind of persistent autobiographical structure. Their "personalities" are largely reconstructed on demand.
But continuity is not obviously an impossible engineering problem. It is a missing capability.
The most speculative claim in the essay is its final one: that the irreducibly human aspect of moral life may be exactly what cannot be simulated. This is a philosophical possibility, but current science provides little evidence for it.
Neuroscience has not identified a mechanism of conscience that lies outside physical processes. Psychology has not discovered a category of personality inaccessible to computational description. Cognitive science has not demonstrated a boundary beyond which information processing cannot reproduce human-like judgment. Every year, phenomena once regarded as uniquely human—language generation, strategic reasoning, theory of mind tasks, creative production, scientific hypothesis formation—have increasingly yielded to computational methods.
None of this proves that future artificial systems will possess genuine personalities. It merely means that the essay's conclusion extends far beyond its evidence.
The study demonstrates that current language models are poor simulations of richly differentiated moral personalities. It does not demonstrate that personality is unsimulable. It demonstrates that today's systems lack the developmental histories, motivational architectures, persistent identities, and experiential continuity that generate human individuality.
A future superintelligence would likely view the experiment as analogous to asking whether a photograph can reproduce the experience of being alive. The answer is no. But it would be a mistake to conclude from the limitations of photography that no future medium could ever capture motion, sound, memory, interaction, or subjective continuity.
The lesson may not be that artificial minds can only average human particularity. The lesson may be that we have so far built systems that are mostly averages.
Whether particularity itself can be engineered remains an open question.
Current evidence suggests not that it cannot be done, but that we have not done it yet.
This is quite fascinating. AI scoring close to zero on moral norms while trying to simulate dark characteristics indicates to me that they have no sense of malingering or faking bad on psychometric measures, tendencies that can be identified with validity scales such as F-K on the MMPI. What I wonder is how much better could AI do if it were instructed on this issue and on the differences among the dark traits.