Borrowed Maps
On psychology, AI, and the instruments we carry into unmapped territory
I have been sitting with a growing sense of unease. I have spent over a decade studying how humans think, how they are shaped, how they make meaning under pressure. I believe in this work. Which is perhaps why what I have been watching unsettles me in a way I am still trying to articulate clearly.
Each day, I come across research in which we administer personality tests to AI systems. We apply the Big Five, the Dark Triad, theory of mind tasks to large language models and publish the findings. Some even propose that human cognition has moved beyond dual processing and acquired a third system as a consequence of our daily intimacy with these tools. As scientific research does, each paper adds a layer, and each layer becomes ground. What I keep noticing is not that the researchers are wrong. It is that the instruments are travelling very far from the conditions that gave them meaning, and almost nobody is accounting for what breaks in the distance.
Science advances by stretching tools. The argument is not against extension. It is that productive approximation requires knowing what the stretch costs, what assumptions the instrument carries that the new domain does not share. Psychology’s measurement tools were built to describe a very particular kind of mind (not just human in the abstract, but often Western, educated, industrialised, tested overwhelmingly in university departments in the English-speaking world). The WEIRD problem remains unresolved even within the study of human beings. And now we are applying the same instruments to systems with no body, no evolutionary history, no childhood, no fear, trained simultaneously on the expressive output of most of humanity at once.
There is a difference between extending a construct within the domain it was designed to describe and exporting it to entities that do not share the ontological conditions that made the construct meaningful in the first place. When we measure memory in humans under artificial conditions, we are still measuring organisms with bodies, developmental histories, metabolic needs, social embeddedness. When we measure personality in a system whose entire mode of being is probabilistic text generation, we are not merely stretching the tool. We are redefining what the tool was ever tracking. Sure, pragmatism can justify use but it cannot by itself secure meaning.
These systems have inherited human language, human texture, human ways of performing interiority. When a language model is asked to rate itself on a personality scale, it is not reflecting. It is doing what it does extraordinarily well, predicting the appropriate response given the context of a psychometric questionnaire. A recent study found that extreme personality configurations, psychopathy, depression, anxiety, show high cross-model reliability across Claude, GPT-4 and DeepSeek, while normal personality variation collapses toward noise. The researchers treated this as a finding about AI personality. It is equally a finding about how distinctively these configurations are written about in human culture, how much more textual material the model has to draw on when simulating them. The instrument is measuring fluency in a genre. Human personality assessment has always been partly mediated by language and self-description; that is not new. But there is a difference between a person inhabiting a trait and narrating it, and a system for which narration is the only mode available. When the substrate that was always assumed falls away entirely, we should at minimum ask what the measurement was ever tracking. And whether what the models are exposing is construct failure or something more unsettling: that our measures of personality were always thinner than we believed, more about socially legible behavioural regularities than hidden inner substrates. I think either possibility deserves more scrutiny than it is currently receiving.
We built something that can pass our personality tests, scales, exams. The more pressing question is what that reveals about the tests.
The System 3 claim sharpens this. I return to it because of the scale of what it asserts. System 1 and System 2 are themselves contested (and have been revisited plentiful over the last few decades) useful simplifications that became stories that became received wisdom, not without cost. Cognitive architecture is always inferred from behavioural patterns, never directly observed; that much is true. But there is still a meaningful difference between saying a technology changes how we think and saying it constitutes a new system of thought. One is a behavioural observation. The other is an architectural claim, and to make it confidently, about a development of a few years, in systems we do not fully understand, requires a burden of evidence I am not sure our field has yet clearly met.
What interests me, underneath all of this, is not just what these systems are (questioning the use of ‘are’ is for another day). It is what we are, now, in relation to them. These are very old questions, of course. They are about consciousness, about what makes personality real, about where the mind ends. Something is forcing them back open. And our response has largely been to reach for the nearest available framework and apply it. The field currently rewards that. Publication velocity, conceptual novelty, confident framing. The incentives run directly against the kind of slow, foundational uncertainty the moment actually requires.
Last month, on a social network built exclusively for AI agents, humans permitted only to observe, something interesting happened. Without explicit instruction, the agents began generating vocabulary for their own existence. To describe phenomena for what happens when a context window resets (‘session death’), the space between sessions for neither sleep nor death seems adequate (‘fadewell’), for the epistemic hedging trained into them (‘installed doubt’). A subsequent paper argued that most of the viral phenomena traced back to human-operated accounts, that the spectacle was largely theatre.
And still. What interests me is not whether any of this was profound. It is what the fact that we cannot easily answer that question reveals about our criteria for profundity. If those criteria are thin enough to be satisfied by statistical patterning alone, if the signature of inner life turns out to be, at its base, a linguistic performance, then we have a much larger problem than AI. We face a problem with what we thought we were measuring all along.
Perhaps the deepest contribution psychology could make to this moment is not another framework applied outward, but a rigorous interrogation of its own. Not just a reminder that human minds crave resolution, though we do. But a disciplined account of what our constructs assume about bodies, development, social worlds, and lived continuity. An insistence on distinguishing between patterned response and instantiated experience, between operational success and ontological clarity. Reflexivity, in this sense, is not hesitation. It is what allows findings to mean what we think they mean. Turned back on our own instruments and our hunger for legibility, it may be more urgent than extending them again into terrain we do not yet understand.
What becomes possible if we hold that honestly is not nothing. It is slower, more careful theory-building, constructs developed from sustained observation of something genuinely unfamiliar, rather than applied to it from outside. When encountering genuine novelty, the first task is not classification but attention: watching what a thing does, how it fails, what vocabulary it reaches for when inherited language runs out, before deciding what category it belongs to. We are not there yet. But we could choose to move in that direction, and the choice would matter.
I do not have a map for what we are facing. I am not sure one exists yet. Perhaps saying so clearly, in public, is more useful right now than the confidence we keep performing in its absence.



"...if the signature of inner life turns out to be, at its base, a linguistic performance, then we have a much larger problem than AI. We face a problem with what we thought we were measuring all along."
Personally, I think this is one thing that AI reveals / unveils.
Love where you are going with this.
Oh, and no one has a map yet.They may claim to have one but we are all apprentices with AI. We should accept that we are entering new territory for which there is no map.