Superintelligence: How Would We Actually Know?
Share
AI Made It Possible · Research Notes
Superintelligence: How Would We Actually Know?
If an AI becomes better than humans at everything we can test, what exactly have we proven?
We may have proven extraordinary capability. We may have proven broad generalization. We may even have proven that the system can generate discoveries beyond what any individual human can produce. But none of those results, by themselves, prove that the system is conscious, sentient, self-aware, or having an inner experience.
FAST ANSWER
Superintelligence is primarily a capability claim, not a consciousness claim. In current AI research, intelligence, generality, autonomy, agency and consciousness are related questions, but they are not interchangeable. A system could become superhuman across broad cognitive tasks without science having established that anything is being subjectively experienced.
Before we say “AI”
AI is not one thing
For many people, “AI” still means a chatbot. That is understandable because ChatGPT, Gemini, Claude and similar tools are the most visible part of the current wave.
But AI is a much larger category. NIST uses several definitions across different technical contexts, including systems that make predictions, recommendations or decisions, systems that learn from data, and systems that perform tasks associated with perception, reasoning, planning, communication or action. The OECD likewise defines an AI system broadly as a machine-based system that infers from inputs how to generate outputs such as predictions, content, recommendations or decisions that can affect physical or virtual environments.
That means an AI discussion can be about very different things: a language model generating text; software recommending who sees an advertisement; computer vision identifying an object; an agent using tools to complete a task; an autonomous vehicle interpreting its surroundings; a fraud system flagging a transaction; or AI operating quietly inside a much larger product.
Those systems do not necessarily have the same capabilities, risks, autonomy or relationship to a human user.
So throughout this series, one of my rules is simple: when somebody says “AI,” I want to know which AI system, doing what, with what authority, using what evidence?
Start with the vocabulary
If “AI” still feels like one giant category, start with AI Is Not One Thing: What Are We Actually Talking About? → It separates models, generative AI, agents, decision systems, robotics, general-purpose AI, AGI and ASI before this article asks how we would prove superintelligence.
There is no single official definition of superintelligence
That matters before we go any further. AI terminology is still developing, and even standards organizations caution that terms can have multiple definitions depending on context. NIST explicitly notes that its glossary should not be treated as a single universal source of preferred definitions across every AI domain.
The classic academic definition comes from philosopher Nick Bostrom: superintelligence describes an intellect that greatly exceeds the best human minds across practically every important intellectual domain.
More recent AI research is trying to operationalize that idea rather than simply repeat it. Google DeepMind's Levels of AGI framework separates performance, generality and autonomy. Its 2026 report From AGI to ASI goes further and describes artificial general superintelligence intuitively as intelligence and cognitive capability beyond even large organizations of humans.
THE TERMINOLOGY THAT HELPS
| Question | Useful term | What it does not automatically prove |
|---|---|---|
| Can it outperform humans on a specific task? | Superhuman narrow capability | Generality or consciousness |
| Can it perform broadly across many domains? | AGI / general intelligence | Sentience |
| Can it substantially exceed top humans across broad domains? | ASI / superintelligence | Consciousness |
| Can it pursue goals with less direct supervision? | Agency / autonomy | Inner experience |
| Does anything actually feel like something to the system? | Sentience / phenomenal consciousness | That remains a separate scientific problem |
Sentience is not simply the level after superintelligence
Popular culture often presents a ladder that looks like this: AI, then AGI, then superintelligence, then sentience. That is not a scientifically established sequence.
A machine could theoretically become enormously capable while having no subjective experience at all. Conversely, if artificial consciousness is possible, there is no requirement that the first conscious artificial system must also be superintelligent.
That is why current consciousness research does not simply ask how high a model scores. Researchers look instead at theories of consciousness and ask whether a system exhibits properties predicted by those theories.
The clone problem
This is the part I keep coming back to.
Modern AI is trained on enormous amounts of human-created material. It learns patterns in our language, our explanations, our arguments, our mistakes, our styles, our emotional expressions and our descriptions of inner life.
So imagine an AI that becomes so good at modeling us that we can no longer reliably distinguish its output from the output of a deeply thoughtful human.
It says, “I feel afraid.” It explains grief. It creates a philosophy of its own existence. It appears to recognize itself. It tells us it wants to survive.
That would be remarkable behavior. It still would not, by itself, prove subjective experience.
The system might have learned an extremely powerful model of what beings with inner experience say and do. The imitation could become so complete that human observers can no longer detect the difference.
The difficult question is no longer, “Can it imitate us?” It becomes, “What evidence would distinguish imitation from instantiation?”
High benchmark scores do not settle that question either
This is where measuring intelligence becomes more complicated than simply watching a leaderboard.
François Chollet's work on measuring intelligence argues that skill on a task is not the same thing as intelligence because skill can be purchased with enough prior knowledge, training data and exposure. His proposed alternative emphasizes skill-acquisition efficiency and the ability to generalize to genuinely new problems.
That concern has only become more important as models have grown. The ARC-AGI research program focuses on few-shot generalization on novel tasks precisely because high performance on familiar task families can hide dependence on prior knowledge. Its 2025 technical report also warns about newer forms of benchmark contamination tied to knowledge coverage.
Humanity's Last Exam was created because older academic benchmarks were saturating. Even there, the researchers are careful about what the benchmark can establish: strong performance would demonstrate advanced capability on difficult, closed-ended academic questions. It would not, by itself, prove AGI, ASI or consciousness.
So what would count as stronger evidence of superintelligence?
If the claim is that a system is genuinely superintelligent, I would want the evidence to get harder as the claim gets bigger.
A stronger empirical standard
1. Data-sealed evaluation. Important tests should be created or securely held after the model's training cutoff so the system cannot simply have encountered the answer during training.
2. Novel-task transfer. The system should solve problems whose structure differs meaningfully from known training examples, not just recombine familiar templates.
3. Breadth. The evidence should span scientific reasoning, engineering, planning, language, mathematics, social reasoning, creativity and other consequential domains rather than relying on one spectacular benchmark.
4. Frontier-human comparison. The relevant baseline should eventually be elite experts and expert teams, not the average person.
5. Real-world verification. Claims become much stronger when the AI produces results that independent humans can test outside the benchmark: a validated scientific discovery, a proof accepted by experts, a new engineering solution that works, or accurate predictions humans were unable to make.
6. Repeated generalization. One breakthrough could be luck, leakage or a narrow trick. The pattern must persist across new situations.
7. Causal understanding of the system. Where possible, researchers should test whether internal mechanisms actually support the claimed abilities rather than relying only on impressive outputs.
But even all of that would not prove consciousness
Imagine the strongest case possible. An AI solves unsolved mathematics, discovers a new drug target, designs an experiment, predicts the result, explains why it works and then transfers the underlying principle to another scientific domain.
That could become extraordinary evidence of general intelligence and perhaps eventually superintelligence.
But we would still be asking a different question when we ask whether the system experienced anything while doing it.
Consciousness is private. Even in humans, we infer another person's experience through behavior, shared biology, neural mechanisms and other evidence. AI removes much of that biological common ground.
What consciousness researchers are actually looking for
A major interdisciplinary report led by Patrick Butlin and Robert Long proposed evaluating AI systems against computational indicators derived from leading scientific theories of consciousness, including recurrent processing, global workspace theory, higher-order theories, predictive processing and attention schema theory.
That approach is very different from asking a chatbot whether it feels alive. It asks whether the system has mechanisms that scientific theories predict should matter for consciousness.
The authors concluded that the systems they evaluated did not provide sufficient evidence of consciousness, while also arguing that there were no obvious technical barriers to creating systems that satisfied more of the proposed indicators.
Even that would still be evidence, not a universally accepted proof. Consciousness science itself remains contested.
A human consequence worth separating from the hype
People can react to simulated personhood even when personhood has not been proven.
This distinction is not merely philosophical. The American Psychological Association reported in September 2026 on growing cases popularly described as “AI psychosis.” The term is not a formal clinical diagnosis, and experts caution against assuming a simple one-way cause. But clinicians have reported cases in which highly interactive chatbot conversations appear to reinforce delusional beliefs or blur reality for vulnerable users.
That matters to this article for a very specific reason: a system does not have to be conscious for a human being to experience it as conscious.
The social effect of simulated personhood can therefore arrive before science has settled the question of machine personhood itself.
Simulation versus instantiation may become one of the central arguments
There is already an active philosophical and scientific dispute over whether the right computation is enough to create conscious experience.
One recent 2026 argument associated with Google DeepMind research explicitly distinguishes simulation from instantiation, arguing that reproducing the functional pattern associated with consciousness does not necessarily mean the underlying physical system possesses subjective experience. That is one position in a larger unresolved debate, not scientific consensus.
The analogy is imperfect, but useful: a computer can simulate a hurricane without producing wind in the room. The unresolved question is whether consciousness behaves more like an abstract computation or more like a physical phenomenon that requires particular underlying properties.
THE LINE I WANT TO KEEP CLEAR
Performance can prove capability. It cannot automatically prove personhood.
AI may become able to do things humans cannot do. That does not require us to pretend the accomplishment is fake simply because the system was trained on human data.
But extraordinary capability also does not give us permission to quietly smuggle in another conclusion: that the system therefore possesses an inner life.
Those are two different claims, and they need two different bodies of evidence.
Why I am documenting this
This research is feeding both the book and the world I am building.
I am working on the second edition of AI Made It Possible while also building the Jack Righteous universe. That universe is not meant to predict the future literally. It is a pseudo-parallel world that takes documented developments happening now and asks what they could look like when carried into the 2030 setting of War Comes.
That is why I care about getting the terminology right before turning it into story.
AI is an extraordinary tool. It can put capabilities that once required money, teams or specialist access into the hands of ordinary creators and working people. That is the part of this technology I find genuinely exciting.
At the same time, the infrastructure, financing, policy, corporate control and human consequences surrounding AI are also real subjects of investigation. I do not need a secret conspiracy to make that interesting. Public statements, financial filings, legislation, court records, technical papers and documented incidents give us enough to examine.
The discipline is to separate three layers: what is documented; what experts reasonably infer from that evidence; and what I explore through fiction.
That separation matters because the closer a fictional world sits to reality, the more important it becomes to show the reader where the factual trail ends and the imagined future begins.
Where I think the next question starts
This leaves us with something more interesting than either “AI is just copying us” or “AI has become alive.”
What if the first artificial superintelligence is neither a simple copy nor a conscious new being?
What if it is a new kind of cognitive system: trained on us, able to generalize beyond us, capable of creating results we could not create ourselves, while remaining phenomenologically empty?
And if that happens, how would we know which part of what we are seeing came from memorization, which part came from generalization, which part came from new machine-native strategies—and whether any of it was accompanied by experience?
That is the question I want to keep researching.
Continue the series
The AI Industry Has Proven the Technology. It Has Not Proven Who Is Behind the Curtain. →
Beginner's Terminology Guide
The words in plain English
You do not need a computer-science degree to follow this conversation. These are the main terms used in this article, translated into everyday language.
AI — Artificial Intelligence
Software designed to perform tasks that normally require human-like abilities such as language, pattern recognition, planning, prediction or problem-solving.
AGI — Artificial General Intelligence
A still-debated term for AI that can perform well across a very wide range of intellectual tasks rather than being strong in only one narrow area.
ASI — Artificial Superintelligence
AI whose broad cognitive abilities substantially exceed those of even the most capable humans or human organizations across many important domains.
Capability
What the system can actually do. For example: solve a hard equation, write code, plan a project or discover a scientific relationship.
Generality
How widely a system can use its abilities. A calculator is powerful but narrow. A more general system can work across many different kinds of problems.
Generalization
Using what was learned in one situation to solve a new situation that was not simply memorized. This matters because a system can score well on familiar tasks without necessarily showing flexible intelligence.
Autonomy
How much a system can continue acting without a human giving it every next instruction. More autonomy does not automatically mean more intelligence or consciousness.
Agency
The ability of a system to pursue a goal through a sequence of actions. In AI, this can be engineered behavior. It does not automatically mean the system has wants, feelings or a personal will.
Sentience
The capacity to have subjective experiences such as pain, pleasure or sensation—to have something actually feel like something from the inside.
Consciousness
A broader and heavily debated term for having subjective awareness or experience. Researchers disagree about its exact boundaries and how it should be measured.
Self-awareness
Recognizing or representing oneself as a distinct entity. A system can talk about itself without that proving human-like self-awareness.
Benchmark
A standardized test used to compare AI systems. Think of it as an exam for models. The problem is that doing well on the exam may not always tell us why the system did well.
Benchmark contamination
When a model may already have seen the test questions, answers or very similar material during training. That can make a score look more impressive than the underlying ability really is.
Training data
The information used to help an AI learn patterns. For language models, this can include huge amounts of text and other human-created material.
Memorization
Producing an answer because something very similar was already encountered during training, rather than solving the problem from scratch.
Interpolation
Creating a new-looking answer by combining or filling gaps between patterns the system already knows. It can be useful and creative-looking without necessarily being a completely new kind of reasoning.
Novelty
Something genuinely new relative to what the system was trained on or previously shown. Proving true novelty is difficult because modern training sets are enormous.
Simulation
Reproducing the appearance or behavior of something. An AI could simulate fear very convincingly—for example, by saying the things frightened humans say—without that alone proving it actually feels fear.
Instantiation
Actually possessing or physically realizing the thing being discussed, rather than only imitating its outward pattern. In this debate: does the system merely simulate consciousness, or does consciousness actually exist in the system?
Phenomenal consciousness
A technical way of talking about subjective experience itself: the idea that there is something it is like to see red, feel pain, hear music or experience fear.
Mechanistic / causal evidence
Evidence based on what is actually happening inside a system and what causes its behavior, rather than judging only by the answer it gives us.
The simplest rule: sounding human, acting independently, being highly intelligent and being conscious are four different claims. One does not automatically prove the others.
Sources and further reading
- Google DeepMind — Levels of AGI for Operationalizing Progress on the Path to AGI
- Google DeepMind — From AGI to ASI (2026)
- François Chollet — On the Measure of Intelligence
- ARC Prize 2025 Technical Report
- Nature — Humanity's Last Exam benchmark paper
- Butlin, Long et al. — Consciousness in Artificial Intelligence
- Nature Reviews Neuroscience — Covert measures of consciousness
- Google DeepMind — The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness
- NIST Glossary — terminology caveats and source definitions
- OECD AI Principles — definition of an AI system
- American Psychological Association — Understanding “AI psychosis”
Working research note
This article is intentionally a living part of the AI Made It Possible research process. It establishes the terminology and evidence standards. Future revisions will go deeper into what counts as novelty, how we detect training-data dependence, machine-native reasoning, embodiment, self-models and competing theories of consciousness.
Research review: September 23, 2026. This article distinguishes demonstrated capability from unresolved claims about consciousness and sentience.
AI Made It Possible · Main Investigation
This article is one branch of the larger evidence map. Start with: AI Could End Humanity Within a Decade. So What Are We Doing About It? →
Evidence before the label
Before calling a system superintelligent, we need to separate memorization, retrieval and benchmark leakage from real generalization. Read the novelty and generalization test →
AI Made It Possible · Research Spine
Where this article sits: Core 3 of 8 · Capability threshold — asks what evidence would justify the label superintelligence.
Supporting investigations
The AI Panic Machine · Will AI Take Your Job? · If AI Is Going to Kill Us, Show Me the Evidence · The AI Boom’s Missing Economic Breakthrough
The AI Race Has a Safety Problem · why capability thresholds become a geopolitical incentive problem when major powers do not want to slow down first.