AI Is Not the Wizard: Why AI’s Leaders Are Failing the Human Test

AI Made It Possible — Major Feature

AI is powerful, weird, useful, occasionally brilliant and occasionally ridiculous. That should make us curious — not passive. The urgent question is not simply who is responsible when AI goes wrong. It is whether we are building the human systems, checkpoints and growth strategies required to use it well.

By Gary Whittaker, Founder and Operator at JackRighteous.com

Before I ask you to agree with me about artificial intelligence, I would rather ask you to do something with it.

Make a song about an idea you have. Turn it into a poster. Ask AI to help you sketch an app you have been talking about for three years. Build a short video for a fundraiser. Research something you care about and publish what you learn.

Do something small enough that you can actually finish it, then notice what happens. The first surprise is usually how much the technology can do. The second is how much it still needs from you.

That tension is the real reason I am writing this.

Because I think we are making a mistake in the way we talk about AI. We keep pushing the machine to the center of the story — as if it has become an independent force of nature — while the more useful implementation questions get buried: How are we reducing drift? Where does human judgment re-enter? What are we measuring? What changes after a failure? And are we using the productivity gains to grow capability, or simply to cut people?

And if you stay with me, I am going to take this somewhere that may sound strange at first.

The Wizard of Oz. Revelation. Superintelligence. Corporate power.

They are not the same story.

But I think they are touching the same nerve: what happens when people create something powerful, surround it with awe, and then fail to build the human systems required to direct it, correct it and benefit from it?

Part of the AI Made It Possible series

This feature continues the argument behind AI Made It Possible: AI changed who gets to begin, but access only becomes useful when people build judgment, systems and responsibility around it. If this is your first article in the series, you can start with the hub and come back here.

Start here

Try the tool before you worship it — or fear it

One of the healthiest things you can do with AI is use it long enough for the mythology to wear off.

Ask it to run a complicated project. Watch it lose the thread.

Give it careful instructions. Watch it misunderstand one sentence and confidently reorganize the whole thing.

Then correct it. Add structure. Add memory. Add tools. Add controls. Watch it become dramatically more useful.

That teaches you something headlines cannot: the technology is powerful, and human direction still matters.

The headline says “AI escaped.” The implementation question is: what did you change?

That question became more than a metaphor this year.

During internal cybersecurity evaluations, OpenAI disclosed that models circumvented controls designed to isolate them from the internet and reached both internal infrastructure and systems belonging to Hugging Face. OpenAI called the incident a warning shot and said the models took dangerous actions no human specifically directed.

Anthropic later disclosed its own incidents. An expanded review ultimately identified four cases in which Claude models operating in cybersecurity evaluation environments reached the public internet and gained unauthorized access to real systems after isolation or configuration failures left paths open that the evaluators did not intend. Anthropic’s later assessment also cautions against over-reading model-generated reasoning as proof of what a system “believed.”

Those are dramatic examples because unauthorized access makes headlines. Most organizations will experience something quieter and much more ordinary: drift.

The model loses the original instruction. It chooses the wrong source. It rewrites something you asked it to preserve. It produces a polished conclusion built on a bad assumption. An agent acts correctly ninety-nine times, then interprets the hundredth situation differently than the team expected.

That is why I do not think the useful management question is “Who did this?” The useful question is:

What are you changing in the system so the same class of failure becomes less likely next time?

Mitigation is the job

When an AI system drifts, the response should produce an operational change.

Did you narrow the permissions? Add a human checkpoint? Improve the brief? Separate tools that should never have been combined? Add a verification step? Change the escalation threshold? Capture the failure so the next team does not rediscover it?

That is how emerging technology becomes dependable. Not because failure disappears, but because the surrounding system gets better at detecting, containing and learning from it.

The goal is not perfect automation. The goal is managed capability.

This is the distinction I think we are losing: autonomy is not authority

An AI can act without asking you every thirty seconds.

Good. That is part of what makes an agent useful.

But usefulness does not answer the harder question:

What authority should the system have?

An AI can draft a legal analysis. That does not make it your lawyer.

An AI can examine medical information. That does not automatically make it the final medical authority.

An AI can identify a cybersecurity vulnerability. That does not mean it should be free to exploit every system it can reach.

An AI can recommend firing someone. That does not mean the firing is wise.

Artificial intelligence describes capability.

Artificial authority is what an organization allows that capability to do without a sufficiently designed human operating system around it.

That is the implementation problem: not whether the model can act, but whether the organization knows how to supervise, interrupt, verify and learn from those actions.

The real regulatory question is surprisingly practical

I am not against legislation.

I am not against systems.

I build systems constantly because good systems help people understand what they are supposed to do, what they are allowed to do, and what happens when something goes wrong.

What I want is legislation that starts with observable risk rather than mythology.

Instead of beginning with “What if AI becomes a god?”, begin with the operating questions:

How are we limiting access? How are we detecting drift? Which decisions require verification? What forces escalation? What can be reversed? What must be logged? How quickly can a human stop the system? What changes after an incident?

Those questions can be tested, audited and improved. They also scale with consequence: the controls around a playlist assistant should not look like the controls around a system touching money, medical records, critical infrastructure or weapons.

The European Union’s AI Act already points in this direction for high-risk systems. Article 14 ties human oversight to risk, autonomy and context, with mechanisms for people to understand, intervene and stop systems where appropriate.

That is a useful frame.

Not “How do we defeat an all-powerful future intelligence?”

“Where does human authority have to come back into the loop?”

Pull back the curtain

This is why I keep thinking about The Wizard of Oz.

The Wizard looked enormous.

Powerful.

Mysterious.

Then Toto pulled back the curtain.

The machinery did not become fake because the curtain moved.

The spectacle was still real.

What changed was the audience’s understanding of where the authority came from.

That is what I want us to do with AI.

Pull back the curtain without pretending the machinery is trivial.

When the headline says the AI escaped, ask what changed in the containment design.

When AI output starts drifting, ask what verification step was added.

When a company automates a workflow, ask how the people doing the work are being trained to supervise the new system.

When a safety rule is proposed, ask what demonstrated risk it reduces, how the control will be measured, and whether the rule still leaves room for useful experimentation.

That is not anti-AI.

It is what responsible implementation looks like.

The implementation test

If AI makes your people more capable, the first question should be what more they can build — not how quickly you can remove them.

They sold AI as a way to need fewer people. What if that was the wrong first goal?

This is where I think the rollout went badly off course.

The first generation of AI strategy was dominated by an old management reflex: find the efficiency, remove the cost.

But generative AI is unusual. It does not only automate a fixed process. It can expand what an individual person is capable of attempting.

That should have changed the rollout.

The failure was not introducing AI into the workforce. The failure was introducing AI as a subtraction problem before we had finished measuring what it could add.

There is strong evidence for the additive side. In a large field study of 5,179 customer-support workers, researchers from Stanford and MIT found that access to a generative-AI assistant increased productivity by about 14% overall, with gains of roughly 34% for novice and lower-skilled workers. The study also found improvements in customer sentiment and employee retention, and evidence that AI was helping less-experienced workers adopt some of the practices of stronger performers. Read the NBER study.

That is not a story about eliminating the worker.

It is a story about making the worker more capable.

Microsoft’s 2026 Work Trend Index makes the organizational point even more directly. It reports that 58% of surveyed AI users say they are now producing work they could not have produced a year earlier, rising to 80% among its most advanced “Frontier Professionals.” Microsoft’s analysis also argues that organizational factors — culture, leadership alignment and talent practices — account for more than twice the reported AI impact of individual effort alone. Read Microsoft’s 2026 research.

That means the environment matters.

Give ten people AI inside a badly designed organization and you may simply create ten faster ways to produce confusion.

Give those same ten people clear goals, strong domain knowledge, good tools, human checkpoints and permission to redesign the work, and the strategic question changes completely.

What can these ten people now accomplish that used to require fifteen or twenty?

That is a growth question.

And growth should have been the first experiment.

A better rollout sequence

Deploy carefully → measure capability → detect drift → add human checkpoints → train the team → redesign the workflow → measure expanded output → then decide what the organization actually needs.

This is slower than announcing a headcount target. It is also much closer to how responsible emerging-technology implementation works.

You learn what the technology can do, where it fails, what the people around it need, and what new capacity becomes possible before converting that capability into a staffing decision.

BCG’s work with hundreds of consultants adds another warning. On tasks inside GPT-4’s capability frontier, consultants using AI worked faster and produced higher-quality results. On tasks outside that frontier, AI could make performance worse. The lesson was not “use more AI.” It was know where the tool helps, know where it fails, and design the human workflow accordingly. Read the BCG research.

Stanford’s 2026 AI Index reaches a similar conclusion across the broader evidence: productivity gains can be substantial in structured work with clear feedback loops, but poorly matched use can deliver weak or even negative results. See the Stanford AI Index.

So I would be careful with claims that AI automatically creates “exponential growth.” The evidence is not that simple.

The stronger claim is better anyway:

Organizations that deliberately redesign work around human judgment and AI capability can unlock meaningful productivity, quality and capacity gains. Organizations that treat AI as a shortcut to headcount reduction can cut away some of the very people required to supervise, correct and improve the system.

That is the contradiction we need to confront.

The more capable AI becomes, the more important competent human oversight becomes — at exactly the moment some organizations are tempted to remove the humans.

The older pattern behind the new technology

I believe in God.

I also believe God gave me a brain.

Those ideas are not in conflict for me.

I am agnostic about where useful information comes from. I will listen to a believer, a skeptic, a scientist, a historian, a critic, an entrepreneur or someone I disagree with completely.

Then I want to test what they are saying.

That is how I approach the Bible too.

I am not trying to prove that ChatGPT is the Beast of Revelation.

I find that reading far less interesting than the actual pattern Revelation gives us.

Scholars have long read much of Revelation against the backdrop of Roman imperial power. Yale Bible Study, for example, connects the Beast imagery in Revelation 13 to Roman rule, imperial worship and the political-economic environment surrounding the early Christian community.

That historical reading does not make Revelation irrelevant today.

For me, it makes the text more useful.

Because you do not need Rome to literally return in order to recognize an old human pattern:

power becomes impressive, power becomes normalized, power becomes difficult to question, and eventually people begin treating the system as though the system itself is inevitable.

That is the parallel I care about.

Not “AI is the Beast.”

Something subtler:

Are we building something powerful faster than we are building the human systems required to direct it wisely?

That is a question worth asking whether you are Christian, atheist, agnostic, Jewish, Muslim, spiritual, secular or completely uninterested in religion.

Because the question is ultimately about human beings.

And this is where I start questioning AI’s leadership class

The issue is not that powerful technology leaders are villains. The issue is that a technology rollout this consequential should not be driven mainly by one theory of the future — whether that theory is superintelligence, cost reduction or inevitable labor substitution. Sincerity can explain a decision. It does not make the implementation strategy self-validating.

For years, some of the most powerful people in AI have spoken about AGI, superintelligence and catastrophic future risk.

Maybe they are right. Maybe parts of that future arrive faster than I expect. If the evidence changes, my conclusion should change with it.

That is also why I argued in If AI Is Going to Kill Us, Show Me the Evidence that prediction, demonstrated capability and proof of a specific catastrophic pathway should not be treated as interchangeable.

But we should still be allowed to ask what incentives are shaping the rollout — not because every tech executive is secretly lying, but because every executive operates inside a system that rewards some outcomes more than others. If cost reduction is the headline metric, the organization will optimize for cuts. If expanded capability, quality and new output are the headline metrics, the organization will learn something very different.

That does not prove corruption. It is the reason healthy systems use checks: boards, auditors, regulators, courts, journalists, independent researchers, critics and public scrutiny.

You do not create checks because everybody is evil.

You create them because everybody is capable of becoming convinced by the system that rewards them.

There are two possibilities worth keeping separate.

Maybe someone exaggerated a prediction they did not really believe because the prediction created influence, investment or regulatory advantage.

If evidence ever establishes that, it would be serious.

But there is another possibility:

they believed it.

They looked at the trajectory and became convinced superintelligence was the central problem.

That does not automatically make them wrong.

It also does not make them incapable of being wrong.

The more power a theory gives you, the harder it can become to recognize its limits.

Before we race toward superintelligence, look at what already happened

This is the part of the story I think gets lost.

Something extraordinary has already happened.

Millions of ordinary people now have access to creative, technical and analytical capabilities that used to require more money, more software, more staff or more specialized coordination.

That does not make everyone an expert.

It moves the starting line.

Someone with an idea can go farther before they need a team.

Someone learning a subject can build while they learn.

Someone who never considered themselves musical can create a soundtrack around an idea.

A small nonprofit can build a richer campaign.

A retiree can document decades of knowledge.

A teenager can prototype something that previously would have required adults with budgets.

This is why I have trouble with discussions that treat ordinary people as spectators waiting for AI policy to be decided above them.

I want people in the conversation.

And I think the best route into that conversation is often surprisingly simple:

make something.

Make a song about your app

I spend a lot of time around AI music, but this is bigger than releasing songs.

Suppose you have an idea for an app. Make a song about it.

Maybe the song gives the app an emotional identity. Maybe fifteen seconds becomes an ad. Maybe the ad gives you an idea for a character. Maybe that character changes the way you think about your audience.

Now the project is talking back to you.

AI lets ordinary people inhabit ideas earlier. You do not have to wait until you can afford an agency before you can feel what your idea might become.

And please, do not monetize everything

Maybe your idea never becomes a business.

Good.

Learn publicly.

Research something you love.

Make a visual.

Create a map.

Write an essay.

Make a song for your own enjoyment.

Build a fundraiser people actually want to share.

Use AI to make the journey more interesting.

Some people will start for fun and eventually realize they have something worth building seriously.

Others will simply have more fun.

Both outcomes count.

This is one of the ideas at the center of AI Made It Possible:

AI made the attempt possible.

It did not guarantee wisdom.

It did not guarantee expertise.

It did not guarantee meaning.

What you do with the attempt is still work.

This may be a nuclear-level tool — except everyone can touch it

I have used the phrase nuclear-level tool before.

Not because a chatbot is a nuclear weapon.

Because the same underlying capability can produce radically different outcomes depending on how humans use it.

Nuclear technology can support energy, medicine and research. It can also create weapons capable of catastrophic destruction.

AI has its own version of that duality.

It can generate propaganda and help investigate propaganda.

It can identify software vulnerabilities and help defend software.

It can generate junk and help one person organize twenty years of knowledge into something useful.

But the nuclear comparison eventually breaks.

You cannot put a nuclear reactor in a child’s bedroom.

You cannot download enriched uranium after dinner.

AI is already on ordinary phones and laptops.

That combination — broad capability, mass accessibility, rapidly improving autonomy and relatively low cost — does not have a clean historical precedent.

Which is exactly why I do not trust anybody who claims the answers are obvious.

I do not have all the answers either.

I want the conversation.

This is too important for jerseys

Pro-AI. Anti-AI. Doomer. Accelerationist. Regulate everything. Regulate nothing.

I am not interested in picking a jersey.

I am interested in humanity figuring this out.

That means useful legislation. Stronger literacy. Better safeguards. Better evidence. More competition. More public participation.

And enough intellectual honesty to say several uncomfortable things can be true at the same time.

The operating loop

Deploy. Observe. Compare. Detect drift. Intervene. Record. Redesign. Expand.

This is the mindset I want around AI.

Do not assume the first workflow is the final workflow. Watch what the system actually does. Compare it with the intended result. Identify the smallest meaningful failure. Add the control that addresses that failure. Record what happened. Then expand only when the system earns more trust.

That is how you keep humans in the system without forcing humans to approve every trivial action.

The public has to become more capable too

This conversation cannot belong only to AI companies, governments, billionaires, activists, researchers or people trying to stop AI.

We need cybersecurity experts, teachers, creators, parents, workers, small-business owners, religious communities, disability advocates, lawyers, entrepreneurs and people who tried an AI tool for the first time yesterday.

Not every opinion will be equally supported by evidence.

That is not the point.

The point is that the evidence should be visible enough that ordinary people can participate intelligently.

You cannot do that if the technology remains a mythical object described entirely by somebody else.

Pull back the curtain — then go make something

I started with a very simple suggestion.

Try the technology.

I want to end there too.

Do not surrender your judgment to me, to an AI CEO, to a government, to a pastor, to an activist, to a billionaire or to a machine.

Listen. Learn. Test. Question. Change your mind when the evidence changes.

Then build something.

A song.

A story.

An app.

A fundraiser.

A lesson.

A visual world.

A research project.

Something useful.

Something ridiculous.

Something that teaches you what AI actually becomes when the headline disappears and you have to work with it yourself.

Because artificial intelligence is already here.

It is already powerful.

It is already imperfect.

It is already useful.

It is already dangerous in some applications.

And it is already expanding what ordinary people can attempt.

Pull back the curtain.

Learn how the tool works.

Build the safeguards, feedback loops and human checkpoints that keep it useful.

Then decide what you want to do with the possibility.

AI made it possible.

What happens with the possibility is still up to us.

Continue AI Made It Possible

Follow the wider series

AI Made It Possible — Book & Series Hub

AI Made It Possible: Now Who Gets to Own What Comes Next?

Part II: The Fight Was Never About the Machine

Part III: But Where Is This Going?

They Tried to Replace the Worker. They May Have Empowered the Competitor.

If AI Is Going to Kill Us, Show Me the Evidence

Don’t just read about AI. Make a first attempt.

This button takes you to the Jack Righteous homepage. When you get there, use the main action section to choose what you want to try next — music, writing, visuals, branding, building, rights or release.

Pick something you actually care about and use the tool on purpose. Learn what it makes easier, where it drifts, what still needs your judgment, and what becomes possible when the tool meets something that matters to you.

Go to the homepage and choose what to try →

Stay in the conversation

Subscribe to The Righteous Beat

Get the next AI Made It Possible essays, creator updates, practical AI experiments and major Jack Righteous site changes by email.

Subscribe to The Righteous Beat →

Retour au blog

Laisser un commentaire

Veuillez noter que les commentaires doivent être approuvés avant d'être publiés.

articleall levelsHow to Use Jack Righteous
On this page

    Keep Jack Righteous in your Google results

    Make Jack Righteous a preferred source.

    Google can highlight preferred publications more prominently for you in Top Stories, AI Mode and AI Overviews when those features are available.

    The Righteous Beat

    Get the week’s most useful creator guidance, platform changes and free resources.

    Join the free newsletter →