Bee Righteous explains the five AI layers: apps, models, data centres, chips and energy.

Inside the AI Factory: The Five Layers Behind Every AI Prompt

Gary Whittaker

Technology profile · AI infrastructure and industrial power · Published July 29, 2026

Direct answer: Every AI prompt depends on five connected layers: energy, chips, infrastructure, models and applications. The application is what the public sees. The lower layers determine who can build, operate, price and control AI at scale.

Type a sentence into an AI tool and the response can feel almost weightless. A prompt goes in. An answer, image, song, video or piece of code comes out. The machinery disappears behind the interface.

But artificial intelligence is not weightless.

Every output begins in the physical world—with electricity, advanced processors, data-centre buildings, cooling systems, fibre networks, storage, trained models, human labour and enormous amounts of capital. The prompt box is not the AI economy. It is the customer-facing door at the top of it.

A useful way to understand the system is as a five-layer factory:

  1. Energy keeps the system operating.
  2. Chips perform the calculations.
  3. Infrastructure houses, connects and cools the machinery.
  4. Models turn compute into capability.
  5. Applications package that capability into products people can use.

The framework is simple. Its implications are not.

Most public attention goes to the top layer because applications are where people experience AI. Yet some of the hardest bottlenecks—and some of the greatest concentrations of power—sit much further down. A user can switch from one chatbot or music generator to another in minutes. A country cannot replace a power grid, semiconductor supply chain or hyperscale data-centre network nearly as easily.

This is the story of the AI factory: what each layer does, who controls it, what the popular diagram leaves out and why it matters to independent creators who may never own a GPU cluster but still depend on the decisions made beneath their tools.

What does “AI compute” actually mean?

AI compute is the processing capacity used to train artificial-intelligence models and run them after they are deployed.

A computer performs calculations. An advanced AI system performs enormous numbers of calculations across large groups of specialized processors. The faster and more efficiently those processors can work together, the more ambitious the model or service can become.

There are two major kinds of AI compute that readers should keep separate.

Training compute

Training is the process used to build or substantially improve a model. It can require large datasets, clusters of processors, repeated experiments and long periods of testing and adjustment.

Training is the expensive event most often associated with the race to build larger or more capable models. But it is not the end of the cost.

Inference compute

Inference is the processing used when the trained model performs a task. Every generated answer, image, song, translation, summary or video requires inference.

Training builds the engine. Inference keeps it running every time someone turns the key.

That distinction matters because the AI industry does not need infrastructure only when a new model is created. It needs continuing capacity whenever millions of people and businesses use that model.

The five-layer AI factory

The five-layer model explains the dependency chain from the physical base to the product in the user’s hand.

Layer Primary function What it produces
Applications Interface, workflow and distribution User experience
Models Pattern recognition and generation AI capability
Infrastructure Housing, networking, storage and cooling Operational capacity
Chips Mathematical processing Compute
Energy Continuous electrical supply Operation

The bottom three layers—energy, chips and infrastructure—are often described as the physical AI factory. Resources enter. Machinery operates. Processing takes place. A usable service emerges.

Models do not create computing power. They consume it. Applications do not eliminate the industrial system. They hide it behind an interface.

That is the first major correction to the way AI is commonly discussed. AI may arrive through software, but it is being built like heavy industry.

Layer one: Energy is the foundation

Energy is the first layer because every processor, server, network, storage device and cooling system depends on continuous electricity.

Data centres do not simply need large amounts of electricity. They need dependable power that can support dense computing loads with minimal interruption. They also require backup systems, transmission capacity, substations, transformers and long-term supply planning.

The International Energy Agency reported that global data-centre electricity demand grew by 17% in 2025, while electricity consumption from AI-focused data centres grew by about 50%. Its central projection has global data-centre electricity use reaching roughly 950 terawatt-hours by 2030, almost double the 2025 level. These are projections, not guarantees, but they show why energy has moved from a background cost to a central strategic concern.

The energy layer includes utilities, generators, transmission operators, grid planners, regulators, governments and the companies buying the power. It also creates a difficult public question: when a private data-centre project requires new generation, transmission lines or grid upgrades, who pays?

The answer will not be the same everywhere. Some projects may finance dedicated infrastructure. Some costs may be shared through utility systems. Some governments may offer incentives. Some communities may receive jobs and tax revenue. Others may carry new risks without receiving a proportionate benefit.

This is why the energy question cannot be reduced to whether AI uses “a lot” of electricity. The more precise questions are:

  • How much power has actually been requested?
  • How much has been contracted or delivered?
  • What new infrastructure is required?
  • Who finances it?
  • What happens if the projected demand does not materialize?
  • What competing public or industrial uses may be affected?

I examine those questions more directly in Who Pays to Power the AI Race?.

The first limit on artificial intelligence may not be intelligence. It may be the ability to deliver enough reliable power, quickly enough, without transferring an unfair share of the cost to the public.

Layer two: Chips manufacture compute

AI chips are the machines that perform the calculations required to train and operate models.

Standard processors can perform many kinds of computing tasks. Graphics processing units and other AI accelerators are especially effective at handling many mathematical operations in parallel. That makes them valuable for the workloads behind modern generative AI.

But the chip layer is not one company and it is not one component.

The supply chain includes chip designers, semiconductor fabrication plants, advanced packaging, high-bandwidth memory, lithography equipment, networking systems, substrates, materials, testing and specialized software. A shortage or delay in one part can constrain the system around it.

NVIDIA is the most visible symbol of this layer because its processors, networking and software ecosystem have become central to large-scale AI deployments. Its financial results show the commercial force behind that position: NVIDIA reported US$75.2 billion in Data Center revenue for the quarter ending April 26, 2026, up 92% from a year earlier.

That number is evidence of extraordinary demand for accelerated computing. It is not proof that every AI investment built on top of those purchases will earn an adequate return.

Manufacturing is another source of concentration. TSMC reported robust AI-related demand throughout 2025 and continued investment in leading-edge processes, advanced packaging and chip-stacking technologies. That illustrates why “who designs the best chip?” is only one part of the question. The ability to manufacture and package advanced processors at scale is a separate strategic advantage.

Chip access can therefore determine who is able to compete before a model is ever trained. A company may have talented researchers, customer demand and a compelling idea, but it cannot scale advanced AI without enough suitable processors and the infrastructure to use them.

This is also why semiconductor policy has become intertwined with trade, national security, manufacturing subsidies and export restrictions. AI capability is partly software. It is also access to some of the most complex manufactured products in the world.

Layer three: Where the AI factory becomes physical

Infrastructure is the system that houses, connects, powers, cools and protects the chips.

A data centre is not simply a warehouse filled with computers. It is a controlled operating environment built around servers, storage, high-speed networking, fibre, electrical equipment, cooling systems, backup power, security and maintenance.

AI processors are connected into large clusters so they can work together. The value of the cluster depends not only on how many chips it contains, but also on how efficiently workloads are distributed, how quickly data moves between processors, how often equipment fails and how effectively heat is removed.

The infrastructure layer brings together hyperscale cloud companies, specialist data-centre operators, property developers, construction firms, engineering companies, network providers, utilities, investors and governments.

It is also the point where private AI ambition meets public systems.

A technology company cannot simply summon a high-capacity grid connection, a new transmission line, a zoning approval, a water allocation or a skilled construction workforce. Those resources depend on institutions, communities and long-term planning that existed before the latest AI model was announced.

That is why the race to control AI begins with the data centre. It is the location where electricity, processors, land, fibre, cooling, capital and government permission must converge.

Cooling is not a minor detail

Advanced processors consume electricity and release heat. If that heat is not removed effectively, equipment performance and reliability suffer.

Facilities may use air cooling, liquid cooling, evaporative systems or combinations of several methods. The amount and type of water use can vary significantly by design, climate, electricity source and operating conditions.

This is where public discussion often becomes too loose. “AI uses water” is directionally true but analytically incomplete. The serious questions are how much, where, through which cooling system, during what period and whether a figure measures withdrawal, consumption or indirect water associated with electricity generation.

Not every data centre has the same water profile. Not every estimate applies universally. Accuracy requires facility-level and region-specific evidence whenever possible.

Layer four: Models turn compute into capability

An AI model is a trained mathematical system that identifies and reproduces patterns in data.

Depending on its design, a model may generate or interpret language, images, music, video, code, scientific information or business predictions.

At a high level, model development involves gathering or licensing data, preparing it, training the model, evaluating its outputs, refining behaviour and deploying the resulting system. Safety controls, human feedback and ongoing testing may be added at several stages.

A model is not the same thing as an application.

The same model may power several company products and third-party services through an API. One application may also switch between several models depending on the task, price or customer plan.

Power at the model layer comes from more than research talent. It also depends on access to training compute, data, capital, distribution, customer feedback, evaluation systems and cloud capacity.

Data is the missing ingredient

The five-layer diagram does not give data its own level, but no serious profile of AI can ignore it.

Models can be trained on public web material, licensed datasets, private company information, user interactions, synthetic data, human feedback, scientific records and copyrighted creative works. The composition and quality of that material influence what a model can do, which languages it handles well, which patterns it reproduces and whose work may have contributed to its development.

For creators, this is not an abstract technical concern. It reaches directly into questions of licensing, consent, attribution, compensation, disclosure and output similarity.

Energy powers the factory. Chips perform the calculations. Data helps determine what the resulting system knows, reproduces and overlooks.

Layer five: Where the industrial system disappears

Applications are the tools people actually use: chatbots, AI music generators, image systems, video tools, coding assistants, workplace software and specialized industry products.

This layer receives most public attention because it is visible and easy to compare. Users notice a new interface, a lower subscription price, a better voice, a faster song or a stronger image model. They do not see the energy contract, chip cluster or cooling system behind it.

Application companies still add real value. They create the interface, workflow, integrations, billing, storage, moderation, customer support, branding and industry-specific experience.

But many remain dependent on model providers, cloud companies, API pricing, app stores and distribution platforms. A product can appear independent to the customer while relying on several powerful suppliers underneath.

This creates a central business question: if strong models become widely available, what makes an application defensible?

The answer may be proprietary customer data, workflow integration, trusted identity, specialized expertise, community, distribution, regulatory approval or switching costs. A thin interface over someone else’s model may be useful, but it can also be easy for a larger platform to copy or bundle.

Who controls each layer?

Layer Main power centres Hardest bottleneck
Energy Utilities, generators, governments and grid operators Reliable supply and connections
Chips Designers, foundries, equipment and memory suppliers Advanced production capacity
Infrastructure Cloud firms, data-centre operators and investors Facilities, networks, land and cooling
Models AI laboratories and major technology companies Compute, data, talent and capital
Applications Platforms and specialist software companies Distribution, trust and retention

The higher layers can look more competitive because new applications appear constantly. The lower layers are slower to build, more capital-intensive, more regulated and harder to duplicate.

This creates the possibility of a market that looks diverse at the surface while becoming concentrated underneath.

Vertical integration makes that concentration more important. A company that controls cloud infrastructure, data-centre capacity, custom chips, model development and customer distribution can move costs and advantages across the whole system. It may deploy products faster, negotiate from a stronger position and make it harder for smaller rivals to compete on equal terms.

The question is not whether every large technology company will win. It is whether control across several layers allows a small group of firms to shape the conditions under which everyone else participates.

What the five-layer model leaves out

The five-layer model is useful because it is simple. It becomes misleading only when the simplicity is mistaken for completeness.

Several forces run through the entire system:

Capital

Money finances chip purchases, land, power contracts, construction, research staff, model training, acquisitions, customer subsidies, legal disputes and political influence.

The AI boom is not being built only through monthly subscriptions. It is supported by corporate cash flow, debt, venture capital, infrastructure investors, sovereign funds and public incentives.

Government

Governments shape every layer through electricity regulation, semiconductor policy, export controls, land-use approvals, tax incentives, research funding, public procurement, privacy rules and copyright law.

The facilities may be privately owned, but many of the conditions that allow them to exist are public decisions.

Labour

The AI factory depends on engineers, electricians, construction workers, fabrication technicians, network specialists, data-centre operators, researchers, annotators, evaluators, moderators, lawyers and support staff.

The output may appear automated. The supply chain remains deeply human.

Natural resources and materials

Data centres occupy land. Power generation and cooling can involve water. Semiconductor production depends on complex material and manufacturing supply chains. Servers eventually become electronic waste.

AI is digital at the point of use and physical at the point of production.

Culture and trust

Applications gain influence only when people accept them. Trust, resistance, regulation and social norms determine which uses become routine and which remain contested.

These forces are not extra layers neatly placed on top. They run through the entire factory.

Abundance at the top, scarcity underneath

AI applications promise abundance: more text, more images, more songs, more video, more code and faster analysis.

But that abundance is produced through scarce resources: advanced chips, reliable electricity, suitable land, grid connections, construction capacity, specialized labour and capital.

Artificial intelligence can make digital output cheaper while increasing competition for the physical resources required to produce it.

The same contradiction appears at the creator level. AI can lower the cost of generating a first draft without removing the cost of judgment, revision, documentation, distribution and attention. I examine that directly in AI Content Is Cheap—Why Being a Creator Costs More.

Artificial intelligence may create abundance at the top of the stack while intensifying scarcity at the bottom.

Who pays—and who bears the risk?

The AI factory distributes risk unevenly.

Technology companies risk spending heavily on infrastructure that may take longer than expected to generate adequate returns. Investors risk backing inflated expectations. Utilities risk building for customers whose demand may change. Governments risk offering incentives that do not produce the promised jobs, tax revenue or strategic value.

Communities may face land, water, power or environmental pressures. Workers may face disruption before new opportunities are clear. Creators and small businesses may build workflows around services whose prices, features or commercial terms can change quickly.

The public question is therefore not simply whether AI succeeds.

It is: when AI succeeds, who receives the gains? When projections fail, who absorbs the losses?

This is also why the industry’s forward-looking claims require discipline. An announced data centre is not an operating data centre. Reserved electricity is not delivered capacity. A benchmark is not a profitable product. A spending commitment is not a public benefit.

My related feature, One Year After Sam Altman Declared the Singularity, Who Actually Benefited?, examines what happened when the promises of rapid transformation collided with infrastructure costs, corporate influence and slower evidence of broad productivity gains.

What does the AI factory mean for independent creators?

Creators operate mainly at the application layer, but they remain exposed to decisions made across every layer beneath it.

A creator may never purchase an enterprise GPU or negotiate a power agreement. Yet chip shortages, cloud costs, model licensing, regional restrictions and infrastructure spending can still reach them through:

  • subscription increases;
  • generation limits;
  • slower service during peak demand;
  • feature removals or replacements;
  • changes to commercial permissions;
  • new disclosure or rights rules;
  • platform acquisitions or shutdowns;
  • pressure to buy several tools instead of one.

Do not confuse access with ownership

A subscription provides access under current terms. It does not guarantee permanent storage, unchanged pricing, continued commercial rights or identical model behaviour.

Preserve your own records

Keep local copies of prompts, lyrics, drafts, exports, stems, licences, contribution records, receipts and important version information. Your project should survive even when a platform changes.

Avoid complete dependence on one tool

You do not need to pay for five alternatives at once. You do need to understand how you would continue if your primary tool became too expensive, removed a feature or changed its terms.

Build value above the application

Your strongest creator assets are not merely access to the same generation button everyone else can press. They are your judgment, direction, documented human contribution, trusted identity, finished work, owned audience and repeatable process.

Creators cannot control the global compute stack. They can control how completely their work depends on any single point inside it.

Is AI becoming a utility?

AI compute is increasingly compared with electricity, telecommunications or cloud infrastructure because many organizations may depend on it without building their own systems.

The comparison is useful, but incomplete.

Electricity delivers standardized power. AI services can produce probabilistic outputs, reflect model biases, moderate information, shape recommendations and operate inside unresolved copyright and privacy disputes.

Compute may become utility-like in importance without becoming utility-like in ownership, regulation or public accountability.

That raises difficult questions:

  • Should governments maintain sovereign or public compute capacity?
  • Can a small number of private companies safely mediate access?
  • What obligations should apply to dominant providers?
  • Does compute access become a prerequisite for economic participation?
  • What should the public receive when public resources support private infrastructure?

What remains uncertain

This profile establishes that AI depends on a rapidly expanding industrial system. It does not establish that every announced project will be completed, every model will become profitable or every forecast will prove accurate.

Several major questions remain open:

  • Will paying demand grow fast enough to justify the current infrastructure build-out?
  • Will inference demand continue rising faster than efficiency improves?
  • Will lower computing costs reduce resource use or encourage even more use?
  • How much proposed electrical and data-centre capacity will actually be delivered?
  • Will governments demand stronger public returns for subsidies and grid support?
  • Will open and smaller models reduce concentration, or will infrastructure ownership remain decisive?
  • Where will long-term profits settle: chips, cloud infrastructure, models, applications or some combination?

There are at least three plausible outcomes.

Demand may catch up and make the build-out look necessary. Infrastructure may run ahead of profitable demand and force a correction. Or AI may become widely used while most economic value concentrates in a small number of infrastructure and platform companies.

The technology can be genuinely consequential and still be overhyped in the short term. Infrastructure can be necessary in the long run and still be built too quickly, too expensively or with too little public accountability.

A profile of power, not just technology

The five layers are more than a technical diagram.

Energy is political power. It determines what can operate and where.

Chips are strategic power. They determine who can train and deploy advanced systems.

Infrastructure is territorial power. It ties AI to land, grids, water, permits and jurisdictions.

Models are informational power. They influence how knowledge is generated, filtered and interpreted.

Applications are cultural power. They determine how people experience and normalize artificial intelligence.

Whoever controls only an application controls a product. Whoever controls several layers may influence the conditions under which an entire market operates.

Look past the prompt. Follow the factory.

The five-layer model is useful because it makes the invisible visible.

Applications sit at the top, where the public can see them. Energy, chips and infrastructure sit underneath, where many of the most consequential decisions are harder to observe.

AI is not arriving only through better software. It is being constructed through power generation, semiconductor factories, data centres, investment agreements, public incentives and political choices.

The question is no longer only what artificial intelligence can do.

It is who will own the systems that make it possible, who will pay to build them and who will have a meaningful voice in deciding what they become.

To understand AI, look past the application. Follow the factory.


Reporting and source note

This profile was reviewed July 29, 2026. It distinguishes completed infrastructure and reported financial results from forecasts, announced projects and company claims. Electricity projections can change as efficiency, construction, grid access and demand evolve.

Frequently asked questions

What does AI compute mean?

AI compute is the processing capacity used to train artificial-intelligence models and run them when users request an output.

What are the five layers of the AI factory?

The five layers are energy, chips, infrastructure, models and applications. Each layer depends on the layers beneath it.

Why do AI data centres need so much electricity?

They use electricity to power dense clusters of processors, servers, storage, networking and cooling systems that may operate continuously.

What is the difference between training and inference?

Training builds or improves a model. Inference is the processing used each time the trained model performs a task for a user or system.

Why should independent creators care about AI infrastructure?

Infrastructure costs and control can affect subscription prices, generation limits, feature availability, model access, platform stability and commercial terms.

Regresar al blog

Deja un comentario

Ten en cuenta que los comentarios deben aprobarse antes de que se publiquen.