Why Smaller AI Factories Matter for the Future of Sovereign AI

Much of the discussion around AI infrastructure is currently focused on scale. New projects are increasingly being discussed in hundreds of megawatts, with some proposed AI campuses extending into gigawatt territory. There are good reasons for this. Training increasingly capable foundation models requires enormous amounts of compute, and concentrating thousands of GPUs in a small number of locations can provide significant operational and economic advantages.
However, this does not necessarily mean that gigawatt-scale campuses are the right architecture for every AI workload. As AI moves from model development into widespread production use, inference will account for an increasingly important part of the infrastructure requirement. Inference has different characteristics from training, particularly when applications depend on responsiveness, data locality and predictable access to compute.
For these workloads, smaller AI factories in the 1 to 5 MW range could become an important part of national AI infrastructure. Rather than concentrating all AI capacity into a small number of very large campuses, countries could develop distributed networks of regional AI facilities located close to cities, industrial centres and the organisations consuming AI services.
This could be particularly important for Sovereign AI. A distributed model provides the opportunity to keep data and AI processing within national or regional boundaries while also placing the infrastructure closer to the workloads it supports.

The attraction of very large AI campuses is understandable. AI infrastructure benefits from scale, particularly for training large models where thousands of accelerators need to operate together. The difficulty is that building a gigawatt AI campus involves much more than finding enough GPUs and constructing a data centre.
Power is becoming one of the largest constraints. The International Energy Agency expects global data centre electricity consumption to reach around 945 TWh by 2030, more than double its 2022 level, with AI being the most important driver of that growth. In the United States, Lawrence Berkeley National Laboratory estimates that data centres could consume between 6.7% and 12% of total US electricity generation by 2028.
At hundreds of megawatts or gigawatt scale, providing power becomes a major energy infrastructure program in its own right. New transmission infrastructure, substations, generation capacity and grid connections may all be required. Projects must also address land availability, cooling, heat rejection, fibre connectivity and planning permission.
The UK Government has already identified grid connectivity as one of the major constraints affecting the development of AI infrastructure. Its AI Growth Zone program requires proposed sites to demonstrate a credible route to at least 500 MW of power availability by 2030, either through the electricity network or through behind-the-meter generation.
The issue is therefore not whether gigawatt campuses can be built. They clearly can and will be. The question is whether concentrating the majority of future AI capacity into a relatively small number of extremely large locations creates unnecessary infrastructure constraints.

Smaller AI factories provide an alternative approach. Instead of concentrating several gigawatts into a small number of locations, part of that capacity could be distributed across many regional facilities.
A 1 to 5 MW AI factory is small compared with a hyperscale campus, but it can still support a substantial amount of accelerated computing infrastructure. Modern GPU systems allow significant inference capacity to be deployed within relatively compact facilities. More importantly, facilities of this size can potentially be deployed in a much wider range of locations.
This creates the possibility of a tiered national AI infrastructure. Large AI campuses can provide the enormous concentrations of compute required for frontier model training and other highly intensive workloads. Regional AI factories can provide inference and enterprise AI services closer to population centres and industry. Smaller edge infrastructure can then support applications requiring even greater locality.
This model is similar to the way other infrastructure has evolved. Capacity does not have to exist in one place to operate as a single service. A distributed orchestration layer can present multiple physical AI factories as a common pool of national compute capacity.
Workloads could then be placed according to their requirements. Factors such as latency, data sovereignty, GPU availability, energy availability, cost and security classification could influence where an application runs.
This creates a very different model from simply building a large data centre and directing every workload towards it.

One of the most important distinctions is between AI training and AI inference.
Training is generally tolerant of physical distance from the eventual user. A model that takes several weeks to train does not usually care whether the GPUs performing that training are 20 miles or 500 miles from the people who will eventually use the model.
Inference is different because it sits directly in the application path. A request is generated, processed by an AI model and a result is returned. Network latency therefore becomes part of application latency.
The UK Compute Roadmap recognises this distinction, stating that AI training is generally less dependent on location while inference can benefit from proximity to data sources and end users. The same roadmap estimates that the UK will require at least 6 GW of AI-capable data centre capacity by 2030.
This raises an important infrastructure question. If several gigawatts of new AI capacity are required, how much should be concentrated into large training campuses and how much should be distributed closer to where inference is actually consumed?
The answer will vary by workload, but it is unlikely that every application will benefit from being served from a small number of enormous centralised facilities.

Smart Cities provide a useful example of why location matters. A modern city may contain thousands of cameras, traffic systems, environmental sensors, connected vehicles and other sources of continuously generated data. Increasingly, AI will be used to interpret this information and make decisions.
Traffic management is an obvious example. Computer vision systems could identify congestion, accidents, pedestrians or unusual behaviour and provide information to traffic management platforms. Similar systems could be used for public transport, infrastructure monitoring, emergency services and environmental management.
Sending all of this raw information to a distant hyperscale facility is not always the most efficient architecture. In many cases, the valuable output of the AI model may be very small compared with the amount of source data.
A video stream might contain enormous quantities of information, while the useful inference result might simply indicate that an accident has occurred at a particular location. Processing the video regionally means that the raw stream does not necessarily need to traverse the national network before it can be analysed.
ETSI’s work on Multi-access Edge Computing specifically identifies low latency, high bandwidth and real-time applications as key reasons for moving compute closer to users and devices. Use cases include video analytics, IoT and vehicle-to-everything communications.
Regional AI factories could provide a similar capability at metropolitan scale. They would not replace small edge devices located directly beside sensors, nor would they replace hyperscale AI campuses. They would provide the infrastructure layer between those two environments.

The discussion around edge and regional computing often focuses entirely on latency, but there is an equally important issue: data movement.
AI applications can generate very large quantities of source data. Video, telemetry, industrial sensors and connected vehicles are good examples. Continuously transferring this information over long distances consumes network capacity and potentially introduces additional cost.
Processing information closer to its source means that organisations can transfer the result rather than the complete dataset. Raw data can remain local where appropriate, while metadata, alerts or aggregated information can be sent to other systems.
This is also relevant to privacy and sovereignty. If sensitive information can be processed within the region where it was generated, there may be less need to move that information between locations.
In simple terms, instead of always moving the data to the AI, we increasingly have the option of moving the AI closer to the data.

Sovereign AI is frequently discussed as a question of where data is stored. Data residency is important, but genuine AI sovereignty involves more than storage location.
Countries and organisations also need to consider where models execute, who controls the infrastructure, which jurisdiction applies to that infrastructure and how dependent critical AI services are on external providers.
The UK Government has stated that domestic data centre capability is important for protecting sensitive data and increasing resilience to global disruption. A distributed AI infrastructure could extend this concept beyond simply ensuring that capacity exists somewhere within the country.
Regional AI factories could support government, healthcare, financial services, universities, manufacturing and other industries while keeping inference within sovereign infrastructure. Certain workloads could remain within specific regions where necessary, while others could move between facilities according to available capacity.
This creates a national AI fabric rather than a collection of isolated data centres.
Sovereignty then becomes a characteristic of the overall architecture.

Power remains a challenge regardless of data centre size, but smaller facilities create more options for how that power can be delivered.
Providing 1 GW of continuous power at a single location is an enormous engineering challenge. Providing a few megawatts is a much more conventional infrastructure problem.
A smaller AI factory could potentially combine grid connectivity with local generation, battery storage and other behind-the-meter technologies. Depending on location and requirements, this could include solar, wind, fuel cells, gas generation, waste-to-energy or local microgrids.
This does not mean that a 5 MW AI factory can simply operate independently of the electricity grid. However, local generation and storage can reduce peak grid requirements and provide additional resilience.
The US Department of Energy and Lawrence Berkeley National Laboratory have highlighted onsite generation, energy storage and data centre load flexibility as potential mechanisms for accommodating rapidly increasing data centre electricity demand. The UK AI Growth Zone framework also explicitly allows behind-the-meter generation to contribute towards meeting power requirements.
At smaller scale, these approaches may be considerably easier to repeat across multiple locations.
There is also an interesting opportunity to make the compute itself energy-aware. Not every inference request or AI workload has the same urgency. Workloads that are not latency-sensitive could potentially be shifted between regional AI factories according to available compute capacity, energy availability and electricity cost.
The result would be an AI infrastructure platform that considers both compute and energy when deciding where workloads should execute.

Another benefit of a distributed architecture is resilience.
Concentrating large amounts of capacity into one location also concentrates risk. Power failures, cooling problems, fibre outages or other infrastructure failures can potentially affect a significant amount of capacity at once.
Uptime Institute continues to identify power as the leading cause of serious and severe data centre outages.
With a distributed architecture, the failure of one facility does not necessarily remove access to the entire AI platform. Workloads could potentially be redirected to another regional AI factory, assuming applications and orchestration platforms are designed to support this.
Capacity can also be expanded more incrementally. Rather than attempting to predict demand years in advance and constructing hundreds of megawatts at once, operators can add regional capacity as demand develops.
A city or industrial region might begin with a 1 MW deployment and expand towards 5 MW as utilisation increases. New locations can be added where demand appears rather than requiring all future demand to be predicted when the original campus is designed.
The network becomes part of the AI platform
A distributed model also changes the role of networking.
If multiple regional AI factories are operated as a common infrastructure platform, connectivity between those locations becomes an important part of the architecture. High-capacity fibre, software-defined networking and intelligent workload placement could allow applications to consume AI infrastructure without needing to know exactly where the underlying GPUs are located.
A workload requiring extremely low latency could be kept within the nearest metropolitan facility. A less time-sensitive workload could be moved to another region with spare capacity. Sensitive workloads could be restricted to approved facilities, while large batch workloads could be directed towards lower-cost capacity.
This could also allow AI infrastructure to respond dynamically to changes in energy availability. Compute capacity in one region might be favoured at certain times because renewable generation or local energy storage is available, while another region could carry more workloads at a different time.
The orchestration layer therefore becomes just as important as the physical GPU infrastructure.
This is not an argument against gigawatt AI campuses
Large AI campuses will remain important. Frontier model training, large-scale research and other highly compute-intensive workloads benefit enormously from having thousands of accelerators concentrated together.
The mistake would be assuming that the infrastructure designed for those workloads should automatically become the architecture for all AI.
The future is more likely to consist of several layers of compute. At one end will be hyperscale AI factories providing enormous pools of training capacity. At the other will be embedded AI running directly inside vehicles, cameras, robots and industrial equipment.
Regional AI factories in the 1 to 5 MW range could occupy the important space between them. They can provide substantial GPU capacity while remaining close enough to cities, businesses and industrial areas to support responsive inference.
They can also be deployed incrementally, integrated with local energy infrastructure and interconnected to provide resilience.

The AI infrastructure debate should therefore move beyond measuring success purely in gigawatts.
The amount of compute a country controls is clearly important, but where that compute is located, how it is powered and how easily workloads can access it will become increasingly important as AI adoption grows.
A country with several gigawatts of AI capacity concentrated in one or two locations certainly has significant computational capability. A country with large national training facilities combined with a network of regional AI factories has something different. It has the beginnings of a distributed AI infrastructure platform.
That architecture could place inference close to Smart Cities, industrial facilities, universities, healthcare systems and businesses while maintaining national control over the infrastructure and the data processed by it.
It could also reduce dependence on a small number of extremely large grid connections by spreading demand geographically and making greater use of behind-the-meter energy, storage and local generation.
Gigawatt campuses will undoubtedly form an important part of the AI infrastructure landscape. However, they should not be the only model.
As AI moves from experimentation towards becoming part of everyday infrastructure, the requirement will increasingly be not simply for more compute, but for compute in the right place.
For Sovereign AI, that may mean thinking nationally about infrastructure while deploying much of it regionally.
The future AI factory does not always need to be measured in gigawatts. In many locations, the most useful AI factory may be the 1 to 5 MW facility located close to the people, data and applications that actually need it.
