AI costs push companies towards on-premise computing
The shift is particularly visible among businesses running persistent inference, model fine-tuning, computer vision and generative AI workloads. Cloud platforms continue to offer speed, flexibility and access to advanced GPUs without heavy upfront investment, but hourly charges can become substantial when expensive accelerators remain active around the clock.
GPU prices available through the IndiaAI Compute Portal illustrate the scale of the calculation facing enterprises. On-demand access to a single Nvidia H100 SXM GPU is listed at about ₹153 an hour, while a 12-month reservation brings the rate down to around ₹117. An H200 SXM GPU is available at roughly ₹140 an hour on demand in some configurations, compared with about ₹100 under longer reservations.
More powerful systems cost considerably more. A two-GPU Nvidia B200 configuration is priced at about ₹581 an hour on demand, while an eight-GPU configuration exceeds ₹2,300 an hour. Prices vary by provider, architecture, reservation period and configuration, making headline GPU rates only one part of the overall expense.
For companies operating workloads continuously, those charges accumulate quickly. An H100 running uninterrupted at ₹153 an hour would generate compute charges of more than ₹1.3 million over a year before storage, networking, data movement and other services are considered. Larger clusters can multiply that figure rapidly.
That arithmetic is strengthening the argument for owning AI infrastructure where utilisation is predictable. Purchasing servers places GPUs directly under enterprise control and removes recurring rental charges, although companies must fund hardware, networking, electricity, cooling, maintenance and technical staff before any savings emerge.
Utilisation is therefore becoming the critical variable.
A GPU server operating only occasionally can become an expensive idle asset. Cloud infrastructure remains attractive for experimental projects, irregular training runs and companies whose AI demand changes sharply from week to week. Capacity can be increased or released within minutes without purchasing equipment that may remain unused.
The equation changes when GPUs operate continuously at high utilisation. Companies running stable production inference or regular model training can potentially spread hardware costs over several years and lower the effective cost of each computing hour.
Hardware ownership nevertheless carries risks that cloud customers largely transfer to their providers. AI processors are evolving rapidly, meaning expensive equipment can lose relative competitiveness long before it physically wears out. New generations from Nvidia, AMD and other suppliers offer higher performance, larger memory and improved energy efficiency, potentially altering the economics of systems purchased only two or three years earlier.
Software compatibility can be equally important. Nvidia's CUDA ecosystem remains deeply embedded across AI frameworks and enterprise applications. Alternative accelerators may offer attractive pricing or memory specifications, but migration can require engineering work, optimisation and testing that reduce theoretical savings.
Performance also varies significantly by workload. A cheaper GPU is not automatically less expensive if a model takes longer to produce the same number of tokens or complete the same training task. Enterprises are increasingly measuring cost per inference, cost per million tokens and useful output per watt rather than relying solely on hourly rental prices.
Cooling and power requirements present another obstacle to bringing AI infrastructure inside corporate facilities. High-density AI racks can demand far more electricity than conventional enterprise servers, with newer configurations requiring specialised liquid cooling and upgraded power distribution.
This is encouraging a third model between public cloud and equipment installed inside company offices. Enterprises can purchase or reserve dedicated GPU infrastructure housed in specialised data centres, combining greater control with professionally managed power, cooling and connectivity.
Hybrid deployment is also gaining ground. Sensitive data and predictable inference can remain on dedicated infrastructure, while cloud GPUs absorb temporary spikes, experimental workloads and unusually large training jobs.
Data governance adds another dimension. Organisations handling financial, healthcare, government or proprietary information may prefer local or dedicated systems because they offer tighter control over where data and models are processed. Cloud providers have responded with private-cloud, sovereign-cloud and dedicated infrastructure options, narrowing some of that distinction.
India's expanding shared compute infrastructure is meanwhile altering the cost equation. More than 38,000 GPUs have been empanelled under the IndiaAI Mission, with another 20,000 planned as part of efforts to broaden access to advanced computing. The programme allows eligible startups, researchers, government bodies and other users to access subsidised or competitively priced GPU capacity without making large capital investments.
Competition among cloud providers, domestic data-centre operators and specialised GPU companies is also pushing enterprises towards more granular purchasing decisions. Reserved capacity, spot pricing, smaller inference accelerators and purpose-built AI processors can reduce bills without forcing companies to abandon cloud infrastructure entirely.