As AI adoption accelerates, many organizations may feel pressure to invest in the newest models, the fastest GPUs, or the largest available cloud environments. However, purchasing powerful technology does not necessarily produce a strong business result. For manufacturers, life sciences companies, and other complex enterprises, AI infrastructure should be planned as a flexible portfolio rather than a one-time hardware purchase.
To get started, the organization must consider the complete environment, and connect those components to measurable business outcomes:
- Business use cases
- Applications
- Data sensitivity
- Computing
- Storage
- Networking
- Deployment options
- Software
- Security
- Employee skills
- Facilities
Below, we have described a general infrastructure planning process, which includes the following stages:
- Build an AI Workload Inventory
- Separate Workload Types
- Establish Performance Requirements
- Test Existing Infrastructure
- Compare Deployment Scenarios
- Run a Controlled Pilot
- Expand in Stages


With this process in mind, let’s dive into the specifics.
(Still uncertain about what AI platform you should use, or whether you should use a local LLM? If so, we recommend that you review our other articles on this topic first: Which AI Platform Is Your Best Conductor? and Are You Considering a Local LLM? )
Start With the Workload, Not the Hardware
Before comparing processors or signing a cloud contract, leadership should determine what the organization actually expects AI to do:
- How many employees, applications, or AI agents will use the system?
- Will the organization perform training, fine-tuning, or primarily inference?
- Which SAP systems, databases, laboratory platforms, manufacturing systems, or other applications will the AI access?
- What information must remain within a controlled environment?
AMD’s AI infrastructure planning guidance recommends beginning with the usage model because different activities create different demands. For instance, a retrieval-augmented generation application may require document processing, permissions, search, and storage in addition to model inference. On the other hand, an AI agent may also need CPU capacity to plan tasks, call tools, query databases, and coordinate workflows.
Because of this, GPU capacity should not be evaluated in isolation. A powerful accelerator could remain underused if the CPUs, memory, storage, or network cannot supply work quickly enough. Hardware also has limited value if the organization’s preferred models, applications, developers, and monitoring tools cannot use it effectively.
Build a Hybrid and Workload-Specific Environment
For many medium to large organizations, the best solution will not be exclusively cloud or exclusively on-premises. An IDC report sponsored by Intel predicted the following:
“By 2028, 75% of enterprise AI workloads will be deployed on hybrid fit-for-purpose infrastructure to turbocharge time to value while optimizing performance, cost, and compliance.”
Cloud services may suit early experimentation, variable demand, or temporary access to advanced accelerators. Meanwhile, on-premises systems may make more sense for predictable workloads, offline operations, or information requiring greater direct control. For instance, a life sciences company could experiment in an approved cloud environment while keeping regulated inference workloads private. A manufacturer might run an edge model near production equipment while using cloud capacity for periodic analysis.
Additionally, organizations should test their existing infrastructure before assuming that every workload requires specialized hardware. CPUs and current enterprise systems may adequately support smaller models, retrieval, data preparation, databases, and some inference. When testing demonstrates a bottleneck, the organization can add accelerators where they provide a measurable advantage.
On a similar note from the IDC report, most enterprises do not need to train a foundation model from scratch. They may be able to use a pretrained model, add retrieval-augmented generation, fine-tune an existing model, or deploy a smaller model for a narrow task. Routing requests among models can further control cost: Simple tasks may go to a smaller model, while difficult or high-risk work is directed to a more capable system.
Run a Controlled Pilot and Expand in Stages
A controlled pilot should measure the entire business workflow, including utilization, latency, failure rates, security events, employee adoption, human-review time, and cost per successful outcome. The organization can then progress from proof of concept to departmental pilot, shared service, production deployment, and enterprise expansion. Each stage should have performance, security, financial, and governance requirements.
When working to expand, AI demand is difficult to predict, especially when an organization is moving from isolated pilots to shared platforms. To mitigate this, rather than building around one forecast, leadership should maintain conservative, expected, and accelerated growth scenarios. Each scenario can estimate users, concurrent requests, model sizes, latency targets, storage growth, cloud expenses, and local power and cooling requirements.
Under its sections on demand forecasting, capacity modeling, and implementation roadmaps, Introl recommends maintaining multiple growth scenarios, updating them quarterly based on actual trends, preserving capacity for spikes, and expanding infrastructure in phases. Although the article’s specific numerical targets should be independently validated, its broader planning principle is useful:
Organizations should continuously compare forecasts against actual demand and adjust their infrastructure plans accordingly.
During any expansion, the organization should have a detailed, well-rounded financial model to have a strong understanding of the best direction. Within this model, total cost of ownership must extend beyond compute. Cost per token can be useful, but cost per accurate and valuable business outcome is the stronger executive metric. Below are some cost categories that should be considered:
| Cost category | Examples |
| Compute | CPUs, GPUs, accelerators, cloud instances |
| Memory and storage | RAM, high-bandwidth memory, vector databases, backups |
| Networking | Switches, interconnects, bandwidth, data transfer |
| Facilities | Rack space, electrical upgrades, power, cooling |
| Software | Model-serving tools, orchestration, monitoring, licenses |
| Personnel | Engineers, administrators, security, governance, support |
| Integration | APIs, identity management, business applications |
| Security | Testing, monitoring, incident response, audits |
| Lifecycle | Maintenance, support contracts, replacement, disposal |
| Downtime | Lost productivity or customer impact |
| Exit costs | Model migration, cloud egress, retraining, application changes |
Connect Infrastructure With AI Governance
Infrastructure determines where AI operates, what information it can reach, and how easily the organization can monitor or stop it. Governance therefore cannot be added after the technical design is complete.
To ensure proper governance, an executive steering committee can set strategy, risk tolerance, and investment priorities. Meanwhile, a cross-functional governance group can maintain policies, approved platforms, risk classifications, and an inventory of AI systems. Additionally, technology teams can manage architecture, capacity, reliability, and portability. Finally, legal, privacy, security, and business leaders evaluate the risks and outcomes of specific use cases.
On a deeper level, every AI system should also have a named business owner. The governance record should document its purpose, users, model and version, hosting location, data sources, integrations, risk level, testing results, human-review requirements, known limitations, cost, and next review date. Higher-risk applications—such as systems affecting employees, finances, regulated operations, quality, or safety—should receive stronger isolation, testing, logging, and human oversight.
Below is a potential governance structure:

If you would like a more in-depth read on this topic, the NIST AI Risk Management Framework discusses organizing AI risk work around four functions: Govern, Map, Measure, and Manage. Additionally, ISO/IEC 42001 outlines requirements for an AI management system that addresses areas such as leadership, policy, risk, data governance, lifecycle controls, monitoring, and continual improvement.
Build for Business Value and Responsible Growth
At the end of the day, the central question is not, “How many GPUs should we buy?”
Rather, business leaders should consider, “What combination of people, processes, platforms, and computing resources can deliver an approved business outcome at an acceptable level of cost and risk?”
By starting with defined workloads, testing existing capacity, using hybrid resources where appropriate, and scaling according to evidence, an organization can avoid both premature overinvestment and infrastructure shortages. When those decisions are connected to governance from the beginning, AI infrastructure becomes controlled foundation for long-term business transformation.
This was the concluding article on our series on How to Build an AI Foundation That Can Scale. Thank you, and stay tuned for future topics!
In the meantime, feel free to take a look at some of our other articles:
Preparing Today for the AI Industrial Shift Tomorrow





Leave A Comment