From Small AI To AI Boxes: Rethinking AI As Infrastructure

Source: SpeedUP Technologies Vietnam.
Written by Lợi Đoàn (Luke), Deputy CEO of SpeedUP Technologies Vietnam.
When we talk about AI, we often think of ChatGPT, chatbots, or virtual assistants. But applications are only the most visible part. Behind an answer delivered in seconds, power, chips, networks, storage, and operational software must all work together.
At Davos in January 2026, NVIDIA founder and CEO Jensen Huang described AI as a “five-layer cake”: energy, chips, data center and cloud infrastructure, models, and applications.
From an operational perspective, this framework raises a question for business leaders: beyond choosing an AI application, does the company understand what it takes to keep that application reliable, secure, and economically viable?
The Five Layers Behind an Answer
The first layer is energy. According to the International Energy Agency’s (IEA) 2025 Energy and AI report, data centers worldwide consumed approximately 415 billion kWh of electricity in 2024, with consumption projected to reach around 945 billion kWh by 2030. These figures cover all data centers, with AI the most significant driver of growth.
Power brings a thermal challenge. In high-density AI server clusters, cooling directly affects the ability to sustain performance.
Liquid cooling options need to be assessed against the configuration and deployment conditions; no single power threshold applies to every server room.
The second layer is chips. GPUs, processors capable of performing many calculations in parallel, play an important role in many AI systems. But powerful chips alone are not enough: memory must accommodate the model, and data must arrive fast enough to keep the chips working.
The third layer is data center and cloud infrastructure. Networks, storage, and orchestration software connect processors into a functioning system. This is where computing capacity on a specification sheet becomes real-world performance, or is lost to bottlenecks.
The fourth layer is models, which process data to generate predictions, content, or inferences. Model selection must account for the task, data quality, and response requirements. A larger model is not automatically a better fit for every job.
The fifth layer is applications, where users see value: finding documents faster, handling customer requests, or supporting product inspections. That value can only be sustained if the system beneath it meets the demands placed on it.
Where A Problem Appears May Not Be Where It Starts
One operational principle is worth remembering: the point where a problem becomes visible and the source of that problem may be two different places.
Imagine an AI system responding slowly. Users may suspect the application or model. But if GPUs are waiting for data from storage, buying more GPUs may simply add more processors to the queue. If a server slows down because it is overheating, the cooling system needs attention.
This does not mean every AI problem comes from hardware. Incorrect answers can also originate in the data, the model, or the application’s design. The mistake to avoid is drawing conclusions before taking measurements.
Bringing AI Closer To Where Data Is Created
Enterprise data resides in factories, transaction systems, and internal document repositories. As AI becomes part of everyday work, the question of where that data is processed becomes more practical.
Consider a camera inspecting products on a production line. Processing its images at the factory can reduce the amount of data transmitted and dependence on an internet connection. For a document retrieval assistant, running on private infrastructure can give the business more control over access permissions and usage logs.
These are the drivers behind deploying AI close to data: reducing latency, limiting data movement, and improving control. The benefits depend on the actual architecture; they do not appear automatically simply because a server is installed on company premises.
Technology providers are also developing distributed computing capabilities. NVIDIA and its telecommunications partners have introduced AI Grid, bringing inference, the use of a model to produce results, closer to where data is generated.
Small AI And the Opportunity For AI Boxes
AI infrastructure often conjures images of thousands of GPUs and enormous data centers. Yet internal assistants, contract analysis, and visual product inspections do not necessarily require that scale.
We can call an approach tailored to these needs Small AI: start with a specific task, select an appropriate model, and then determine the resources required. “Small” refers to the scope of deployment and the scale of infrastructure, not necessarily the value it delivers.
One way to deploy it is an AI Box, a “miniature AI factory” on company premises. It integrates servers and GPUs, networking, storage, a resource management platform, models, and applications into a single system.
A business can begin with one or a few servers serving a user group or workflow, then expand once demand has been measured.
The decision process should move from the task to the model, and from the model to suitable GPUs and infrastructure. Buying the biggest GPU first and then finding work for it is a recipe for waste.
A Compact System Still Needs a Solid Foundation
An AI Box makes AI infrastructure more accessible to businesses, but it still needs reliable power, cooling, maintenance, and someone accountable for operations.
For high-density configurations, immersion cooling, submerging compatible equipment in a dielectric fluid, can help manage heat in a smaller footprint. It is an architectural option to evaluate alongside air cooling and direct-to-chip liquid cooling.
Even when immersion cooling is integrated into an AI Box, the business must still plan for power delivery, heat rejection, safety, and maintainability. Packaging the equipment together does not remove these physical requirements.
An AI Box can also operate alongside the public cloud. Steady workloads that need processing close to data can run on-premises, while experiments or demand spikes can use rented resources.
Lenovo’s 2025 cost analysis suggests that owned infrastructure may offer advantages when used consistently, while the cloud suits short-term or variable workloads. This is a vendor analysis with a limited cost scope. Businesses still need to account fully for electricity, cooling, staffing, software, and maintenance in their own circumstances.
Security Must Run Through Every Layer
Security needs to be present at every layer, from equipment, networks, and data to models and the actions an application is authorized to perform.
At the application layer, OWASP identifies the risk of “prompt injection”: malicious content in documents or external sources can cause AI to follow unintended instructions. The risk increases when an assistant has permission to send emails, access data, or take actions within a system.
Businesses therefore need to limit access privileges, apply input and output controls, maintain logs, and require human approval for high-risk actions. A single layer of protection should not be considered sufficient.
Từ mua công nghệ đến xây năng lực
Trước khi mở rộng, lãnh đạo cần biết AI đang cải thiện công việc nào, dữ liệu nằm ở đâu và kết quả được đo ra sao. Quyết định đầu tư phải đi cùng đánh giá điểm nghẽn, chi phí vận hành và trách nhiệm bảo vệ dữ liệu.
Small AI mở ra một cách bắt đầu vừa sức: đưa năng lực xử lý đến gần dữ liệu, kiểm chứng giá trị trong phạm vi nhỏ rồi mở rộng. AI Box chỉ có ý nghĩa khi phục vụ được công việc thực tế bằng một hệ thống ổn định, an toàn và có thể quản trị.
Ứng dụng là nơi doanh nghiệp nhìn thấy giá trị. Hạ tầng giúp giá trị ấy được tạo ra bền vững. Trong kỷ nguyên AI, hạ tầng không còn là phần hậu trường của chiến lược công nghệ; nó trở thành một phần của chính chiến lược kinh doanh.
About the Author: Lợi Đoàn (Luke)With more than 25 years of experience in researching, operating, and implementing IT infrastructure, cloud computing, and AI infrastructure solutions, including over 15 years with FPT Corporation - Lợi Đoàn (Luke) was among the key contributors who laid the foundation for the FPT HI GIO Cloud platform. He is also a pioneering expert in the research and development of immersion cooling solutions in Vietnam. He currently serves as Deputy CEO of SpeedUP Technologies Vietnam.About SpeedUP Technologies Vietnam
SpeedUP Technologies Vietnam is a technology company specializing in automation solutions, the AI Tank ecosystem, AI Box infrastructure, advanced immersion cooling, and AI-powered enterprise platforms.