AI Infrastructure
Right-sized AI compute for organisations that are not hyperscale
Large enterprises have their own teams and budgets, and the global cloud providers serve their own scale well. What many mid-sized Thai organisations still lack is anyone designing a system that fits the size of their business. That gap is the one we set out to close.
Why now
The demand exists and the budget exists — what is missing is someone who can finish the job
PDPA and data-residency requirements
Customer, patient and financial data cannot be sent out for processing
GPU cloud costs that are high and hard to predict
For sustained use, owning the hardware pays back within a few years
Generative AI and RAG reaching real use
Inference has to be low-latency and close to the data
Compute hardware is almost entirely imported
Long lead times, so the partner has to be able to source and import
A national shortage of data-centre skills
Organisations have the hardware but nobody to run it — managed service is required
Where we sit
Between the box shop and the mega-project SI
We take work from a single server up to roughly two racks — the band a general IT shop cannot design for, and a large SI cannot justify taking.
| General IT shop | Ethernet (Thailand) | Large SI | |
|---|---|---|---|
| Project size taken | A machine | 1 server – 2 racks | 10 racks and up |
| Architecture design | ✗ | ✓ | ✓ |
| On-site power & cooling assessment | ✗ | ✓ | ✓ |
| Software stack deployment | ✗ | ✓ | ✓ |
| Agility and response time | High | High | Low |
| Minimum project cost | Low | Low–medium | Very high |
Common problems
What customers usually arrive with
Private AI assistant / RAG
They want an LLM over internal documents, but the data cannot leave
Inference server, vector DB and a web UI, all inside the organisation
AI development platform
Data scientists are competing for the same machine
A shared GPU pool with queuing and quotas
Computer vision / smart factory
Manual QC is slow and inconsistent
Edge nodes plus a central training node, with a retraining pipeline
Medical imaging / research
Patient data has to stay in the hospital
A closed internal system, designed around PDPA
VDI / graphics workstations
The design team needs powerful machines and needs to work anywhere
vGPU-based virtual workstations
Rendering / simulation
The render queue is long and deadlines slip
A small in-house GPU render farm
Small cloud / MSP
They want to sell GPU capacity but cannot fund a large build
A multi-tenant GPU pod with metering and billing
Reference architectures
Three tiers that grow into each other without a rebuild
Every tier runs the same software stack and the same network pattern, so a customer who starts at the first tier grows into the next without relearning anything and without throwing the first one away.
Tier 1 — Starter
A first PoC, a team of 5–20
- Nodes
- 1
- GPUs
- 2–4
- Approximate power
- 3–6 kW
- Network
- 10/25GbE
- Cooling
- Air (the existing server room)
- Location
- In the office
Tier 2 — Growth
Production use across a department
- Nodes
- 2–4
- GPUs
- 8 per node
- Approximate power
- 10–25 kW
- Network
- 100GbE RoCE
- Cooling
- Air + in-row
- Location
- Server room or colocation
Tier 3 — Private AI cloud
Serving the whole organisation, or resold as a service
- Nodes
- 6–16
- GPUs
- 8 per node
- Approximate power
- 40–100 kW
- Network
- 400GbE Spectrum-X or InfiniBand
- Cooling
- Air or liquid-assisted
- Location
- Colocation
Lead time and budget depend on the models chosen and on availability at the time — talk to us for figures against your actual requirement.
What we source
Chosen against the workload, not against the price list
NVIDIA data center GPUs
- NVIDIA L4 — low-power inference and video analytics, fits a general-purpose server
- NVIDIA RTX PRO 4500 / 6000 Blackwell Server Edition — generative AI, inference, graphics and simulation
- NVIDIA L40S — a cost-effective balance of AI and graphics work
- NVIDIA H200 NVL — fine-tuning, larger LLMs and HPC
- NVIDIA HGX / GB-series — for large builds
AI networking
- NVIDIA Spectrum-X Ethernet platform
- NVIDIA ConnectX SuperNIC and BlueField DPU
- NVIDIA Quantum InfiniBand
- Data-centre fibre and structured cabling
Servers, storage and supporting infrastructure
- Rack servers supporting 2, 4 or 8 GPUs
- NVMe all-flash, parallel file systems and NAS
- Racks, PDUs and UPS
- In-row cooling and DCIM
On-premise AI has only recently become realistic for mid-sized organisations, because the entry point is no longer an 8-GPU system — it can start with a 2U server in the server room that is already there.
Import & compliance
The real hardware, on time, and correctly declared
Buying high-performance compute today is not only a question of price. What sinks projects is usually late delivery and unclaimable warranty, not a wrong specification.
Sourcing and allocation management
Sourced through authorised channels, which removes the risk of mismatched or unwarrantable stock, with lead-time status reported transparently throughout.
Trade compliance
Correct customs classification and import documentation; compliance with export-control requirements on high-performance compute, including end-user and end-use verification; and documentation for BOI-privileged buyers.
Logistics and inspection
Transit insurance, shock-controlled handling for high-value equipment, condition inspection on arrival, photographed serial numbers, and an asset register.
Warranty and RMA
Warranty registered in full with the manufacturer, and claims handled by us as the intermediary — the customer never has to deal with an overseas RMA desk.
After delivery
Three levels, chosen by how critical the system is
| Standard | Business | Mission-critical | |
|---|---|---|---|
| Service hours | 8x5 | 12x5 | 24x7 |
| On-site attendance | On request | Next business day | As agreed in the SLA |
| Spare parts | Queued | Held in country | Advance replacement |
| Preventive maintenance | Annual | Twice a year | Quarterly |
| Remote monitoring | — | ✓ | ✓ with proactive alerting |
| Performance reporting | — | Quarterly | Monthly |
Every level includes a health check 30 days after handover, a direct line to the engineering team, and complete as-built documentation.
Already have a project in mind?
Send us the outline. We will come back with the questions worth answering before we book a site visit.

