Compute · Self-built data center · Hosting
Self-built data center
Easy AI owns a self-built data center. We check your workload and duration against available resources and agree who does what.
ForEnterprise technical teams · Infrastructure partners
Workload request
- Workload
- Inference service
- Use
- Internal knowledge-base Q&A
- Model
- Open-source LLM, self-hosted
- Duration
- Per project
- Start
- Next quarter
- Asking
- Whether your resources fit
Ways to deploy
The three can be combined, for example a cloud trial that moves to our data center once it is stable.
Sizing guide
Rough GPU memory for serving each maker’s latest open models in-house, from their official weights. A starting point for a request; measure on the target hardware.
| Model | Total / active parameters | Official weights | Inference deployment |
|---|---|---|---|
| Language models | |||
| Kimi K3 | 2.8T / 104B | MXFP4, ~1.56 TB | One server of 8 × 288 GB-class; or two of 8 × 141 GB-class |
| Qwen3.8-2.4T | 2.45T / 95B | FP8, ~2.5 TB | From two servers of 8 × 192 GB-class; 288 GB-class for long context |
| DeepSeek-V4-Pro-0813 | 1.6T / 49B | FP4 + FP8, ~893 GB | From one server of 8 × 141 GB-class; 192 GB-class for long context |
| GLM-5.3 | 753B (MoE) | FP8, ~756 GB | One server of 8 × 141 GB-class; or two of 8 × 80 GB-class |
| DeepSeek-V4.1-Flash | 552B / 16B | FP4 + FP8, ~510 GB | One server of 8 × 141 GB-class; 8 × 80 GB-class is tight |
| MiniMax-M3 | 428B / 23B | MXFP8, ~444 GB | One server of 8 × 80 GB-class; BF16 needs 8 × 141 GB-class |
| Video generation | |||
| MiniMax-H3 | 33B dense | BF16, ~144 GB in all | Official example on 4 GPUs; from four 80 GB-class |
| Smaller models | |||
| 70B and smaller | Dense models | BF16, ~2 GB per billion parameters | One or two 80 GB-class GPUs; one with INT4 |
- MoE models need memory for all parameters: the active count sets compute per token, but every expert must be loaded.
- Add KV cache and 10–20% headroom. These language models take million-token context; more concurrency and longer context need more, possibly another server.
- Video models need more memory and time as resolution, length and concurrency grow; H3’s total covers the transformer, text encoder and video decoder. Classes describe memory per GPU only, not specific models in stock.
Use cases
Self-hosting
An open-source model needs long-running inference
Training
Resources for a few weeks or months
Data rules
Data location and network isolation requirements
Resource partners
Compute or data-center resources to cooperate on
Scope
Discussable directly
- Whether resources fit the workload
- Deployment environment and networking
- Duration and cooperation model
- Resource supply cooperation
Confirm separately
- Location, equipment and available capacity
- External rental and operations scope
- Certifications, backups and SLA terms
- Migration and exit at the end of the term
Part of the data center runs Easy AI’s own business, so not all of it is available. Owning a data center does not mean third-party models run in it.
How we work
- 1
Describe the workload
Inference or training, model size, duration.
- 2
Check resources
Whether compute, storage and network fit.
- 3
Split responsibilities
Access, operations, backups and exit.
Quote inputs
Guides
All guides- How to describe infrastructure needsWorkload type, model size, duration and data location.
- Splitting data-center responsibilitiesHardware, access, backups and exit: who owns each.
- Sizing GPU memory for a self-hosted modelEstimate memory for weights, KV cache and headroom, weigh quantization, then measure on real hardware.
- Training, fine-tuning and inference comparedHow the three workloads differ in memory, storage, network and duration, and what to state for each.