Compute · Self-built data center · Hosting

Self-built data center

Easy AI owns a self-built data center. We check your workload and duration against available resources and agree who does what.

ForEnterprise technical teams · Infrastructure partners

Easy AI data centerSample

Workload request

Workload
Inference service
Use
Internal knowledge-base Q&A
Model
Open-source LLM, self-hosted
Duration
Per project
Start
Next quarter
Asking
Whether your resources fit
Company and figures are placeholders. Send us yours.

Ways to deploy

  • Public cloud GPUs

    On-demand compute in the cloud

    Suits
    Short trials and spiky workloads
    Start
    Available as soon as you sign up
    Billing
    Hourly or metered; costly when it runs for long
    Data
    Stays in the cloud provider’s environment
  • Self-built data center

    Data-center resources matched to the workload

    Suits
    Long-running inference and periodic training or fine-tuning
    Start
    Resources confirmed by configuration and duration
    Billing
    Quoted by configuration and term
    Data
    Data-center location and network isolation can be agreed
  • On-premises

    Hardware in your own environment

    Suits
    Data that must not leave the company
    Start
    Longer purchasing, racking and deployment
    Billing
    Up-front hardware, operations on top
    Data
    Stays entirely in-house

The three can be combined, for example a cloud trial that moves to our data center once it is stable.

Sizing guide

Rough GPU memory for serving each maker’s latest open models in-house, from their official weights. A starting point for a request; measure on the target hardware.

ModelTotal / active parametersOfficial weightsInference deployment
Language models
Kimi K32.8T / 104BMXFP4, ~1.56 TBOne server of 8 × 288 GB-class; or two of 8 × 141 GB-class
Qwen3.8-2.4T2.45T / 95BFP8, ~2.5 TBFrom two servers of 8 × 192 GB-class; 288 GB-class for long context
DeepSeek-V4-Pro-08131.6T / 49BFP4 + FP8, ~893 GBFrom one server of 8 × 141 GB-class; 192 GB-class for long context
GLM-5.3753B (MoE)FP8, ~756 GBOne server of 8 × 141 GB-class; or two of 8 × 80 GB-class
DeepSeek-V4.1-Flash552B / 16BFP4 + FP8, ~510 GBOne server of 8 × 141 GB-class; 8 × 80 GB-class is tight
MiniMax-M3428B / 23BMXFP8, ~444 GBOne server of 8 × 80 GB-class; BF16 needs 8 × 141 GB-class
Video generation
MiniMax-H333B denseBF16, ~144 GB in allOfficial example on 4 GPUs; from four 80 GB-class
Smaller models
70B and smallerDense modelsBF16, ~2 GB per billion parametersOne or two 80 GB-class GPUs; one with INT4
  • MoE models need memory for all parameters: the active count sets compute per token, but every expert must be loaded.
  • Add KV cache and 10–20% headroom. These language models take million-token context; more concurrency and longer context need more, possibly another server.
  • Video models need more memory and time as resolution, length and concurrency grow; H3’s total covers the transformer, text encoder and video decoder. Classes describe memory per GPU only, not specific models in stock.
Sizing GPU memory for a self-hosted model

Use cases

  • Self-hosting

    An open-source model needs long-running inference

  • Training

    Resources for a few weeks or months

  • Data rules

    Data location and network isolation requirements

  • Resource partners

    Compute or data-center resources to cooperate on

Scope

Discussable directly

  • Whether resources fit the workload
  • Deployment environment and networking
  • Duration and cooperation model
  • Resource supply cooperation

Confirm separately

  • Location, equipment and available capacity
  • External rental and operations scope
  • Certifications, backups and SLA terms
  • Migration and exit at the end of the term

Part of the data center runs Easy AI’s own business, so not all of it is available. Owning a data center does not mean third-party models run in it.

How we work

  1. 1

    Describe the workload

    Inference or training, model size, duration.

  2. 2

    Check resources

    Whether compute, storage and network fit.

  3. 3

    Split responsibilities

    Access, operations, backups and exit.

Quote inputs

  • Workload type
  • Model size
  • Compute and storage
  • Network needs
  • Duration
Contact us

Contact us

Topic: Infrastructure capability

WeChatLi___CaB6

Contact page