Mumbai · Pune · Chennai · Noida

GPU capacity
that lives
in India

Rent H200s, H100s and L40S by the second from our own Tier III floors. Your data never leaves the country, and you never pay to get it out.

Per-second billing · Zero egress · No commitment on on-demand

MSL Sovereign GPU Infrastructure

Running production workloads on MSL

Vantara AIKettle LabsNorthline BankPraxis HealthOrbit StudiosMeridian Auto
U34

From nothing to training

Four commands.
No sales call.

Sign up with a company email, add a payment method, and the API is live. Enterprise procurement exists if you want it; it just isn't in the way.

  • CLI, REST and Terraform: the same resource model in all three
  • OCI images: bring your own container or start from our CUDA bases
  • Persistent volumes: attach the same NVMe volume to any pod in the region
  • Idle timeout: pods stop themselves so a forgotten notebook can't cost you a weekend
bash: provision an H100 pod
# install and authenticate
$ pip install msl-cli
$ msl auth login

# eight H100s, a 2 TiB scratch volume, Mumbai
$ msl pods create \
    --gpu h100-sxm --count 8 \
    --image msl/pytorch:2.4-cu124 \
    --volume scratch:2Ti --region bom1

✓ pod-7fk29d running in 38s
✓ ssh msl@7fk29d.bom1.mslproducts.com

The floor, in numbers

4 sites

Mumbai · Pune · Chennai · Noida

18 MW

Contracted IT load

1.38

Design PUE

99.99%

Uptime commitment

U22

Why teams move here

The boring
reasons matter

Latency

Under 10 ms to your users

Inference served from Mumbai reaches most of western India in single-digit milliseconds. A US region cannot do that, whatever the GPU costs there.

Residency

Data stays where the law wants it

RBI localisation, the DPDP Act and sectoral rules all point the same way. Our regions are in-country and audited, so residency stops being a design constraint.

Egress

Nothing to pay on the way out

Moving a checkpoint set out of a hyperscaler can cost more than training it. We don't meter egress at all: not on object storage, not on pods.

Support

Engineers in your timezone

The person who answers at 3 a.m. IST can see the rack, the switch and the PDU. Escalation is a corridor, not a ticket queue in another hemisphere.

Built for

Workloads we
see every day

01

Model training

Multi-node runs on 400G InfiniBand with NCCL tuned and checkpointing to local NVMe.

02

Inference at scale

Autoscaling endpoints with scale-to-zero, so idle traffic costs nothing overnight.

03

Fine-tuning

Single-node A100 and L40S pods for LoRA and full fine-tunes, priced for iteration.

04

Render and simulation

RTX 6000 Ada farms for VFX, CAD and CFD, with shared project volumes.

Capacity is live now

Start on one GPU.
Grow to a cluster.

On-demand needs a card and an email. Reserved capacity needs a conversation, usually a short one.