$ ./ai-factory/run-ai

Run:ai & Slurm Orchestration_

Most GPU clusters run under half utilized: capital sitting idle. NVIDIA Run:ai and Slurm give every team compute, guarantee production its share, and lend idle cycles out automatically.

// CLUSTER 01 · 32x GPU · LIVE ALLOCATION
aaaabbaaccbbbccc½baa½½
research · quota 10production · quota 8data-sci · quota 6lendable capacityutilization 64% · queue 0 · idle lent out automatically

What's in scope

$ runai queue --cluster 01
01

Fractional GPU Sharing

Slice GPUs across notebooks and light inference so experimentation never monopolizes hardware.

[active]
02

Quotas & Fair-Share Scheduling

Guaranteed allocations per team and project, with idle capacity lent out automatically.

[active]
03

Slurm & Batch Scheduling

HPC-style batch queues for training and research, run alongside Run:ai on the same clusters.

[active]
04

Cluster Consolidation

One orchestration layer across DGX systems, cloud GPUs, and existing Kubernetes estates.

[active]
05

Utilization Analytics

Per-team and per-project visibility into who uses what, and what it costs.

[active]
06

Day 2 Co-Administration

The scheduler is run and tuned alongside your operations team long after go-live.

[active]

Have a project in mind?_

Talk to an engineer →