← All projects Case study Operational
GPU timeshare
Multi-tenant scheduling for one RTX 3090
A constraint-aware scheduler for one RTX 3090 — a durable 24-hour plan, dispatch against physically free memory, and preemption that has to be earned by a committed beneficiary.
Problem
Speech, vision, language, and benchmark workloads all want one 24 GB GPU. A process that says it has released memory may still hold it, and a queue of first-come requests cannot keep deadlines or always-on services honest.
My contribution
I designed, built, and operate the coordinator — the planner, the dispatcher, the park/thaw protocol, enforcement, metrics, and the incident corrections — and migrated every consumer on the host onto it.
Outcome
After the 28 August 2026 rollout, 2,771 of 2,835 deadline jobs finished on time. The two failures it exposed became specific corrections — one guarded by a new alert — rather than retries.
Figure — drawn from the record
- Always-on services
- Image generation
- Speech recognition
- OCR
- Embeddings
- Other batch work
Figure data as a table
| Window | Always-on services | Image generation | Speech recognition | OCR | Embeddings | Other batch work |
|---|---|---|---|---|---|---|
| 00:00–03:00 | 1.4 | 3.5 | 3.0 | 0.3 | 0.1 | 4.9 |
| 03:00–06:00 | 1.0 | 6.1 | 0.4 | 0.6 | 1.6 | 0.0 |
| 06:00–09:00 | 0.0 | 8.7 | 1.4 | 0.5 | 0.0 | 6.0 |
| 09:00–12:00 | 0.0 | 4.2 | 5.6 | 0.0 | 0.2 | 8.9 |
| 12:00–15:00 | 0.0 | 0.9 | 3.7 | 0.0 | 1.9 | 0.0 |
| 15:00–18:00 | 0.0 | 1.2 | 2.6 | 0.0 | 1.8 | 2.1 |
| 18:00–21:00 | 0.0 | 6.4 | 0.4 | 0.0 | 2.3 | 0.0 |
| 21:00–24:00 | 0.0 | 13.5 | 0.0 | 0.0 | 0.0 | 0.0 |
Slices started by family: Always-on services 0 · Image generation 333 · Speech recognition 508 · OCR 334 · Embeddings 44 · Other batch work 56. Source: Timeshare Coordinator v2 ledger (v2_slices, v2_actions), extracted read-only by scripts/evidence/timeshare_day_figure.py; extracted 2026-09-27.
The central decision is to keep three truths separate: intent (jobs, requirements, and leases survive restarts), the plan (a persisted 24-hour schedule), and dispatch, which trusts neither and admits work only against directly observed free memory. A promise to release VRAM is not proof that it is free.
Design
- Say why a job deserves the GPU. Jobs declare finite work, a periodic guarantee, best-effort opportunity, or always-on availability; one eligibility model feeds planning, admission, preemption, and metrics.
- Plan a day, keep every version. Each replan writes an immutable schedule: deadlines first (Moore–Hodgson-style selection when not everything fits), then guarantees and opportunity by remaining debt.
- Earn preemption. A running job is interrupted only for a beneficiary that owns a live slot in the persisted plan, rechecked before acting.
- Treat physical changes as protocols. Parking (a CUDA checkpoint to host RAM), thawing, and termination are durable actions with recorded attempts, so a crash leaves evidence instead of a silent side effect.
- Keep a kill net. A separate audit service owns unapproved GPU processes; 31 Prometheus rules watch the loops, deadlines, and write contention.
What operation taught
The soak after the 28 August rollout found two real failures. Heartbeat errors clustered around physical actions because a record held SQLite’s writer lock through the whole checkpoint call; it now commits first, and an alert guards the fix. One thaw timed out with about 11.9 GiB still resident because its restore budget came from memory at thaw time; parking now records the pre-checkpoint footprint as a floor. A follow-up removed a deadlock: a later queued job was reserved ahead of the interrupted job whose partial restore held the memory it needed.
Concurrent packing looked perfect in shadow — 6,003 schedules, zero errors — and I kept it out of enforcement until recovery, preemption, and activation share one capacity model.
Why it matters
Separate what was promised, what was planned, and what physically happened, and check each independently: a scheduler is only as trustworthy as its account of where they disagree.
Built with
- Python
- FastAPI
- SQLite
- systemd
- NVML
- cuda-checkpoint
- Prometheus