← All projects Case study Operational

GPU timeshare

Multi-tenant scheduling for one RTX 3090

A constraint-aware scheduler for one RTX 3090 — a durable 24-hour plan, dispatch against physically free memory, and preemption that has to be earned by a committed beneficiary.

2026 Sole designer, engineer, and operator In production — plans and enforces every GPU consumer on the host

card capacity00:0024:00
Cover — drawn from the record. Stacked declared VRAM through 2026-08-31 (UTC), ten-minute buckets, against the 24,576 MiB card.
  1. Problem

    Speech, vision, language, and benchmark workloads all want one 24 GB GPU. A process that says it has released memory may still hold it, and a queue of first-come requests cannot keep deadlines or always-on services honest.

  2. My contribution

    I designed, built, and operate the coordinator — the planner, the dispatcher, the park/thaw protocol, enforcement, metrics, and the incident corrections — and migrated every consumer on the host onto it.

  3. Outcome

    After the 28 August 2026 rollout, 2,771 of 2,835 deadline jobs finished on time. The two failures it exposed became specific corrections — one guarded by a new alert — rather than retries.

Figure — drawn from the record

One day on one GPU — 31 August 2026 (UTC): stacked declared VRAM by workload family, with preemption, park, and thaw events. The full data is in the transcript below. 0 GiB 8 GiB 16 GiB 24 GiB RTX 3090 · 24.0 GiB 00:0003:0006:0009:0012:0015:0018:0021:0024:00 preempt park thaw
  • Always-on services
  • Image generation
  • Speech recognition
  • OCR
  • Embeddings
  • Other batch work
One day on one GPU — 31 August 2026 (UTC). 1,275 slices started; 35 preemptions prepared, 17 parks, 49 thaws. Peak stacked declared VRAM 21.5 GiB of 24.0 GiB. Declared (planner-reserved) VRAM, time-weighted per bucket; not NVML-observed usage.
Figure data as a table
Mean declared VRAM (GiB) per three-hour window, 2026-08-31 UTC
WindowAlways-on servicesImage generationSpeech recognitionOCREmbeddingsOther batch work
00:00–03:001.43.53.00.30.14.9
03:00–06:001.06.10.40.61.60.0
06:00–09:000.08.71.40.50.06.0
09:00–12:000.04.25.60.00.28.9
12:00–15:000.00.93.70.01.90.0
15:00–18:000.01.22.60.01.82.1
18:00–21:000.06.40.40.02.30.0
21:00–24:000.013.50.00.00.00.0

Slices started by family: Always-on services 0 · Image generation 333 · Speech recognition 508 · OCR 334 · Embeddings 44 · Other batch work 56. Source: Timeshare Coordinator v2 ledger (v2_slices, v2_actions), extracted read-only by scripts/evidence/timeshare_day_figure.py; extracted 2026-09-27.

The central decision is to keep three truths separate: intent (jobs, requirements, and leases survive restarts), the plan (a persisted 24-hour schedule), and dispatch, which trusts neither and admits work only against directly observed free memory. A promise to release VRAM is not proof that it is free.

Design

  • Say why a job deserves the GPU. Jobs declare finite work, a periodic guarantee, best-effort opportunity, or always-on availability; one eligibility model feeds planning, admission, preemption, and metrics.
  • Plan a day, keep every version. Each replan writes an immutable schedule: deadlines first (Moore–Hodgson-style selection when not everything fits), then guarantees and opportunity by remaining debt.
  • Earn preemption. A running job is interrupted only for a beneficiary that owns a live slot in the persisted plan, rechecked before acting.
  • Treat physical changes as protocols. Parking (a CUDA checkpoint to host RAM), thawing, and termination are durable actions with recorded attempts, so a crash leaves evidence instead of a silent side effect.
  • Keep a kill net. A separate audit service owns unapproved GPU processes; 31 Prometheus rules watch the loops, deadlines, and write contention.

What operation taught

The soak after the 28 August rollout found two real failures. Heartbeat errors clustered around physical actions because a record held SQLite’s writer lock through the whole checkpoint call; it now commits first, and an alert guards the fix. One thaw timed out with about 11.9 GiB still resident because its restore budget came from memory at thaw time; parking now records the pre-checkpoint footprint as a floor. A follow-up removed a deadlock: a later queued job was reserved ahead of the interrupted job whose partial restore held the memory it needed.

Concurrent packing looked perfect in shadow — 6,003 schedules, zero errors — and I kept it out of enforcement until recovery, preemption, and activation share one capacity model.

Why it matters

Separate what was promised, what was planned, and what physically happened, and check each independently: a scheduler is only as trustworthy as its account of where they disagree.

Built with

  • Python
  • FastAPI
  • SQLite
  • systemd
  • NVML
  • cuda-checkpoint
  • Prometheus