← Back to the project

Evidence file · approved

GPU timeshare — failure-injection inventory

A redacted, statically counted inventory of failure-injection and incident-scenario coverage around the single-GPU scheduler.

Review state
Owner-approved artifact
Reviewed
7 Aug 2026
Record
gpu-timeshare-failure-injection

Published evidence artifact — approved as-is by the owner, 2026-08-07

Substantiates (portfolio metrics):

  • gpu-timeshare failure-injection — “failure-injection tests”, including incident-scenario suites
  • Supports the case-study claim: “The test suite drills the ugly paths: consumers that lie about yielding, processes that die mid-pause, leases that outlive their owners.”

Where the source data lives (generic): the test tree of the private GPU-coordinator repository. This inventory was produced by static parsing of test files (function names and docstrings); nothing was executed.

Redactions applied

  • Removed all filesystem paths; test files are named by basename only.
  • Redacted the name of one undocumented private consumer service that appears in an incident docstring (described as “a resident vision-LLM consumer”). The life-coach consumer is named because it is documented on the portfolio.
  • No hostnames, ports, or service endpoints are reproduced; incident dates and internal job/run numbers are kept as evidence of real-incident provenance.
  • Test function names and docstring content are otherwise quoted or closely paraphrased from source.

Suite totals (statically counted, not run)

  • 81 test files, 944 test functions in the coordinator test tree.
  • The failure-injection / incident-scenario subset inventoried below: 13 files, 103 test functions. (Adjacent OOM-observability and recovery suites exist beyond this subset and are not itemized here.)
  • Multiple suites are regression-pinned to dated production incidents (2026-06-26, 2026-06-28, 2026-07-08, 2026-07-09, 2026-07-25, 2026-07-27), i.e. the scenarios were injected because they actually happened once.

Categorized inventory

1. Drain races — “released” is not “drained”

test_vram_drain_race.py (3), plus the drain-race case in test_v2_incident_scenarios.py. Scenario: a predecessor releases its lease but its CUDA context has not torn down, so device telemetry still shows the VRAM occupied.

  • test_ack_slot_rejected_when_vram_not_drained — a successor’s slot commit must be rejected (vram_not_drained) when physical VRAM is still occupied by a just-released predecessor, even though the ledger row is gone (the TOCTOU window).
  • test_ack_slot_admitted_when_vram_drained — control: real drain admits normally.
  • test_ack_slot_rejected_when_telemetry_down — fail closed: if GPU telemetry is unavailable the physical re-check cannot run, so admission is refused rather than admitting blind.
  • test_v2_drain_race_vacating_lease_blocks_new_dispatch — a lease marked “vacating” still counts as reserved VRAM until actually released; a dispatch needing that VRAM is refused, not double-booked.

2. Consumers that lie or collide — process identity as proof

test_pid_conflict.py (13) and the escalation case in test_v2_incident_scenarios.py. Scenario (2026-06-28 incident): two unrelated processes share an idempotent lease and both report a PID; the second silently overwrote the first, leaving real GPU work in an unregistered tree the enforcer then killed.

  • Unrelated live PID reporting second gets a 409 conflict; the winner keeps the lease (test_unrelated_pid_rejected_when_existing_alive, test_conflict_does_not_overwrite_existing_pid).
  • Overwrite is allowed only with proof: the prior holder is dead (start-time mismatch) or the new PID is a descendant of the reporter (test_dead_existing_pid_overwrite_allowed, test_descendant_pid_overwrite_allowed), with /proc ancestry checks covered down to “PID 1 is everyone’s ancestor”.
  • test_v2_warn_kill_escalation_requires_grace_and_process_identity — the warn-then-kill chain must carry explicit grace proof, and the destructive stage fails closed without a reported (pid, starttime), so no unrelated process can receive the signal.

3. Death mid-pause and thaw failure — never claim “thawed”

test_v2_vacating_reap_and_thaw_breaker.py (33) and test_v2_cuda_checkpoint_timeout.py (10). Scenarios from a 2026-07-09 overnight triple defect (a pid-less vacating lease from the life-coach consumer blocked every dispatch from 03:22 to 08:00) and a 2026-07-27 stranded-VRAM incident (a resident vision-LLM consumer’s 8 GB checkpoint restore hit a flat 15 s timeout and was re-frozen while still holding VRAM).

  • Pid-less or dead-pid vacating leases are reaped after a grace window instead of blocking dispatch forever; live identities are untouched; PID reuse is detected and failed (test_pidless_vacating_reaped_after_grace, test_dead_pid_vacating_reaped, test_pid_reuse_is_failed).
  • A thaw circuit breaker opens after N consecutive failed physical thaws, deduplicates its events, suppresses further planner-driven thaw churn, and resets on success (test_breaker_opens_after_limit_and_dedupes_events, test_success_resets_the_count).
  • Checkpoint restore budgets scale with the process’s memory footprint (counting swapped pages), are capped, and fall back safely when the footprint is unknown; on restore failure the process is re-frozen — the contract is “never claim thawed” (test_restore_budget_scales_with_footprint, test_resume_still_refreezes_on_restore_failure).

4. Orphaned and ghost leases — leases that outlive their owners

Cases across test_v2_incident_scenarios.py, test_v2_vacating_reap_and_thaw_breaker.py, and test_repair_v2_ghost_credit.py (2).

  • test_v2_dead_pid_release_frees_capacity_for_next_dispatch — a lease whose process has died is released via a dead-PID observation and its VRAM becomes dispatchable, not stuck reserved forever.
  • test_ghost_beneficiary_restores_its_preemption_victim (2026-07-25 incident, jobs 1122/1127) — when a job that financed a preemption turns out to be a ghost, the victim it parked is resumed at the moment the waste is provable; idle parks not tied to the ghost are left alone; the convenience resume still respects the thaw breaker and a kill switch.
  • test_ghost_quota_slice_invents_no_credit and the guarded one-time repair tests — ghost slices earn no accounting credit, and the historical repair is audited, identity-checked, and refuses a second run.

5. Heartbeat pathologies — flapping, stale, and lost

test_heartbeat_lifecycle.py (8) and test_v2_heartbeat_hardening.py (6). Scenario (2026-06-26 bug): heartbeat threads exited on any HTTP 404, so a force-released-and-requeued allocation orphaned its client.

  • Heartbeats survive non-terminal 404s and exit only on terminal states, across all three client lease types; the slice loop clears the lost-event flag so one lost slice cannot kill the next.
  • test_v2_stale_heartbeat_churn_does_not_corrupt_schedule_state — a flapping heartbeat (stale/fresh/stale/fresh) must not accumulate duplicate active leases, crash replanning, or leave inconsistent state.
  • Hardening: monotonic timestamps, bounded retry of database-busy collisions, a retryable 503 after the busy budget, and stable bounded heartbeat jitter.

6. OOM classification and inferred-OOM backoff

test_oom_detection.py (4) and test_v2_inferred_oom_backoff.py (5).

  • The wrapper classifies CUDA-OOM signatures in child stderr and emits a typed event with request, PID, and attempted VRAM — observational only, the child’s exit code propagates unchanged; the signature is the signal even when a framework catches OOM and exits 0.
  • The inferred-OOM damper (2026-08-03 finding): a client that OOMs, cancels, and resubmits with escalated VRAM never trips explicit OOM counters, so the streak is derived from durable job/lease history — three escalating crash cycles back off; a long successful run, a reused VRAM figure, or a stale streak releases the damper.

7. Client give-up and missed-deadline policy

test_gpu_client_giveup.py (8) and test_v2_deadline_missed_policy.py (5).

  • Give-up contract (2026-07-08 incident): a bounded caller must cancel its live queue entry before raising (an abandoned grant burns a 120 s ghost slot), and the terminal state is cancelled, never a fake release; polling survives transient outages and raises only after grace expires.
  • Missed-deadline policy (owner decision 2026-07-09): a deadline job whose window is blown runs best-effort late, appended after every on-time job so running late never causes another miss; an explicit strictness flag is rejected at admission instead of silently dropped.

8. Migration and repair guards

test_migrate_v2_restore_uncertain.py (2) — the partial-restore attribution migration is registered, adds its columns, and backfills the latest failed physical thaw, so uncertain restores stay visible across schema evolution.