Evidence file · approved
GPU timeshare — failure-injection inventory
A redacted, statically counted inventory of failure-injection and incident-scenario coverage around the single-GPU scheduler.
- Review state
- Owner-approved artifact
- Reviewed
- 7 Aug 2026
- Record
- gpu-timeshare-failure-injection
Published evidence artifact — approved as-is by the owner, 2026-08-07
Substantiates (portfolio metrics):
- gpu-timeshare
failure-injection— “failure-injection tests”, including incident-scenario suites- Supports the case-study claim: “The test suite drills the ugly paths: consumers that lie about yielding, processes that die mid-pause, leases that outlive their owners.”
Where the source data lives (generic): the test tree of the private GPU-coordinator repository. This inventory was produced by static parsing of test files (function names and docstrings); nothing was executed.
Redactions applied
- Removed all filesystem paths; test files are named by basename only.
- Redacted the name of one undocumented private consumer service that appears in an incident docstring (described as “a resident vision-LLM consumer”). The life-coach consumer is named because it is documented on the portfolio.
- No hostnames, ports, or service endpoints are reproduced; incident dates and internal job/run numbers are kept as evidence of real-incident provenance.
- Test function names and docstring content are otherwise quoted or closely paraphrased from source.
Suite totals (statically counted, not run)
- 81 test files, 944 test functions in the coordinator test tree.
- The failure-injection / incident-scenario subset inventoried below: 13 files, 103 test functions. (Adjacent OOM-observability and recovery suites exist beyond this subset and are not itemized here.)
- Multiple suites are regression-pinned to dated production incidents (2026-06-26, 2026-06-28, 2026-07-08, 2026-07-09, 2026-07-25, 2026-07-27), i.e. the scenarios were injected because they actually happened once.
Categorized inventory
1. Drain races — “released” is not “drained”
test_vram_drain_race.py (3), plus the drain-race case in
test_v2_incident_scenarios.py. Scenario: a predecessor releases its lease
but its CUDA context has not torn down, so device telemetry still shows the
VRAM occupied.
test_ack_slot_rejected_when_vram_not_drained— a successor’s slot commit must be rejected (vram_not_drained) when physical VRAM is still occupied by a just-released predecessor, even though the ledger row is gone (the TOCTOU window).test_ack_slot_admitted_when_vram_drained— control: real drain admits normally.test_ack_slot_rejected_when_telemetry_down— fail closed: if GPU telemetry is unavailable the physical re-check cannot run, so admission is refused rather than admitting blind.test_v2_drain_race_vacating_lease_blocks_new_dispatch— a lease marked “vacating” still counts as reserved VRAM until actually released; a dispatch needing that VRAM is refused, not double-booked.
2. Consumers that lie or collide — process identity as proof
test_pid_conflict.py (13) and the escalation case in
test_v2_incident_scenarios.py. Scenario (2026-06-28 incident): two
unrelated processes share an idempotent lease and both report a PID; the
second silently overwrote the first, leaving real GPU work in an
unregistered tree the enforcer then killed.
- Unrelated live PID reporting second gets a 409 conflict; the winner keeps
the lease (
test_unrelated_pid_rejected_when_existing_alive,test_conflict_does_not_overwrite_existing_pid). - Overwrite is allowed only with proof: the prior holder is dead
(start-time mismatch) or the new PID is a descendant of the reporter
(
test_dead_existing_pid_overwrite_allowed,test_descendant_pid_overwrite_allowed), with/procancestry checks covered down to “PID 1 is everyone’s ancestor”. test_v2_warn_kill_escalation_requires_grace_and_process_identity— the warn-then-kill chain must carry explicit grace proof, and the destructive stage fails closed without a reported (pid, starttime), so no unrelated process can receive the signal.
3. Death mid-pause and thaw failure — never claim “thawed”
test_v2_vacating_reap_and_thaw_breaker.py (33) and
test_v2_cuda_checkpoint_timeout.py (10). Scenarios from a 2026-07-09
overnight triple defect (a pid-less vacating lease from the life-coach
consumer blocked every dispatch from 03:22 to 08:00) and a 2026-07-27
stranded-VRAM incident (a resident vision-LLM consumer’s 8 GB checkpoint
restore hit a flat 15 s timeout and was re-frozen while still holding VRAM).
- Pid-less or dead-pid vacating leases are reaped after a grace window
instead of blocking dispatch forever; live identities are untouched; PID
reuse is detected and failed (
test_pidless_vacating_reaped_after_grace,test_dead_pid_vacating_reaped,test_pid_reuse_is_failed). - A thaw circuit breaker opens after N consecutive failed physical thaws,
deduplicates its events, suppresses further planner-driven thaw churn, and
resets on success (
test_breaker_opens_after_limit_and_dedupes_events,test_success_resets_the_count). - Checkpoint restore budgets scale with the process’s memory footprint
(counting swapped pages), are capped, and fall back safely when the
footprint is unknown; on restore failure the process is re-frozen — the
contract is “never claim thawed”
(
test_restore_budget_scales_with_footprint,test_resume_still_refreezes_on_restore_failure).
4. Orphaned and ghost leases — leases that outlive their owners
Cases across test_v2_incident_scenarios.py,
test_v2_vacating_reap_and_thaw_breaker.py, and
test_repair_v2_ghost_credit.py (2).
test_v2_dead_pid_release_frees_capacity_for_next_dispatch— a lease whose process has died is released via a dead-PID observation and its VRAM becomes dispatchable, not stuck reserved forever.test_ghost_beneficiary_restores_its_preemption_victim(2026-07-25 incident, jobs 1122/1127) — when a job that financed a preemption turns out to be a ghost, the victim it parked is resumed at the moment the waste is provable; idle parks not tied to the ghost are left alone; the convenience resume still respects the thaw breaker and a kill switch.test_ghost_quota_slice_invents_no_creditand the guarded one-time repair tests — ghost slices earn no accounting credit, and the historical repair is audited, identity-checked, and refuses a second run.
5. Heartbeat pathologies — flapping, stale, and lost
test_heartbeat_lifecycle.py (8) and test_v2_heartbeat_hardening.py (6).
Scenario (2026-06-26 bug): heartbeat threads exited on any HTTP 404, so a
force-released-and-requeued allocation orphaned its client.
- Heartbeats survive non-terminal 404s and exit only on terminal states, across all three client lease types; the slice loop clears the lost-event flag so one lost slice cannot kill the next.
test_v2_stale_heartbeat_churn_does_not_corrupt_schedule_state— a flapping heartbeat (stale/fresh/stale/fresh) must not accumulate duplicate active leases, crash replanning, or leave inconsistent state.- Hardening: monotonic timestamps, bounded retry of database-busy collisions, a retryable 503 after the busy budget, and stable bounded heartbeat jitter.
6. OOM classification and inferred-OOM backoff
test_oom_detection.py (4) and test_v2_inferred_oom_backoff.py (5).
- The wrapper classifies CUDA-OOM signatures in child stderr and emits a typed event with request, PID, and attempted VRAM — observational only, the child’s exit code propagates unchanged; the signature is the signal even when a framework catches OOM and exits 0.
- The inferred-OOM damper (2026-08-03 finding): a client that OOMs, cancels, and resubmits with escalated VRAM never trips explicit OOM counters, so the streak is derived from durable job/lease history — three escalating crash cycles back off; a long successful run, a reused VRAM figure, or a stale streak releases the damper.
7. Client give-up and missed-deadline policy
test_gpu_client_giveup.py (8) and test_v2_deadline_missed_policy.py (5).
- Give-up contract (2026-07-08 incident): a bounded caller must cancel its
live queue entry before raising (an abandoned grant burns a 120 s ghost
slot), and the terminal state is
cancelled, never a fakerelease; polling survives transient outages and raises only after grace expires. - Missed-deadline policy (owner decision 2026-07-09): a deadline job whose window is blown runs best-effort late, appended after every on-time job so running late never causes another miss; an explicit strictness flag is rejected at admission instead of silently dropped.
8. Migration and repair guards
test_migrate_v2_restore_uncertain.py (2) — the partial-restore
attribution migration is registered, adds its columns, and backfills the
latest failed physical thaw, so uncertain restores stay visible across
schema evolution.