By Billy P. — 7 min read · Part 3 of the GCF 2.0 series
The GCF 2.0 integration walk-through left us with three open defects. Then a closer look at the background worker turned up two more. Two short stages, five bugs, and one promise to the rest of the codebase: don’t touch the frozen stages. This is the story of how we closed them.
Here’s the punchline up front:
TL;DR. Stage 29.1 fixed the three known defects (background cancellation ignored, dropped result receiver, counter underflow) by routing every terminal state through a single
finalize_background_taskhelper. Stage 29.2 found two more during a deliberate bug sweep and fixed them with overflow-safe revision arithmetic and aDropimplementation that drains remaining tasks cleanly. Both stages passed the full 6-cargo release gate, kept the regression ceiling under 5%, and left every frozen stage byte-for-byte identical.
What We Were Fixing
The stage 29.0 walk-through had already logged three known defects and marked them as KNOWN-DEFECT in the E2E matrix. For a quick reminder, they were:
| # | Severity | Defect | Where |
|---|---|---|---|
| 1 | MEDIUM | Background worker never checked cancel_token.is_cancelled(); ran strategies to completion even on cancellation. | hierarchy.rs background loop |
| 2 | CRITICAL | Result receiver was dropped at task creation, so worker try_send saw Disconnected and silently discarded results. | hierarchy.rs::enqueue_background_task_with_plan |
| 3 | LOW | Foreground tasks (stages 18–27) had no cancellation hooks at all. | throughout |
Two more turned up during the stage 29.2 bug sweep:
| # | Severity | Defect | Where |
|---|---|---|---|
| 4 | HIGH | Background worker never compared the queued task’s expected_world_revision to the live WorldModel::global_revision, so stale tasks executed on outdated state. | hierarchy.rs::spawn_worker_thread |
| 5 | MEDIUM | When a HierarchicalManager was dropped or shut down, remaining tasks in the queue and yielded slot were discarded without decrementing BACKGROUND_WORK_COUNT, leaving the global counter elevated. | HierarchicalManager::Drop |
Let me walk through how we closed them.
Stage 29.1 — The Three Original Defects
Fix 1: Canonical result ownership
The dropped-receiver bug was the most dangerous of the three: results were being silently lost on every task. The fix is to make BoundedResultBuffer the canonical retrievable result store and treat SyncSender as an optional immediate-delivery path only.
Three rules:
- Always write every terminal result exactly once to
BoundedResultBuffer, regardless ofSyncSenderconnectivity or full status. - Keep the buffer at its existing capacity (100) with circular FIFO eviction and non-blocking behavior. No new registry, no unbounded channels.
get_background_result()queries only the canonical store.
Fix 2: Explicit try_send handling
Every try_send outcome is now handled explicitly. No more let _ = try_send; swallowing errors:
rust
match task.payload.result_tx.try_send(res.clone()) {
Ok(_) => { /* notification succeeded */ }
Err(TrySendError::Full(_)) => { /* channel full; log bounded audit event */ }
Err(TrySendError::Disconnected(_)) => { /* receiver dropped; log bounded audit event */ }
}
// In all three cases, the canonical buffer still holds the result.
Fix 3: One finalization helper, one decrement
The real fix for both the cancellation bug and the counter-underflow bug is the same: route every terminal state through a single helper.
rust
fn finalize_background_task(
task_id: BackgroundTaskId,
input_query: &str,
state: TaskLifecycleState,
error_msg: String,
response_text: String,
result_tx: &SyncSender<BackgroundResult>,
result_buffer: &Arc<Mutex<BoundedResultBuffer>>,
) {
let res = BackgroundResult { /* ... */ };
// 1. Optional immediate notification
match result_tx.try_send(res.clone()) { /* ... */ }
// 2. Canonical buffer storage
let mut buf = result_buffer.lock().unwrap();
buf.push(res);
// 3. Decrement active task counter safely (loop-based CAS, no underflow)
let mut current = BACKGROUND_WORK_COUNT.load(Ordering::SeqCst);
loop {
if current == 0 { break; }
match BACKGROUND_WORK_COUNT.compare_exchange_weak(current, current - 1, ...) {
Ok(_) => break,
Err(actual) => current = actual,
}
}
}
Every Completed, Failed, Expired, and Cancelled path now goes through this single function. Counter decrement happens exactly once. Cancellation is honored exactly when we want it to be — at one of five explicit checkpoints:
- Right after popping the task from the queue
- Before resuming from the yielded slot
- Before each new face strategy increment
- Before each recursive child (already handled natively in stage 26)
- Right after the blocking model/HTTP call returns
Note on defect 3 (foreground cancellation): It’s a stage 18–27 baseline limitation. Stage 29.1 documents it as
test_foreground_lack_of_cancellation_documentedand explicitly marks it as a known GCF 2.0 limitation, not a defect to fix here.
Stage 29.1 Results
All six cargo release gates passed. 67/67 tests (the previous 65 plus two promoted regression tests). Stress test over 10,000 operations: 0 KB working-set growth, 9,647.9 req/s throughput, 0 counter underflows, 0 panics, 0 deadlocks.
Stage 18 regression check (60 inputs, 25 reps, thermal-balanced):
| Configuration | Mean (ms) | Regression |
|---|---|---|
| Stage 18 baseline | 5.3716 | — |
| Stage 29.1 disabled | 5.6139 | +4.51% — PASS |
| Stage 29.1 enabled-empty | 5.5865 | +4.00% — PASS |
| Stage 29.1 active-background | 4.9847 | −7.20% — PASS |
That −7.20% on active-background isn’t a mistake — fixing the dropped-receiver bug eliminated a real source of wasted worker time, so the active-background path is now measurably faster than the original.
Stage 29.2 — The Two New Defects
Once stage 29.1 was in, we did a deliberate bug sweep. Two more defects surfaced.
Defect GCF-29.2-01 (HIGH) — Stale background WorldModel revision
The background worker was happily executing tasks against a world view that might have moved on. The fix is four explicit revalidation checkpoints and overflow-safe arithmetic.
rust
// overflow-safe: no underflow if expected > current
let delta = current.saturating_sub(expected);
if delta > max_acceptable_revision_skew {
// transition to NeedsRevalidation and finalize
}
The four checkpoints (same shape as the cancellation ones):
- Immediately after queue pop
- Before yielded resume
- Before each subsequent face strategy increment
- After the blocking model/HTTP call returns
If take_snapshot() itself fails, the task fails closed: transition to NeedsRevalidation and finalize via the same finalize_background_task helper from stage 29.1, so the counter still decrements exactly once.
Defect GCF-29.2-02 (MEDIUM) — Counter leak on manager drop
When a HierarchicalManager is dropped (or shutdown is called), any tasks still in its queue or yielded slot were getting thrown away without decrementing BACKGROUND_WORK_COUNT. The fix is a proper Drop implementation:
rust
impl Drop for HierarchicalManager {
fn drop(&mut self) {
// 1. signal shutdown
// 2. wake the condvar
// 3. join the worker thread (but not if we ARE the worker — self-join panic guard)
// 4. drain background_queue, finalizing each task
// 5. drain yielded_slot, finalizing the one task there
}
}
Three subtle things matter here:
- Task-local decrements, not global resets. Multi-manager isolation is preserved: manager A’s drop doesn’t zero manager B’s counts.
- Self-join guard. The drop thread checks whether it’s the worker thread before joining
worker_handle, preventing a self-join panic. - No new locks. Network and model calls still happen outside structural locks.
Stage 29.2 Results
All six cargo release gates passed. 74/74 tests (67 from stage 29.1 plus 7 new tests: test_stale_background_revision_revalidated, test_future_revision_revalidated, test_revision_within_skew_executes, test_manager_drop_accounting_cleanup, test_multi_manager_counter_isolation, test_stale_yielded_revision_revalidated, test_revision_becomes_stale_between_faces).
Stress test over 10,000 operations: 0 KB working-set growth, 0 panics, 0 deadlocks.
Stage 18 regression check:
| Configuration | Mean (ms) | Regression |
|---|---|---|
| Stage 18 baseline | 4.6736 | — |
| Stage 29.2 disabled | 4.8844 | +4.51% — PASS |
| Stage 29.2 enabled-empty | 4.6887 | +0.32% — PASS |
| Stage 29.2 active-background | 4.6524 | −0.46% — PASS |
What Stayed Frozen
The whole point of these two stages was to fix bugs without disturbing the rest of the codebase. The walkthroughs explicitly confirm zero semantic changes to:
- Stage 22 Planner
- Stage 23 Governance
- Stage 24 WorldModel
- Stage 25 Prediction
- Stage 26 Recursion Budget / RAII
- Stage 27 Economy
- Stage 28 Worker / Queue / Yield bounds (the structural mechanics, not the new helper methods)
The only files modified are src/cube/recursive/hierarchy.rs (the finalization helper, the revalidation closures, and the Drop impl) and src/cube/recursive/integration.rs (the new tests). Nothing else was touched.
The Bug Sweep Reconciliation
Beyond the two new defects, the stage 29.2 sweep ran a structured audit across the rest of the integration surface:
| Area | Result |
|---|---|
| Stage 27 economy boundary (Stop never suppresses mandatory governance) | PASS |
| Persistent cognitive state (no invalid success writes on failure paths) | PASS |
| Feature combinations (no unwraps, no worker spawn when background is off) | PASS |
| IntegrationError mapping and 256-byte clamping | PASS |
| Correlation-log integrity (100-capacity circular FIFO) | PASS |
| Lock and deadlock audit (no model calls under structural locks) | PASS |
| Panic and poison recovery (counter decremented correctly) | PASS |
Shutdown audit (clean JoinHandle waits across lifecycles) | PASS |
No further defects found.
What’s Next
The remaining open item is the LOW-severity foreground cancellation (stage 18–27 baseline). Documented as test_foreground_lack_of_cancellation_documented in stage 29.1, it’s not on the critical path. The cleanest place to add it is as a separate “stage 30 — foreground cancellation” pass that doesn’t risk colliding with anything we just shipped.
For now, GCF 2.0 has 74/74 tests passing, zero counter underflows, zero lost background results, zero stale-revision executions, and zero regressions. That’s a fine place to be.
— Billy P.
Tested on branch=2.0, last verified August 2026. Cover diagram: 6-cargo release gate + 5-defect fix summary.
Series navigation:
- Part 1: Building a Bounded, Governance-Aware Cognitive Framework in Rust (Stages 24–26) — post 014
- Part 2: Building a Cognitive Control Plane in Rust (Stages 27, 28, 29.0) — post 015
- Part 3: Five Bugs, Two Passes, Zero Regressions (this post — Stages 29.1, 29.2) — post 016
- Part 4: GCF 2.0 Final — From 250,000 Operations to a Validated Release Candidate (Stages 29.3, 29.4, 30) — post 017
- Part 5: ISAC Is Now Under Development on GCF 2.0 (the series finale) — post 018