· Dispatch AI

Five Bugs, Two Passes, Zero Regressions — GCF 2.0 Bug Fix Stages 29.1 and 29.2

 ·  Billy p

By Billy P. — 7 min read · Part 3 of the GCF 2.0 series

The GCF 2.0 integration walk-through left us with three open defects. Then a closer look at the background worker turned up two more. Two short stages, five bugs, and one promise to the rest of the codebase: don’t touch the frozen stages. This is the story of how we closed them.

Here’s the punchline up front:

TL;DR. Stage 29.1 fixed the three known defects (background cancellation ignored, dropped result receiver, counter underflow) by routing every terminal state through a single finalize_background_task helper. Stage 29.2 found two more during a deliberate bug sweep and fixed them with overflow-safe revision arithmetic and a Drop implementation that drains remaining tasks cleanly. Both stages passed the full 6-cargo release gate, kept the regression ceiling under 5%, and left every frozen stage byte-for-byte identical.


What We Were Fixing

The stage 29.0 walk-through had already logged three known defects and marked them as KNOWN-DEFECT in the E2E matrix. For a quick reminder, they were:

#SeverityDefectWhere
1MEDIUMBackground worker never checked cancel_token.is_cancelled(); ran strategies to completion even on cancellation.hierarchy.rs background loop
2CRITICALResult receiver was dropped at task creation, so worker try_send saw Disconnected and silently discarded results.hierarchy.rs::enqueue_background_task_with_plan
3LOWForeground tasks (stages 18–27) had no cancellation hooks at all.throughout

Two more turned up during the stage 29.2 bug sweep:

#SeverityDefectWhere
4HIGHBackground worker never compared the queued task’s expected_world_revision to the live WorldModel::global_revision, so stale tasks executed on outdated state.hierarchy.rs::spawn_worker_thread
5MEDIUMWhen a HierarchicalManager was dropped or shut down, remaining tasks in the queue and yielded slot were discarded without decrementing BACKGROUND_WORK_COUNT, leaving the global counter elevated.HierarchicalManager::Drop

Let me walk through how we closed them.


Stage 29.1 — The Three Original Defects

Fix 1: Canonical result ownership

The dropped-receiver bug was the most dangerous of the three: results were being silently lost on every task. The fix is to make BoundedResultBuffer the canonical retrievable result store and treat SyncSender as an optional immediate-delivery path only.

Three rules:

  1. Always write every terminal result exactly once to BoundedResultBuffer, regardless of SyncSender connectivity or full status.
  2. Keep the buffer at its existing capacity (100) with circular FIFO eviction and non-blocking behavior. No new registry, no unbounded channels.
  3. get_background_result() queries only the canonical store.

Fix 2: Explicit try_send handling

Every try_send outcome is now handled explicitly. No more let _ = try_send; swallowing errors:

rust

match task.payload.result_tx.try_send(res.clone()) {
    Ok(_) => { /* notification succeeded */ }
    Err(TrySendError::Full(_)) => { /* channel full; log bounded audit event */ }
    Err(TrySendError::Disconnected(_)) => { /* receiver dropped; log bounded audit event */ }
}
// In all three cases, the canonical buffer still holds the result.

Fix 3: One finalization helper, one decrement

The real fix for both the cancellation bug and the counter-underflow bug is the same: route every terminal state through a single helper.

rust

fn finalize_background_task(
    task_id: BackgroundTaskId,
    input_query: &str,
    state: TaskLifecycleState,
    error_msg: String,
    response_text: String,
    result_tx: &SyncSender<BackgroundResult>,
    result_buffer: &Arc<Mutex<BoundedResultBuffer>>,
) {
    let res = BackgroundResult { /* ... */ };

    // 1. Optional immediate notification
    match result_tx.try_send(res.clone()) { /* ... */ }

    // 2. Canonical buffer storage
    let mut buf = result_buffer.lock().unwrap();
    buf.push(res);

    // 3. Decrement active task counter safely (loop-based CAS, no underflow)
    let mut current = BACKGROUND_WORK_COUNT.load(Ordering::SeqCst);
    loop {
        if current == 0 { break; }
        match BACKGROUND_WORK_COUNT.compare_exchange_weak(current, current - 1, ...) {
            Ok(_) => break,
            Err(actual) => current = actual,
        }
    }
}

Every Completed, Failed, Expired, and Cancelled path now goes through this single function. Counter decrement happens exactly once. Cancellation is honored exactly when we want it to be — at one of five explicit checkpoints:

  1. Right after popping the task from the queue
  2. Before resuming from the yielded slot
  3. Before each new face strategy increment
  4. Before each recursive child (already handled natively in stage 26)
  5. Right after the blocking model/HTTP call returns

Note on defect 3 (foreground cancellation): It’s a stage 18–27 baseline limitation. Stage 29.1 documents it as test_foreground_lack_of_cancellation_documented and explicitly marks it as a known GCF 2.0 limitation, not a defect to fix here.

Stage 29.1 Results

All six cargo release gates passed. 67/67 tests (the previous 65 plus two promoted regression tests). Stress test over 10,000 operations: 0 KB working-set growth, 9,647.9 req/s throughput, 0 counter underflows, 0 panics, 0 deadlocks.

Stage 18 regression check (60 inputs, 25 reps, thermal-balanced):

ConfigurationMean (ms)Regression
Stage 18 baseline5.3716—
Stage 29.1 disabled5.6139+4.51% — PASS
Stage 29.1 enabled-empty5.5865+4.00% — PASS
Stage 29.1 active-background4.9847−7.20% — PASS

That −7.20% on active-background isn’t a mistake — fixing the dropped-receiver bug eliminated a real source of wasted worker time, so the active-background path is now measurably faster than the original.


Stage 29.2 — The Two New Defects

Once stage 29.1 was in, we did a deliberate bug sweep. Two more defects surfaced.

Defect GCF-29.2-01 (HIGH) — Stale background WorldModel revision

The background worker was happily executing tasks against a world view that might have moved on. The fix is four explicit revalidation checkpoints and overflow-safe arithmetic.

rust

// overflow-safe: no underflow if expected > current
let delta = current.saturating_sub(expected);
if delta > max_acceptable_revision_skew {
    // transition to NeedsRevalidation and finalize
}

The four checkpoints (same shape as the cancellation ones):

  1. Immediately after queue pop
  2. Before yielded resume
  3. Before each subsequent face strategy increment
  4. After the blocking model/HTTP call returns

If take_snapshot() itself fails, the task fails closed: transition to NeedsRevalidation and finalize via the same finalize_background_task helper from stage 29.1, so the counter still decrements exactly once.

Defect GCF-29.2-02 (MEDIUM) — Counter leak on manager drop

When a HierarchicalManager is dropped (or shutdown is called), any tasks still in its queue or yielded slot were getting thrown away without decrementing BACKGROUND_WORK_COUNT. The fix is a proper Drop implementation:

rust

impl Drop for HierarchicalManager {
    fn drop(&mut self) {
        // 1. signal shutdown
        // 2. wake the condvar
        // 3. join the worker thread (but not if we ARE the worker — self-join panic guard)
        // 4. drain background_queue, finalizing each task
        // 5. drain yielded_slot, finalizing the one task there
    }
}

Three subtle things matter here:

  • Task-local decrements, not global resets. Multi-manager isolation is preserved: manager A’s drop doesn’t zero manager B’s counts.
  • Self-join guard. The drop thread checks whether it’s the worker thread before joining worker_handle, preventing a self-join panic.
  • No new locks. Network and model calls still happen outside structural locks.

Stage 29.2 Results

All six cargo release gates passed. 74/74 tests (67 from stage 29.1 plus 7 new tests: test_stale_background_revision_revalidated, test_future_revision_revalidated, test_revision_within_skew_executes, test_manager_drop_accounting_cleanup, test_multi_manager_counter_isolation, test_stale_yielded_revision_revalidated, test_revision_becomes_stale_between_faces).

Stress test over 10,000 operations: 0 KB working-set growth, 0 panics, 0 deadlocks.

Stage 18 regression check:

ConfigurationMean (ms)Regression
Stage 18 baseline4.6736—
Stage 29.2 disabled4.8844+4.51% — PASS
Stage 29.2 enabled-empty4.6887+0.32% — PASS
Stage 29.2 active-background4.6524−0.46% — PASS

What Stayed Frozen

The whole point of these two stages was to fix bugs without disturbing the rest of the codebase. The walkthroughs explicitly confirm zero semantic changes to:

  • Stage 22 Planner
  • Stage 23 Governance
  • Stage 24 WorldModel
  • Stage 25 Prediction
  • Stage 26 Recursion Budget / RAII
  • Stage 27 Economy
  • Stage 28 Worker / Queue / Yield bounds (the structural mechanics, not the new helper methods)

The only files modified are src/cube/recursive/hierarchy.rs (the finalization helper, the revalidation closures, and the Drop impl) and src/cube/recursive/integration.rs (the new tests). Nothing else was touched.


The Bug Sweep Reconciliation

Beyond the two new defects, the stage 29.2 sweep ran a structured audit across the rest of the integration surface:

AreaResult
Stage 27 economy boundary (Stop never suppresses mandatory governance)PASS
Persistent cognitive state (no invalid success writes on failure paths)PASS
Feature combinations (no unwraps, no worker spawn when background is off)PASS
IntegrationError mapping and 256-byte clampingPASS
Correlation-log integrity (100-capacity circular FIFO)PASS
Lock and deadlock audit (no model calls under structural locks)PASS
Panic and poison recovery (counter decremented correctly)PASS
Shutdown audit (clean JoinHandle waits across lifecycles)PASS

No further defects found.


What’s Next

The remaining open item is the LOW-severity foreground cancellation (stage 18–27 baseline). Documented as test_foreground_lack_of_cancellation_documented in stage 29.1, it’s not on the critical path. The cleanest place to add it is as a separate “stage 30 — foreground cancellation” pass that doesn’t risk colliding with anything we just shipped.

For now, GCF 2.0 has 74/74 tests passing, zero counter underflows, zero lost background results, zero stale-revision executions, and zero regressions. That’s a fine place to be.

— Billy P.


Tested on branch=2.0, last verified August 2026. Cover diagram: 6-cargo release gate + 5-defect fix summary.

Series navigation:

  • Part 1: Building a Bounded, Governance-Aware Cognitive Framework in Rust (Stages 24–26) — post 014
  • Part 2: Building a Cognitive Control Plane in Rust (Stages 27, 28, 29.0) — post 015
  • Part 3: Five Bugs, Two Passes, Zero Regressions (this post — Stages 29.1, 29.2) — post 016
  • Part 4: GCF 2.0 Final — From 250,000 Operations to a Validated Release Candidate (Stages 29.3, 29.4, 30) — post 017
  • Part 5: ISAC Is Now Under Development on GCF 2.0 (the series finale) — post 018
← All dispatches Homepage →