By Billy P. — 12 min read
If you’ve ever built a complex reasoning system in Rust, you know the tension: the language wants you to prove every allocation is safe, but a real cognitive framework has to be flexible enough to expand into child work, contract when governance denies a request, and survive a 100,000-operation stress run without leaking a byte.
Over the past few months I’ve been shepherding GCF 2.0 (our General Cognitive Framework) through three of its most demanding stages. This post walks through what we built, what we learned, and why — and it includes enough concrete code and benchmarks that you can apply the patterns to your own systems.
Here’s the punchline up front, before we dive in:
TL;DR. Across three stages we shipped a thread-safe World Model, a deterministic Predictive Cognition layer, and a Recursive Cognitive Expansion engine. All three are bounded, governance-gated, and verified at scale. The hot paths run in fractions of a microsecond and the 100,000-operation stress runs grow memory by less than a megabyte.
The Shape of the System
GCF 2.0 is built as a sequence of stages, each owning a single concern. The early stages handle telemetry, invariance detection, persistent state, adaptive planning, and governance. Stages 24, 25, and 26 — the focus here — close out a coherent phase by adding environmental memory, predictive control, and recursive scaling.
The three stages form a single pipeline:
| # | Stage | Purpose | Outputs |
|---|---|---|---|
| 24 | World Model Engine | Represent the external environment decoupled from internal state | BTreeMap-backed, capacity-bounded world store with atomic persistence |
| 25 | Predictive Cognition | Estimate outcome, confidence, and cost before execution | PredictionResult with explainability factors and per-path estimates |
| 26 | Recursive Cognitive Expansion | Scale hard reasoning via bounded child subcubes | RAII-guarded expansion; deterministic, precedence-ordered result merging |
A few design rules hold across all three:
- Stage 25 is advisory. It never spawns work or terminates execution on its own.
- Stage 22 is the planner. It remains the sole authority for cell allocation and expansion authorization.
- Stage 23 is the gatekeeper. No mutating call into the World Model and no child-cube spawn happens without a
Proceeddecision from governance. - Child cubes are isolated. They receive an in-memory snapshot of the world and are forbidden from writing back.
Let’s look at each stage.
Stage 24: The World Model Engine
Most cognitive systems need a way to represent facts about the outside world — the current state of a UI, the contents of a database, the result of a tool call. You could stuff this into the same store as your internal memory, but you’ll regret it the first time you need to reset the world without resetting your memory.
We built a deterministic, stateful, thread-safe World Model that lives next to (but not inside) the internal cognitive state. The key decisions:
- BTreeMap, not HashMap. Sorted traversal gives us reproducible snapshots and serializations, which matters for tests and for atomic file replacement.
- Explicit ownership. The model is owned by the orchestrator, not a global singleton. This makes unit tests trivial and lifecycle management clean.
- Atomic replacement on disk. JSON is written to
gcf_world_model.json.tmpand then atomically moved into place viastd::fs::rename. The existing file is never half-written. - Pre-validated load. Before swapping live state, we deserialize into a temporary
WorldSnapshotand verify every entry fits the configured capacity and size limits. If the load fails, the live state is left untouched. - Proceed-only governance. The core model knows nothing about policy. The orchestrator calls governance first, and only a
Proceeddecision allows a mutating call.
The Public API
rust
pub fn is_world_model_enabled() -> bool;
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]
pub struct WorldEntry {
pub value: String,
pub entry_revision: u64,
pub provenance: String,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum WorldUpdateKind {
Assert { key: String, value: String, provenance: String },
Revise { key: String, value: String, expected_revision: u64, provenance: String },
Invalidate { key: String, expected_revision: u64 },
}
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
pub enum WorldUpdateOutcome {
Inserted, Revised, Invalidated, Unchanged,
Conflict, CapacityExceeded, Rejected,
}
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]
pub struct WorldSnapshot {
pub global_revision: u64,
pub entries: BTreeMap<String, WorldEntry>,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum WorldModelError { IOError, MalformedData, RevisionOverflow }
pub struct WorldModel { inner: RwLock<WorldModelInner> }
impl WorldModel {
pub fn new(config: WorldModelConfig) -> Self;
pub fn update(&self, update: WorldUpdateKind)
-> Result<WorldUpdateOutcome, WorldModelError>;
pub fn take_snapshot(&self) -> Result<WorldSnapshot, WorldModelError>;
pub fn get_entry(&self, key: &str)
-> Result<Option<WorldEntry>, WorldModelError>;
pub fn save_to_file(&self, path: &Path) -> Result<(), WorldModelError>;
pub fn load_from_file(&self, path: &Path) -> Result<(), WorldModelError>;
pub fn is_dirty(&self) -> bool;
}
A few details worth highlighting:
WorldEntrydoes not store the key. The BTreeMap key is the sole authoritative storage — no duplication.last_persisted_revisionis tracked alongsideis_dirty. If a concurrent write invalidates the saved snapshot during a slow file I/O, the dirty flag stays set and the next save cycle handles it.
Verification
The walkthrough confirmed everything we needed:
running 33 tests
test world_model::tests::test_invalid_config ... ok
test world_model::tests::test_assert_and_get ... ok
test world_model::tests::test_revise_and_invalidate ... ok
test world_model::tests::test_capacity_and_size_limits ... ok
test world_model::tests::test_persistence_and_prevalidation ... ok
...
test result: ok. 33 passed; 0 failed; 0 ignored; finished in 2.67s
| Operation | Latency (µs / op) |
|---|---|
| Insertion (Assert) | 0.1518 |
| Lookup (Read) | 0.0860 |
| Revision (Revise) | 0.1317 |
| Snapshot Creation | 95.7374 |
| Conflict Detection | 0.1289 |
| Invalidation (Invalidate) | 0.0851 |
The 100,000-operation stability run was the part that made me relax: memory growth was 4 KB over the entire run, throughput held at ~492k ops/sec, and not a single panic or deadlock.
Stage 25: Predictive Cognition
Once the framework had a world to look at, the next question was obvious: can we predict, before we commit to a reasoning path, whether it’s likely to succeed and what it’ll cost?
Stage 25 is a deterministic, explainable, resource-bounded predictor. It runs before the Stage 22 planner allocates any cells, so it has no circular dependency on planning budgets. The inputs are passed in explicitly via a PredictionEvidence struct:
rust
pub struct PredictionEvidence<'a> {
pub task_fingerprint: &'a crate::task_fingerprint::TaskFingerprint,
pub world_snapshot: &'a crate::world_model::WorldSnapshot,
pub invariants: &'a [crate::invariance::Invariant],
pub memory_record: Option<&'a crate::cube::memory::MemoryRecord>,
pub hardware_snapshot: &'a crate::planner::HardwareSnapshot,
pub cognitive_budget: &'a crate::planner::CognitiveBudget,
}
The Confidence Model
The final confidence score is a deterministic weighted sum:
C = wmem · Cmem + winv · Cinv + wworld · Cworld
The three component scores are computed as:
- Memory (C_mem) —
min(1.0, attempts / target_attempts) * successes / attempts, with0.5as the neutral default if there’s no usable history. - Invariance (C_inv) — count of invariants whose descriptions contain a keyword from the task fingerprint, normalized by
target_invariants. An empty slice returns0.5. - World (C_world) —
M_world / max(1, total_keywords), with a-0.5penalty per conflicting world entry.
We classify outcomes against T_success (default 0.80) and T_partial (default 0.50):
| Outcome | Rule |
|---|---|
LikelySuccess | C ≥ T_success AND no conflicts |
LikelyPartial | T_partial ≤ C < T_success AND no conflicts |
LikelyFailure | C < T_partial OR any conflict |
InsufficientEvidence | attempts < 2 AND M_inv == 0 AND M_world == 0 |
Why This Matters
The PredictionFactor enum makes the result explainable. A planner or operator can ask “why did the predictor say this?” and get an answer like PriorFailure, MissingWorldEvidence, or InsufficientHardwareBudget. That’s not a nice-to-have — it’s the only way to debug a real reasoning system.
A subtle point. The predictor is advisory on purpose. The Stage 22 planner still owns cell allocation. The predictor’s job is to surface information; the planner’s job is to use it.
Configuration Validation
A word of advice from the trenches: validate your configuration at construction. Weights must be finite and sum to 1.0 within 1e-5. Thresholds must satisfy 0 < T_partial < T_success < 1.0. Zero or negative bounds are rejected with PredictionError::InvalidConfig. This catches an entire category of bugs before they ever reach a test.
Stage 26: Recursive Cognitive Expansion
This is the stage I was most nervous about. We were adding the ability for the framework to expand itself — to take a hard reasoning task and break it into bounded child cognitive subcubes, each of which can in turn spawn grandchildren.
The risk surface was real:
- Runaway spawning could lock up the host.
- A child cube could write back to the parent’s world state.
- A panic mid-spawn could leave budgets in an inconsistent state.
- Recursive cycles could trivially send the system into infinite expansion.
The RAII Safety Guard
The centerpiece is ExpansionReservationGuard, a Rust RAII guard that holds reservations against the recursion budget. Its Drop implementation does exact state restoration — no partial refunds, no leaks.
rust
pub struct ExpansionReservationGuard<'a> {
pub budget: &'a mut RecursionBudget,
pub parent_node: &'a mut RecursiveNodeState,
pub parent_context: &'a RecursionContext,
pub request: RecursiveExpansionRequest,
pub node_id: u64,
pub parent_id: Option<u64>,
pub cells_reserved: usize,
pub node_reserved: bool,
pub execution_started: bool,
// Captured pre-reservation values for exact restoration
pub saved_remaining_nodes: usize,
pub saved_children_spawned: usize,
pub saved_active_capacity: usize,
pub saved_cumulative_work: usize,
}
impl<'a> Drop for ExpansionReservationGuard<'a> {
fn drop(&mut self) {
if !self.execution_started {
// Abandoned BEFORE child execution starts: exact restore of all metrics
self.budget.remaining_nodes = self.saved_remaining_nodes;
self.parent_node.children_spawned = self.saved_children_spawned;
self.budget.simultaneously_active_cell_capacity = self.saved_active_capacity;
self.budget.cumulative_recursive_cell_work_remaining = self.saved_cumulative_work;
} else {
// Child execution started and terminated: restore only active cell capacity
self.budget.simultaneously_active_cell_capacity = self.saved_active_capacity;
}
}
}
The two-branch logic matters: if execution never started, we refund every counter. If execution started, only the active cell capacity is restored (because some of the work has already been consumed by the child).
Atomic Reservation
The approve_expansion transaction performs every validation, every governance check, and every lineage check before any state is mutated. Then — and only then — does it commit.
The checks, in order, are:
- Cooperative cancellation —
token.is_cancelled()short-circuits. - Stage 22 authorization —
plan.allow_recursive_expansionmust be true. - Request bounds — task size, cell count, keyword count, keyword length.
- Depth —
context.depth < budget.max_depth. - Lineage cycle — no ancestor has the same task type and keywords.
- Node/branch/cell eligibility — counters must allow the reservation.
- Stage 23 governance — actual
evaluate_governancewith real GCF triggers. - Atomic calculation —
checked_add/checked_subfor every counter. - Monotonic node ID —
node_id = budget.next_node_id(never reused). - Commit — only now do we update the live state.
Why this order? If we updated state and then a check failed, we’d have to roll back. By computing every new value first and committing last, we make the entire reservation either succeed completely or fail with zero side effects.
The Child Executor
Once the guard is committed, the child runs under a scoped HierarchicalManager that preloads an isolated, in-memory WorldModel from a snapshot. The child cannot write back to the parent’s world, and it cannot read from the parent’s memory — it gets an empty ephemeral CubeMemory.
This is what “isolation” looks like in practice:
rust
pub fn new_child_scoped(
world_snapshot: crate::world_model::WorldSnapshot,
prediction_engine: crate::planner_predict::PredictionEngine,
) -> Self {
let isolated_model = std::sync::Arc::new(
crate::world_model::WorldModel::from_snapshot(
world_snapshot,
crate::world_model::WorldModelConfig::default(),
)
);
Self {
memory: crate::cube::memory::CubeMemory::new_empty(),
face_memory: crate::face_memory::FaceMemory::new(),
synthesis_engine: SynthesisEngine::new(),
autonomous_system: AutonomousSystem::new(),
input_gate: InputSufficiencyGate::new(),
validation_engine: None,
world_model: isolated_model,
prediction_engine,
}
}
Bounded Topology
The configuration caps the worst case explicitly:
rust
impl Default for RecursionConfig {
fn default() -> Self {
Self {
max_depth: 3,
max_total_nodes: 10,
max_children_per_node: 3,
max_cells_per_child: 4,
max_total_cells: 16,
enabled: true,
}
}
}
Given max depth D and branching B, the worst-case node count is:
Nworst =
| BD+1 – 1 |
| B – 1 |
With our defaults, that’s 40 nodes — but the config cap clamps it at 10. The math is there to prove the cap is sufficient; the cap is there to enforce it.
Result Merging
When a child finishes, its output is merged into the parent using a deterministic status-precedence ladder:
| Precedence | Status | Notes |
|---|---|---|
| 1 | Cancelled | Cooperative cancellation, top priority |
| 2 | ChildFailed | Hard failure in a descendant |
| 3 | BudgetExhausted | No remaining cells, nodes, or capacity |
| 4 | DepthLimitReached | Tree depth limit hit on a descendant |
| 5 | CycleDetected | Repeated task type + keywords in lineage |
| 6 | GovernanceDenied | Stage 23 denied the descendant spawn |
| 7 | PlannerDenied | Stage 22 denied expansion on the descendant |
| 8 | Partial | Output, branching, or child-output limit truncated |
| 9 | Completed | Default lowest-precedence successful state |
The crucial rule: a higher-precedence status can never be silently overwritten by a lower-precedence one. If a child’s Cancelled status is followed by a MergedOutputLimitExceeded while merging output, the final status stays Cancelled.
Stress Test Results
The stability workload (100,000 operations) was the moment of truth:
| Metric | Value |
|---|---|
| Total Time | 171.106 s |
| Throughput | 584.43 ops/sec |
| Initial Memory | 6,352 KB |
| Final Memory | 7,056 KB |
| Net Memory Growth | 704 KB (completely bounded) |
| Panics / Leaks | 0 |
The microbenchmarks were even better:
| Operation | Latency |
|---|---|
validate_expansion_tree() | 0.5933 µs / tree |
MergeAccumulator merge | 0.2420 µs / op |
The Stage 18 regression check came in at 0% — well under the 5% ceiling.
Patterns I’d Take to Any System
Looking back at this work, four patterns are worth carrying into any large Rust system:
- Advisory layers, not autonomous ones. The predictor never spawned work. The expansion manager never overrode the planner. The world model never bypassed governance. Keeping each layer advisory to the layer above it kept the entire system auditable.
- Atomic reservations with captured pre-state. Any code path that mutates multiple counters should snapshot them, run all checks, compute all new values, and commit last. The
ExpansionReservationGuardpattern generalizes to almost any multi-resource allocation in Rust. - Validated configuration at construction. Don’t ship magic numbers. Validate every config field at startup, reject zero/negative bounds, and use safe fallbacks. You will sleep better.
- Bounded memory is a non-negotiable. Every collection has a max size, every counter has a max value, every recursion has a depth cap. The 100k stability run with sub-megabyte growth is the kind of evidence that makes other engineers willing to deploy your code.
Where This Goes Next
Stages 24–26 close out a coherent phase of GCF 2.0. A few things on my mind for the next iteration:
- Cross-stage end-to-end test. Lock in the contract that all three stages together work as a pipeline.
- Wire
PredictionFactor::HighHistoricalFrameCostinto an early-exit hook so the predictor’s advice actually saves cycles in production. - Promote
MergeAccumulatorto a shared utility. Its precedence logic is general and would help other multi-source aggregations in the codebase.
If you’re building a reasoning system in Rust — or any system with budget, governance, and recursion — I’d love to compare notes. The full implementation plans for all three stages are coming in a longer-form report, but the patterns above should give you a head start.
— Billy P.
Appendix: Quick Reference
| Stage | Headline Test | Headline Performance | Status |
|---|---|---|---|
| 24 | 33/33 tests pass | 491,800 ops/sec, 4 KB growth over 100k | PASS |
| 25 | Determinism + evidence validation | Sub-microsecond predict() latency | SPEC & IMPLEMENTED |
| 26 | Full suite, 0 warnings | 0.59 µs/tree, 0.7 MB growth, 0 panics | PASS |
Reposted from the GCF 2.0 engineering notebook. Tested on branch=2.0, last verified August 2026.