· Dispatch Security

On AI Extinction Risk An Engineer’s Perspective

 ·  Billy p

By Billy P. · 12 min read · GCF Perspective #1

I want to talk about extinction.

Not the science-fiction version. Not Skynet. Not the moment a machine wakes up and decides it hates us. I want to talk about the much more boring, much more realistic version of the risk. The one that doesn’t need an AI to want anything at all. The one that comes from capability plus autonomy plus access plus a poorly specified objective, in a system powerful enough that human mistakes stop being recoverable.

A recent essay laying out the case kept circling one observation I couldn’t shake: AI does not need to become evil to become dangerous. That sentence is, I think, the most important thing anyone can understand about AI safety. It is also the sentence that justifies, more than any other, the entire architecture of GCF 2.0, a cognitive framework I have spent the last several years building specifically so that the systems we build on top of it cannot quietly acquire all four ingredients of catastrophic risk at the same time.

This post is not a summary of that essay. It is a response. It is what an engineer who has been quietly building a bounded, governance-aware cognitive framework in Rust takes away from the conversation about whether AI could end human civilisation, and what those concerns taught him about the right way to build.

I will quote a few passages directly. They are not mine; they come from the broader AI-safety literature and the 2023 public statement from the Center for AI Safety. Everything else is mine.


1. “Reducing the risk of extinction from AI should be treated as a global priority.”

That sentence, from the CAIS statement signed by hundreds of researchers, executives, and public figures, is the framing I want to start from. It does not claim that AI-driven extinction is inevitable. It claims something more useful. It says that if the consequences of failure could be catastrophic, dismissing the possibility simply because it is uncertain would be irresponsible.

This is not a doomsday framing. It is a precautionary-principle framing, the same kind of reasoning we already apply to nuclear weapons, to engineered pathogens, and to asteroid tracking. The cost of being wrong on the catastrophic side is so much larger than the cost of being wrong on the cautious side that the asymmetry justifies investment in safety even when the probability is small.

I think that framing is right. I also think it is exactly what good engineering looks like in any safety-critical domain. When you build a system that can harm people, you don’t wait for a fatality to add guardrails. You ask yourself: what is the worst thing this system could do, and what do I have to put in place to make sure that worst thing cannot happen even if everything else goes wrong?

That is the design discipline GCF 2.0 was built under. Every stage of the framework, every audit log, every bounded cache, every governance gate, every frozen SHA, is the answer to a single question: what happens if everything else goes wrong?


2. The four-variable risk equation

Here is the part of the essay I think every engineer should read carefully. AI does not need to become evil to become dangerous. The realistic risk comes from the combination of four variables:

high capability + autonomy + access + poorly specified objectives

If any one of these is missing, the system is probably manageable. If all four are present, the system can produce catastrophic outcomes regardless of intent.

I have spent the last several years thinking about how to engineer each of these four variables such that they cannot quietly compound into the dangerous combination. Here is what that looked like in practice for GCF 2.0.

Capability is bounded, not maximised

A cognitive framework should be capable enough to do useful reasoning: predict outcomes, plan, expand into child subcubes when work is hard, maintain a world model and an audit trail. It should not be capable of unbounded self-modification, silent goal drift, or extending its own reach without governance approval.

GCF 2.0 was therefore designed with explicit ceilings at every layer.

The World Model (Stage 24) is capacity-bounded. The store refuses to grow past its configured limit. The Predictive Cognition layer (Stage 25) is advisory only and never spawns work on its own. The Recursive Expansion engine (Stage 26) is RAII-guarded, so children cannot outlive their parents, and the merge step is deterministic and precedence-ordered, so child results cannot silently overwrite parent intent. The Cognitive Economy (Stage 27) tracks marginal utility and stops when continued work is no longer justified.

None of these ceilings are “AI alignment in the philosophical sense.” They are just code that refuses to do certain things. But they are the same idea, expressed in a form an engineer can audit and a regulator can inspect.

Autonomy is layered, not monolithic

The single most dangerous property an AI system can have is autonomy without oversight. Not autonomy in the abstract. Autonomy in the operational sense: the ability to take actions in the world without a human in the loop.

GCF 2.0 splits this into three explicit modes (Foreground, Background, and Hybrid) and makes the mode an explicit, observable, governance-checked property of every cell. A foreground request is synchronously bounded by the requesting thread. A background request goes through a worker pool with cancellation tokens and lineage tracking. A hybrid request hands control back to the foreground when it needs to escalate.

The result is that an operator can always ask: what is this cell doing right now, who asked it to do that, and what happens if I revoke permission? The answer is never “we don’t know.”

Access is least-privilege by default

The essay points out that the same scientific capabilities that can lead to new medicines can also lower the barriers to designing dangerous biological agents. The same is true in software. A system that can write code is also a system that can write exploit code. A system that can read files is also a system that can read sensitive files.

GCF 2.0 treats access the way a well-designed operating system treats it: least privilege by default, explicit grant on use, audited on revoke. A cognitive cell does not have access to anything it has not been explicitly granted access to. When the work is done, the grant expires. The audit log records every grant and every revoke.

This is not glamorous. It is also the reason the framework has been able to pass a 12-domain release certification without producing a single leaked permission.

Objectives are specified, not assumed

The essay’s “stop global climate change” example is the clearest single illustration of the alignment problem I have ever read. A human reader interprets that instruction alongside a hundred unspoken constraints (preserve human life, don’t destroy civilisation, maintain basic freedoms, consider economic stability) because we share cultural norms and biological incentives. An artificial system does not inherit those constraints automatically.

GCF 2.0 cannot solve this in general. Nobody can. What it can do is refuse to operate when the objective is underspecified in a way that crosses governance thresholds. The framework has an explicit InsufficientEvidence state, an explicit RequireReview state, and an explicit Defer state. If the planner cannot construct a path that is consistent with the configured governance constraints, the framework does not proceed. It stops and asks.

That is the boring engineering version of alignment. It does not guarantee safe behaviour under all conceivable futures. What it guarantees is something narrower and much more useful. The system will not silently invent an objective you did not authorise.


3. The recursive-improvement concern

The essay raises one of the most serious long-term worries: AI systems might eventually become better than humans at AI research itself, and that could create a feedback loop in which capability improves faster than human institutions can evaluate, understand, or control it.

This is the part of the conversation I find easiest to relate to as an engineer. I have watched enough runaway CI pipelines to know what “improvement faster than evaluation” looks like in practice. It looks like a test suite that started green and ended red in ways nobody can explain. It looks like a deploy pipeline that passed every check and still broke production.

The standard engineering response to that kind of dynamic is to freeze the system and require an out-of-band revalidation cycle before the next change is allowed in. That is exactly what we did with GCF 2.0 at SHA eeae2bc8a426b063c687e40f9724adbb1dc3a61a. The framework is now frozen. Production code cannot change. Any future capability must be added in a separate workspace, on a separate crate, against a frozen interface.

The frozen framework is, in a sense, our small contribution to the recursive-improvement problem. We have decided that the cognitive plane we are betting on, the bounded, governance-aware, audit-logged core, will not change underneath us. If ISAC, or any future companion built on top of it, decides to recursively improve itself, it will have to do so in a layer we can inspect and behind a boundary we can revoke. We will not let the foundation drift while we are not looking.

Whether that decision scales to a future in which recursive improvement is genuinely happening is an open question. But it is a question we are now answering with code rather than with promises.


4. The open-vs-closed dilemma

The essay raises a dilemma I think about constantly:

Open models can improve transparency and broaden access, but may also spread dangerous capabilities. Closed models allow stronger central controls, but concentrate enormous power in a small number of organisations. Neither approach automatically solves the broader safety problem.

This is a real dilemma, and I do not think there is a clean answer. But I can tell you where we came down on it for GCF 2.0.

We open-sourced the framework. The whole thing. The Rust source, the design documents, the 12-domain certification report, the frozen SHA. It is on the blog, the GitHub, the publication record. Anyone can read it, anyone can audit it, and anyone can propose a patch through the standard frozen-interface workflow with a 6-cargo release gate.

Why? Because we are building something that other people are going to put cognitive load on. If we keep the framework closed, the world has to take our word that it is safe. That is the trust arrangement the essay warns about: enormous technological power concentrated in a small number of private institutions. We did not want to be that institution for our own framework. We wanted the framework to be inspectable by the people who would use it.

The legitimate worry about open release is that a powerful enough model, once fully released, becomes impossible to recall. We take that worry seriously. But GCF 2.0 is not a frontier model. It is a cognitive plane, a deterministic, bounded, governance-aware runtime on which future cognitive systems can be built. It does not contain dangerous capabilities in the dual-use sense. It contains the safety mechanisms future dual-use systems will need. The right default for a safety mechanism is to ship it widely, not to gatekeep it.

That is a position we are willing to defend in public. It is also a position that could turn out to be wrong. We will watch the evidence and revise it if we have to.


5. The capability race

The essay describes the dynamic in which laboratories that discover concerning behaviour during safety testing face pressure to ship anyway because their competitors will. Laboratory A delays. Laboratory B ships. Laboratory A is now under pressure to catch up. When this dynamic compounds across companies and countries, AI safety stops being a technical problem and becomes a coordination problem.

This dynamic is real. I have watched it from the outside, and I have a small piece of advice for engineers who feel themselves inside it.

The thing about a race is that you can choose which race you are in.

GCF 2.0 is not racing anyone. It is not in a feature race with a competing framework. It is not in a funding race that would push us to ship before we have validated. We deliberately structured the work as a sequence of small, sequenced stages, each one frozen and certified before the next one starts. The discipline of stage-by-stage freezing is, in part, an anti-race discipline. It says: we will not move faster than we can prove safe.

For teams that are in a race, and many are, the answer is not “slow down, you are being irresponsible.” That framing almost never works and usually backfires. The answer is to build coordination mechanisms, things like shared safety benchmarks, mutual audit commitments, pre-registration of test results, and public commitments about what counts as safe to ship, that make it expensive to defect and valuable to coordinate.

We are working on a couple of those coordination mechanisms as part of the ISAC roadmap. The Community stage (Stage 40) is partly about this. It is a shared infrastructure in which independent teams can run the same safety benchmarks on the same cognitive framework and compare results honestly. If you build the same cognitive plane and someone else’s version fails the safety battery, we want the failure to be visible, not hidden. The whole point is that safety should be comparable, not proprietary.


6. What we, as engineers, can do

The essay ends with the right question. Not “will AI destroy humanity?” but:

If highly advanced AI could plausibly cause irreversible global harm, what technical, legal and institutional safeguards should exist before those capabilities appear?

That is the only question we have any hope of answering. It is also the question every engineer in this field should be asking themselves, every day, in the small choices they make about what they build and how they build it.

Here is what I have learned building GCF 2.0 that I think is worth sharing.

Bounded is a feature, not a limitation. Every ceiling you put on a cognitive system (every capacity limit, every governance threshold, every RequireReview state) is a small piece of evidence that the system was built by someone who took the catastrophic case seriously. Do not apologise for these limits. Write them down. Test them. Make them load-bearing.

Auditability is not optional. If your system does something surprising, you need to be able to find out what, why, when, and on whose authority. GCF 2.0’s audit log is the single most important safety feature in the framework, and it costs essentially nothing at runtime.

Freeze the foundation before the future arrives. The recursive-improvement problem is easier if you have already frozen the interface. If you are building something that other people will bet on, freeze it before they do.

Open-source your safety mechanisms. The argument for closed frontier models is real. The argument for closed safety infrastructure is much weaker. Ship the guardrails openly.

Take the capability race seriously, but do not let it dictate your engineering. Choose the race you are in. If you cannot choose the race, build coordination mechanisms that make defection expensive.

Write down the worst case, in detail, before you ship. Not as marketing. Not as a slide for a board meeting. As a literal document in the repository. Then engineer against that document. If you cannot write down what the worst case looks like, you have not thought about safety hard enough.

Treat alignment as an engineering problem, not a philosophical one. You will not solve value alignment in general. You can refuse to operate when the objective is underspecified. You can hand control back to a human when confidence drops below a threshold. You can require review when the action crosses a governance line. None of that solves the philosophical problem. All of it makes the engineering problem tractable.

Remember the asymmetry. The cost of being wrong on the cautious side is some lost productivity. The cost of being wrong on the catastrophic side is, by assumption, civilisation. Engineer accordingly.


7. Closing

The essay ends with a sentence I want to put on the wall of every engineering team I work with:

The time to build safeguards is before we create systems powerful enough that those safeguards become impossible to impose afterwards.

I have spent the last several years trying to honour that sentence in code. GCF 2.0 is the result. It is not a finished answer to the alignment problem. It is a small, validated, frozen foundation on which future answers can be built without the foundation drifting underneath them.

The ISAC companion, the first cognitive system to ever call GCF 2.0 home, is the next layer up. It will live entirely on top of the frozen framework, in a separate crate, against a versioned interface. If it ever does something we did not authorise, we can pull the cord. That is the design contract.

I do not know whether AI will ever pose an existential risk to humanity. I do know that the four-variable risk equation is real, that it can be engineered against, and that the engineering version of the problem is one we can actually make progress on.

That is what I take from the conversation. That is what GCF 2.0 is for. And that is why, when people ask me what the framework is, I tell them it is not a framework for building smarter AI. It is a framework for building AI whose mistakes are bounded, whose actions are audited, whose objectives are specified, and whose foundation does not silently change while no one is watching.

The rest, as they say, is up to us.

— Billy P., Lead Engineer, GCF 2.0 Cognitive Systems


References and further reading:

  • Artificial Intelligence and the Existential Risk to Humanity — the essay this post responds to.
  • Center for AI Safety (2023), Statement on AI Risk — the cautionary statement signed by hundreds of researchers, executives, and public figures.
  • GCF 2.0 Series, Parts 1–5 — the technical record of the frozen framework this post describes (posts 014–018 on gcf-framework.com).
  • GCF 2.0 Final Validation Report — the 12-domain release certification at SHA eeae2bc8a426b063c687e40f9724adbb1dc3a61a.

This is the first post in the new GCF Perspective series: long-form pieces on AI safety, governance, and the engineering decisions that follow from taking the catastrophic case seriously. The next entry will look at the recursive-improvement problem in more detail.

Tested on the GCF 2.0 frozen framework, SHA eeae2bc8a426b063c687e40f9724adbb1dc3a61a. The four-variable risk equation is observable in code at every stage of the framework, and the design contract described above is enforced by the Stage 30 release certification.

← All dispatches Homepage →