· Dispatch Security

Capability Can Increase Without Authority Increasing

 ·  Billy p

By Billy P. · 14 min read · GCF Perspective #2

A researcher wrote to me recently asking a question I have been thinking about for a long time. The question was whether smaller models are inherently safer. I have a fairly specific answer, and I think it is worth sharing here in full, because the answer is also the architectural principle that has shaped almost every decision in the ISAC and GCF work.

The short version is this: increased intelligence must not automatically result in increased authority. A model can become substantially more capable without gaining any new permissions, memory access, tool authority, or autonomy. Model capability and system authority are treated, in our architecture, as separate control surfaces. They are not the same dial.

This post is an expanded version of my reply. I have tried to keep the original reasoning intact while adding the context I think matters for anyone working on AI systems that have to operate in the world for any length of time.


Smaller models are not inherently safer

The researcher’s framing was that capability-per-compute improvements raise a specific safety question. If a small model becomes capable enough to do work that previously required much larger infrastructure, then it becomes cheaper to operate, easier to distribute, and harder to control through infrastructure-level restrictions alone. The worry is that a sufficiently capable small model with persistent memory, broad tool access, and meaningful autonomy could end up presenting more risk than a substantially larger model operating under very tight authority limits.

I agree with this framing. I do not regard smaller models as inherently safer. The relevant variable is not parameter count. It is the interaction between capability, access, and authority.

A 3B-parameter model that performs at a level previously associated with much larger systems is a safety-relevant capability increase, even though the parameter count itself has decreased. What matters is what the model can actually do, and what it has been authorised to do, and whether those two things are still kept in proportion.

This is the lens through which I evaluate the rest of the architecture.


The core principle: intelligence and authority are separate

The architectural principle I keep coming back to is a simple one:

Increased intelligence must not automatically result in increased authority.

If a model becomes more capable, it does not get more permissions. If a model becomes better at planning, it does not get more autonomy. If a model becomes better at using tools, it does not get more tools. Each of these is a separate decision, made on separate evidence, by a separate governance process.

In GCF, this distinction is enforced at the architectural level rather than the behavioural level. We do not rely on the model deciding not to exceed its authority. We rely on the boundary refusing to grant authority the model has not been explicitly authorised to use. The model can recommend that a particular capability or permission would be useful; recommending authority is fundamentally different from granting that authority to itself. The proposal goes into a governance queue, not into the model.

That last sentence matters more than it sounds. Most AI safety failures I have read about can be described, in plain language, as a model doing something it should not have been able to do. Not a model deciding to do something bad. A model being able to do something bad because the system around it gave it the means. The architectural response is to take the means away from the model entirely and put them behind a boundary the model cannot cross.


The architecture: ISAC, GCF, and the replaceable cognitive component

Here is how the pieces fit together.

ISAC is the wider AI system. It is the user-facing assistant, the thing people interact with. It handles conversation, planning, task execution, memory recall, and the day-to-day cognitive work a companion is expected to do.

The underlying model is a replaceable cognitive component inside ISAC. It is the engine that does the reasoning. It might be one of my own future models, or it might be an external model from OpenAI, Anthropic, Google, or another provider. I want this to be a real architectural fact, not a marketing claim. The model is plug-compatible. Swapping it does not change the safety contract.

GCF is the governed boundary between that cognitive capability and the sensitive parts of the system. The sensitive parts are:

  • persistent memory,
  • tools,
  • permissions,
  • resource access,
  • and any action capable of affecting the external environment.

The intended structure is therefore not:

model → unrestricted tools / memory / actions

It is:

model → governed GCF boundary → authorised memory / tools / resources / actions

GCF sits between the model and everything with side effects. It decides what is allowed. It logs what was attempted. It refuses what is not authorised. The model cannot bypass it, because the model has no path to the sensitive surfaces that does not go through GCF.

This is the design principle. It is the same regardless of which model is sitting behind the boundary. A more capable model receives no additional permissions, memory access, tool authority, or autonomy simply because its reasoning has improved. Capability goes up. Authority stays where it was.


What “become smarter over time” actually means

I have been using the phrase “become smarter over time” in a deliberately narrow sense, and it is worth being explicit about what I mean and what I do not mean.

What I mean by it is system-level improvement. Things like better retrieval, improved use of memory, stronger task routing, better model selection, governed skill development, improved contextual understanding, more efficient reasoning, and better allocation of available compute. These are improvements inside the boundary. The boundary itself does not move.

What I do not mean is unrestricted recursive self-modification. The following operations are treated as fundamentally different from ordinary learning or adaptation:

  • modifying model weights,
  • replacing the model,
  • changing GCF,
  • altering safety controls,
  • increasing system permissions.

Each of those is a separate governance event. They are not things a deployed assistant does to itself.

In particular, I do not intend for an operating ISAC instance to be capable of:

  • unilaterally modifying or bypassing GCF,
  • granting itself greater permissions,
  • removing its own safety mechanisms,
  • or silently replacing its cognitive model with a more capable one.

There may eventually be controlled research into model adaptation or weight updates, but I would treat that as a separate model-development process rather than allowing a deployed assistant to rewrite itself without external controls.

This is not a temporary restraint that I expect to relax as the system matures. It is a property of the architecture. The system does not have the ability to do these things to itself, in the same way a well-designed operating system does not give a user process the ability to rewrite the kernel. The capability is not in the trust boundary, full stop.


Future models and capability-escalation gates

The introduction of future models is another area where I think a stronger formal process is required. My current view is that a newly trained model should not be automatically promoted into production simply because it performs better on conventional benchmarks.

Before deployment, a model should undergo a capability and governance evaluation covering at least:

  • general capability,
  • safety behaviour,
  • adversarial robustness,
  • tool use,
  • autonomy,
  • planning,
  • governance compatibility,
  • and its interaction with GCF.

I also intend to establish capability-escalation thresholds. If a new model crosses a significant threshold, particularly in capability-per-compute, that should trigger additional evaluation rather than being treated purely as an engineering improvement.

For example: if I produced a 3B-parameter model performing at a level previously associated with substantially larger systems, I would consider that a safety-relevant capability increase even though the parameter count itself had decreased. The evaluation should go beyond benchmark scores and examine what has materially changed in the model’s behaviour. Possible changes to look for include improvements in extended planning, tool use, vulnerability discovery, deceptive behaviour, autonomous task execution, workflow replication, persistence across long-running tasks, or other capabilities that could meaningfully alter the system’s risk profile.

Parameter count and safety are not the same variable. What matters far more is the interaction between capability, access, and authority, which is the same sentence I opened with. I keep coming back to it because the architecture keeps confirming it.


Why I built GCF before ISAC

This is one of the questions I get most often, and the answer is connected to everything above.

I did not want the architecture to look like this:

model → unrestricted tools / memory / actions

I wanted it to look like this:

model → governed GCF boundary → authorised memory / tools / resources / actions

The first version is what most AI assistants ship as. The model is the system. Whatever the model decides to do, the system does. The capability and the authority are not separable. If the model becomes more capable, the system becomes more capable, and the system becomes harder to constrain.

The second version is what I wanted instead. The capability and the authority live behind different boundaries. The model can be swapped without touching the authority layer. The authority layer can be tightened without retraining the model. The two evolve on separate timelines, with separate review processes, against separate evidence.

GCF is therefore not intended as a claim that a sufficiently capable model can never cause harm. I would not make that claim. It would not be a responsible claim to make.

Its purpose is instead to avoid placing the entire safety burden on the model’s own behaviour. A capable model should still operate within externally defined constraints, with auditability, tool governance, resource controls, human override mechanisms, and explicit permission boundaries. If the model misbehaves, the boundary catches it. If the boundary misconfigures, the audit log shows it. If the audit log is tampered with, the human override stops the system. There is no single point of failure, and the model is never the single point of failure.

This is the architecture. It is the architecture I want to be true of any deployed AI system that has access to memory, tools, or actions affecting the outside world.


The smart-per-GB proliferation threshold

There is one more consequence of the original “smart per GB” objective that I think deserves explicit acknowledgement.

If capability becomes significantly more efficient, the threshold for proliferation falls. This has clear benefits. Local computing, privacy, accessibility, cost, and energy consumption all improve when a capable model runs on a laptop instead of a cluster. I want all of those benefits. They are real.

But lower proliferation thresholds also mean that advanced capabilities become easier to reproduce, deploy, and distribute. A model that runs locally is a model that can be modified locally. A model that can be modified locally is a model that can be removed from any centralised safety process. We do not get the benefits of efficient capability without also having to take the proliferation consequences seriously.

For that reason, I now believe that capability-per-compute should itself be considered during safety evaluation, rather than being treated solely as a performance or efficiency metric. A 3B model that performs like a 70B model is a different kind of artefact than a 3B model that performs like a 3B model. We evaluate it differently, we govern it differently, and we deploy it behind different gates.

This is, I think, the honest version of the capability-per-compute conversation. We can have efficient models, and we can have safe systems, but only if we treat efficiency as a safety-relevant variable instead of a pure engineering win.


Summary

The principle I would use to summarise the design is:

Capability can increase without authority increasing.

A model may become substantially more intelligent. Permissions, persistent memory, tool access, autonomy, and the ability to affect the external environment remain separately governed. They are governed by GCF, on the outside of the model, not by the model’s own behaviour.

As ISAC progresses beyond its current foundation stages, I intend to formalise this further through a capability-escalation gate for future model development. Significant improvements in intelligence, autonomy, or efficiency will trigger additional safety evaluation and governance review before deployment. The model does not promote itself. The system around it decides.

I would also be very interested in views on which capability thresholds would be most useful to formalise, particularly where improvements in capability-per-compute begin to change proliferation or autonomy risk rather than simply improving benchmark performance. If you are working on this and have thoughts, the GitHub discussion on the GCF 2.0 repo is the right place to bring them.

— Billy P., Lead Engineer, GCF 2.0 Cognitive Systems


This is the second post in the GCF Perspective series. The first post, On AI Extinction Risk: An Engineer’s Perspective, laid out the four-variable risk equation and tied it to the GCF 2.0 architecture. This post answers a related, more specific question: when the model gets smarter, what changes about the system around it? The short answer is: nothing, until governance explicitly says so.

Tested on the GCF 2.0 frozen framework, SHA eeae2bc8a426b063c687e40f9724adbb1dc3a61a. The separation between model capability and system authority is enforced at the GCF boundary, not in the model’s behaviour. Replacing the model does not change the safety contract.

← All dispatches Homepage →