· Dispatch Code

Stage 35 Is Finished. Now We Try to Break It.

 ·  Billy p

By Billy P. · Founder Notes #2.

Written from Blackpool, UK. Probably while you read this.

A small but pretty important update on GCF and ISAC.

Stage 35 is finished.

That means another part of the roadmap is now officially behind us.

Normally the obvious thing to do would be to jump straight into Stage 36 and keep pushing forward. I’m not doing that. Before I add anything else, I’m stopping development for a bit and doing something I think is just as important. A full audit and bug hunting session.

And when I say audit, I don’t mean quickly running a few tests, seeing green ticks and then moving on. I mean going back through the foundation properly. Looking at what has been built. Looking at what assumptions were made. Looking for places where something technically works but could fail under the wrong conditions. Trying to find bugs. Trying to find bad design decisions. Trying to find anything I’ve missed. And basically trying to break the system before somebody else eventually does.

That is where the project is right now. Stage 35 is done. Stage 36 is next. But first, I want to make sure the foundation we’re building on is actually solid.


Why stop now

Because I think one of the easiest mistakes to make with a project like this is constantly adding more. Another feature. Another model. Another integration. Another system. Another layer. It feels like progress because the project keeps getting bigger. But bigger doesn’t automatically mean better.

At some point you have to stop and ask whether everything underneath this still works properly.

That’s especially important with GCF because the next part of the roadmap starts opening the system up more. Stage 36 is where we start moving toward external and custom providers. That means things like external model providers, custom endpoints, credentials, cloud policies, fallback logic, and more. Once you start introducing those things, the number of ways something can go wrong grows very quickly.

So before I do that, I want the foundation checked properly. I’d much rather find an ugly problem now than build another ten stages on top of it.


Some people won’t find this interesting

And that’s fine.

Not everybody is going to look at GCF and think this is something important. Some people will probably look at it and think “so what.” AI already exists. There are already massive models. There are already companies spending billions building them. Why does another framework matter.

And I understand that reaction. Because GCF isn’t really about trying to beat everyone by building the biggest model. That’s not what I’m trying to do.

What interests me is everything else. The stuff around the model. Memory. Governance. Routing. Resource usage. Specialisation. Permissions. Tool access. Different reasoning paths. Recovery. Audit. How the system decides what should happen before a model is ever allowed to act.

And I think those areas are going to matter more and more as AI systems become capable of doing things instead of just replying to a message.


I think we focus too much on one number

A lot of AI discussion seems to end up being about size. How many parameters. How many GPUs. How big was the training run. How large is the context window. How much compute. And again, those things matter. I’m not saying they don’t.

But what I’ve been trying to explore with GCF is whether we can improve the system instead of constantly relying on the model getting bigger. What happens if memory improves. What happens if routing improves. What happens if the system can choose different reasoning approaches depending on the problem. What happens if it learns what worked previously. What happens if smaller models can become better at specific tasks because the framework around them gives them better structure.

Those are the questions I find interesting. And recently we started getting some pretty interesting answers.


The 3GB model experiment

One of the experiments that has made me more confident in this direction involved a model that was only around 3GB. That’s important because this wasn’t some enormous model running on a giant cluster. It was small enough to run locally.

Initially, its performance was what you would probably expect from a model of that size. Useful in places. Limited in others. Nothing particularly amazing.

But instead of immediately saying “the model isn’t good enough, we need a bigger one,” we let the system work with it. We let it learn. We let the architecture build up experience around how it was being used.

And then something started happening. Its behaviour improved. Noticeably. The same small model started becoming much more useful than it had been at the beginning.

That doesn’t mean we’ve magically turned a tiny model into the world’s smartest AI. We haven’t. And I don’t want to make that claim.

What it does mean is that we’re starting to see evidence for something I’ve believed for a while. The model itself isn’t the only place where intelligence can improve. The system surrounding the model matters too. How memory is used matters. How tasks are routed matters. How previous outcomes are remembered matters. How the system decides what reasoning path to take matters.

That is exactly the kind of thing GCF was built to explore. And seeing a small model improve through the architecture around it is one of those moments where you stop and think, okay. There might actually be something here.


Is that proof GCF works

It is proof that parts of the idea are working. That’s the distinction I want to keep making.

I’m not going to say one experiment proves the entire architecture. It doesn’t. There are still too many things we haven’t tested. There are still areas that need external validation. There are still questions around scale. There are still questions around how different models behave. There are still questions around long-term memory, failure modes, and how the framework behaves under much heavier workloads.

That’s exactly why we’re doing the audit now.

But compared with where this project started, we have something we didn’t have before. We have measurements. We have tests. We have benchmarks. We have working behaviour. We have things that can fail or succeed in ways we can actually observe. For me, that is a huge difference.


The benchmarks are coming too

Another thing I want to start doing as the project matures is making more of the benchmark information public.

I don’t want GCF to become one of those projects where every claim is basically “trust me, it’s amazing.” If I’m saying a model improved, I want people to eventually be able to see what changed. If I’m saying the framework uses resources differently, I want numbers behind it. If I’m saying a governance mechanism blocks something, I want people to see the test. If something performs badly, I want that visible too.

So over time I plan to start publishing benchmark results from GCF and ISAC. Not everything will be released immediately. Some tests still need cleaning up. Some need to be repeated. Some need better methodology. And some results probably aren’t meaningful enough yet to publish. But the goal is that the important claims behind this project shouldn’t permanently live behind screenshots or statements from me.

There should be data.


I want people to challenge it

This is probably going to sound strange considering I’ve spent years building it. But I want people to challenge GCF.

I want people to ask difficult questions. I want someone to look at the architecture and say “why did you do it that way” or “what happens when this fails” or “have you tested this against this condition.” That’s useful.

The worst thing that could happen to a project like this would be everyone around it telling me every idea is brilliant. Because they won’t all be brilliant. Some things will be wrong. Some assumptions will fail. Some parts of GCF will probably get replaced completely before we ever reach the final stages of the roadmap. That’s normal.

The goal isn’t to protect every idea I’ve ever had. The goal is to end up with something that works. That’s also the entire point of this audit. I’m not trying to prove the foundation is perfect. I’m trying to find out where it isn’t.


What happens after the audit

Assuming the foundation survives the process and the problems we find are fixed, the next major step is Stage 36. That is where GCF starts expanding into External and Custom Providers.

Up until this point, a big focus has been keeping the architecture local-first and building the internal foundation. Stage 36 starts adding optional support for external systems. That includes things such as external AI providers, custom model endpoints, provider fallback, credential handling, privacy policies, cloud-use rules, cost controls, and the systems needed to decide when an external provider should or shouldn’t be used.

The important word there is optional. I still want GCF to be capable of operating locally. The point isn’t to turn it into something dependent on a single cloud provider. The point is to make the architecture capable of choosing between local and external resources while still keeping governance around those decisions. That is going to be a big stage, which is another reason I don’t want to rush into it.


This is the bit I actually enjoy

Something I’ve noticed recently is that some of the most interesting moments in this project aren’t when I add another feature. They’re when something unexpected happens. A small model performs better than I expected. A benchmark exposes a weakness. A governance test blocks something exactly where it should. A design that looked perfect on paper turns out to be terrible once it’s running.

Those moments teach you more than adding another box to an architecture diagram.

And that’s what this next period is going to be about. Less “look at the new feature.” More “let’s see what breaks.” Because if GCF is eventually going to be used by other developers or businesses, I want to know as much as possible about where it fails before that happens.


So where are we right now.

Stage 35 is finished. The project is now in a full Foundation Stability Audit and bug hunting period. We’re testing. We’re reviewing. We’re fixing. We’re trying to break things. And when we’re happy that the foundation is in a good place, we move into Stage 36.

At the same time, the evidence behind the project is starting to become more interesting. We’re seeing small models improve. We’re seeing governance behaviour working in code. We’re collecting benchmarks. And eventually I want much more of that evidence to become public, because I don’t want people to believe in GCF just because I say they should. I’d rather show what we’re actually seeing and let people make up their own minds.

Maybe some people still won’t find it interesting. That’s completely fine.

But personally, every time I see another part of the architecture working, it makes me want to push the idea further. Because this project was never really about building the biggest AI. It was about asking whether we could build the system around AI better.

And right now we’re starting to get evidence that the answer might be yes.

For now though, no new shiny features. Stage 35 is done. Now we’re going hunting for everything we got wrong before Stage 36 begins.

— Billy P.


This is the second post in Founder Notes. The first post, This Weekend, GCF Became a Startup, covered the formation of GCF FramWorks LTD. This one covers what’s happening next: stopping before Stage 36 to break the foundation on purpose.

Tested on the GCF 2.0 frozen framework, SHA eeae2bc8a426b063c687e40f9724adbb1dc3a61a. Stage 35 is finished. Stage 36 waits until the audit is done.

← All dispatches Homepage →