The Next Preconfiguration Prototype: A 12-Week Beta on the Agent Platforms Themselves
The Alpha answered one question. Can one short spec be compiled into every agent platform’s setup file, checked against each platform’s rules, and its setup run on a clean machine with the project’s tests? On one Linux machine, with sample repositories, the answer was yes. The next prototype, the Beta, takes on the harder question: do the files work on the platforms themselves, for real repositories, and does a team’s agent fail less often on them?
The Beta is planned at about 12 weeks with two people. It keeps the same engine and the same four commands: detect, build, check and verify. What it adds is what real use needs, from runs on each platform to more languages and a check on every pull request. Then teams’ own agents work on generated setups for four weeks.
For someone trying it, the change is easy to picture. Today you unzip the Alpha, run it on a sample repository and watch verify prove it in a container. After the Beta you’d run preconfig detect --write on your own repository, review the draft, commit the files, and let Copilot and Cursor start their next sessions on a machine that was proven the day before.

What Gets Added
Runs on the platforms. Today the files pass each platform’s schema and tools, and the setup script runs in a container. The Beta runs each file where it’s meant to run: the workflow on Copilot’s cloud agent, the Dockerfile on Cursor’s, the dev container in Codespaces, cloud-init on a fresh cloud server, and the script in the agents that take one. Every difference from verify gets explained or fixed.
Every install path. Node.js, Go, PostgreSQL other than 16, Redis 8 and Python versions installed with uv are all generated today, but the Alpha’s test machine couldn’t reach their sources. The Beta runs each one on an open network and again behind a company proxy with its own mirrors.
More languages and services. Java, Ruby and Rust as runtimes, MySQL as a service, and private package sources, each on every target and proven with verify.
More targets. Agents keep adding ways to read setup from the repository. The Beta adds a target when a platform’s format settles, starting with Codex, Claude Code and Jules.
Checks on every pull request. A GitHub Action that runs check on every pull request touching the setup, runs verify when the spec changes, and comments with what broke.
Keeping up, as a routine. Each platform’s documents and schemas compared with the knowledge base every week, with a public changelog of what moved.
The Demo, for Real
The Alpha’s demo used a sample service. The Beta uses real repositories and real agents. Three to five teams run their coding agents on setups written by Preconfiguration for four weeks, starting from detect drafts that a person reviews.
| What gets counted | Why it matters |
|---|---|
| Sessions that failed or stalled because of the setup | The cost the project exists to cut |
| Minutes of setup per session, before and after | What teams pay for in every session |
| Time to write and review a spec | Whether the spec is cheaper than the files it replaces |
| Setup files that broke without anyone noticing | The silent failures check is meant to catch |
| Draft accuracy of detect | How much of the work detect really saves |
| What teams would pay for | Whether the business case holds |
Six Places It Runs
| Place | What runs there | Why |
|---|---|---|
| GitHub Copilot’s cloud agent | The setup workflow, on pull requests and before sessions | The strictest rules and the most silent failures |
| Cursor’s cloud agents | The Dockerfile build, install and start | A build on Cursor’s side, with its own schema |
| Codespaces and VS Code | The dev container with its services | The dev container standard, prebuilds included |
| A cloud virtual machine | cloud-init at first boot | A real first boot |
| Codex, Claude Code on the web, Jules | The setup script | The agents that take a script |
| A Mac and a Windows PC | The engine, and verify with Docker Desktop | Binaries that compile but have never run |
Twelve Weeks
| Weeks | Work |
|---|---|
| 1 to 2 | The platforms set up; every Alpha test and measurement run again, with every install path on an open network and behind a proxy |
| 3 to 6 | Each target on its platform with 20 real repositories, and the fixes they call for |
| 5 to 8 | Java, Ruby, Rust and MySQL; new targets where formats have settled |
| 7 to 10 | The GitHub Action; detect on 100 public repositories |
| 9 to 12 | Four weeks with the teams, then the Beta release and a report with every figure measured again |
Two people: a lead engineer, and a second engineer who knows CI and containers well. The platforms’ subscriptions, a small cloud machine and CI minutes come to roughly $1,500 to $2,500 over the twelve weeks (an estimate).
What Counts as Done
- Each of the five targets has run on its own platform for 20 real repositories, with every difference from verify explained or fixed.
- Every install path in the spec has run on a clean machine, on an open network and behind a proxy.
- Java, Ruby, Rust and MySQL work on every target, proven with verify.
- Teams’ agents worked for four weeks on generated setups, with setup failures and setup minutes counted.
- The GitHub Action checks every pull request that touches the setup.
- Every Alpha figure measured again, and every planted bug, old and new, caught.
- Fuzzing runs every night with no open crash.
What Waits Until After
A hosted service, the team features, bases other than Ubuntu, Windows runners for Copilot and GPU setups all wait until after the Beta. So does part of the license decision: whether the spec format is published openly. The Beta’s job is narrower: to show on the platforms themselves, with real repositories and real agents, that a spec and its proof make agent sessions fail less. The roadmap has the rest.