Cricket Mode: small commands for better AI coding
Cricket Mode is a small behavior layer for AI coding agents. I wanted a consistent way to keep a task focused, challenge what was built, ask for evidence, and understand the result without relearning that routine every time I changed tools.
An agent can produce a convincing explanation while missing the actual requirement. It can also turn a small request into several new abstractions, report that something works without running it, or leave me with code I can use but cannot explain. Cricket Mode turns those moments into specific things I can ask for.
I built it around seven small commands. Each has one job, and I can use one without committing to a whole workflow:
| Command | What I ask it to do |
|---|---|
/pitch | State the problem, the smallest plan, and its boundaries before editing. |
/challenge | Try to break the completed work and separate confirmed findings from unchecked concerns. |
/prove-it | Exercise the changed behavior and close with PASS, FAIL, or PARTIAL plus evidence. |
/senpai | Teach how the work fits together and what pattern I can recognize next time. |
/scrub | Identify leftover scaffolding, redundant wrappers, and other residue with specific cleanup suggestions. |
/drift | Find concrete disagreements between code, tests, documentation, and configuration. |
/chirp | Restate the previous explanation in shorter, plainer language. |
Those names use the contract’s slash spelling. In Codex, the installed skills use $pitch, $prove-it, and the other dollar-prefixed names. The behavior is meant to stay the same.
The distinction between challenge and verification matters. A review can find a suspicious path without proving it failed. A passing test can establish one behavior without establishing everything around it. Cricket asks the agent to say what it inspected, show what it ran, and leave the remaining uncertainty visible. The review commands report findings; they do not automatically rewrite the project.
Portability became the next problem. Maintaining a separate version of every instruction for every tool would make it easy for the commands to drift apart. I separated the meaning into core/COMMANDS.md and kept the adapters focused on each host’s discovery, invocation, and file layout.
The Python installer copies an adapter and a regular copy of the shared contract into the target project. Cursor uses .cursor/skills/; Codex and Antigravity use .agents/skills/. Because those last two share a location, the installer refuses to overwrite the other adapter or an existing unstamped skill. Updating refreshes the installed contract without replacing the skill files.
Verification needed its own boundaries, too. I kept three kinds of evidence separate:
| Evidence | What it establishes |
|---|---|
| Implementation | The contract, adapters, and tooling exist in the repository. |
| Scripted checks | Installation, contract copies, adapter wiring, and the bundled reply checks behave as expected. |
| Recorded live sessions | A named host performed the stated behavior in a particular session. |
The shared conformance checker reads example replies from fixtures and checks things such as required labels, scope boundaries, and a single closing status. It does not call a model or open an editor. A reply can contain the expected words and still miss the point, so a passing fixture is useful evidence about the checker rather than a guarantee about future agent output.
The repository records live pitch and prove-it sessions in Cursor, Codex, and Antigravity. In those sessions, the pitch came before editing, verification used real execution or filesystem checks, and normal prompts did not trigger Cricket. Those observations cover the recorded cases; they do not establish that every command behaves perfectly on every task.
One Antigravity pitch session also surfaced challenge as a used skill. Its documented skill frontmatter lacks an equivalent to the explicit invocation controls used by the other adapters. I kept that limitation in the documentation. Carrying the same contract between hosts still leaves differences in how those hosts select and load instructions.
That became the main lesson of the build: writing an instruction, checking its packaging, and observing an agent follow it are three different accomplishments. A portfolio claim should say which one the evidence supports.
Portable v1 is complete for the seven installed commands. The current reference skills and contract also include an optional /yolo sequencer for one pitch → build → challenge → verify pass, followed by teaching and cleanup only after PASS. It stops at findings, failed verification, or growing scope. The portable installer still copies only the original seven commands.
A separate Python policy module recommends a work lane and a small set of commands from caller-supplied task information. It does not run the checks, invoke commands, select models, or retry work, and it has not been validated inside a host application. Broader live checks are the useful next step before expanding those capabilities.
Cricket Mode helps me ask for a smaller change, a stronger check, or a clearer explanation at the moment I need it. The agent still has to do the work, and I still have to assess the evidence.
Explore the repository ↗ · Command contract ↗ · Architecture and recorded sessions ↗ · What the checks cover ↗