Short notes on what shipped, what broke, and the next thing to work on. Newest week first.
Work so far ·
Build log, week of Oct 5, 2026
Making the work easier to find. I’ve spent a lot of time building things lately. This week I finally spent some time on what happens when somebody tries to find them.
TinyDesk was sitting too far down the Builds page. Home sent people toward an essay before it showed them the practical AI work. The projects were there, but the order wasn’t helping explain what I’m doing.
Shipped
Put TinyDesk, TinyTalk, and Commit City first on Builds, with direct case-study links from Home.
Added a short introduction connecting my IT operations background to the AI tools and automation I’m building.
Finished the Lab notes and replaced public editing placeholders with actual copy and project images where available.
Reworked the GitHub profile README around the same projects and updated the site’s Cricket portrait.
Added TinyTalk’s current-profile resolver, explicit /recall, and /memory inspection, with selection and status information in /sources.
Broke / learned
Writing the project descriptions made me slow down and be specific. TinyTalk’s /sources command shows which memories were supplied with an answer. It doesn’t prove which ones the model used. TinyDesk can draft a reply, but a person still reviews and saves it. Those details matter when I’m explaining what I built.
The portfolio needed to make those decisions visible, too.
TinyTalk found a memory. Was it the right one?
I wanted TinyTalk to remember things between conversations. Then I had to deal with the fact that remembering something and knowing whether it’s still true are two different problems.
An old conversation can contain an earlier name, an assistant’s assumption, or a fact I’ve since replaced. I don’t want all of those competing equally every time I ask a question.
TinyTalk now reads five defined profile fields directly from its structured relationships and saved facts. Old conversations come back through an explicit /recall request. /memory shows the current profile state, and /sources shows the records supplied with the last successful answer.
The unfinished-update handling from the previous week matters here, too. If a save only gets partway through, the proposed value shouldn’t quietly become a confirmed fact.
The application has to decide what counts as current, historical, conflicting, or unfinished before asking the model to write a useful answer. Otherwise, I’m handing it a pile of text and hoping it sorts everything out.
This is still a small memory experiment with explicit limits. But I can inspect more of what it’s doing now, which makes the next failure a lot easier to investigate.
Next
Keep the build log connected to the actual experiments and problems behind the finished pages. Check navigation on smaller screens, collect more real project captures, and practice explaining one project from input to output without needing the README to do all the talking.
Valid JSON. Still a questionable answer. TinyDesk gave me a pretty useful reminder this week: a model can return exactly the format I asked for and still give me a reply I need to fix.
Some drafts asked for information the ticket already contained. Others needed a closer look at what had actually happened in the activity history. Getting the response into the right shape was only one part of the job.
Built / tested
Built TinyDesk’s ticket queue, editable fields, activity history, internal notes, and reply composer, with separate save actions.
Worked through local model comparisons with Llama, Phi, and Qwen, then connected Grok and tested it on fictional help-desk tickets.
Kept category suggestions and reply drafts under human review, with a clear distinction between local and cloud analysis.
Fixed a context-limit problem by keeping the complete ticket history in the browser and constructing a smaller request for the model.
Connected TinyTalk to Grok while retaining its identity instructions, memory, and recent conversation. Added source visibility and worked through the distinction between current facts and historical memories.
Audited TinyTalk and improved fact replacement, unfinished-update handling, and request-size limits.
Broke / learned
Grok produced more usable drafts in our small checks, but I still had to read them. The comparison also made the evaluation question clearer: would I actually use this answer while working the ticket?
TinyTalk raised a related problem. Finding an old conversation doesn’t make everything in it a current fact. Retrieval needs rules about what the stored information means.
Where that left things
TinyDesk needed more work on repeated questions and draft consistency. TinyTalk needed clearer boundaries between current facts, old conversations, and unfinished memory updates. Both gave me specific problems to investigate next.
Three wizards and a chatbot that grew a memory. I gave different coding models the same wizard idea and ended up with noticeably different results. That was interesting enough to keep going.
The experiment moved from a pixel wizard into a Three.js version, then into the less glamorous work of making spell casting behave consistently. Effects, timing, projectiles, impacts, and the camera all had to agree about what was happening.
Continued the 3D wizard through several passes on casting behavior, effects, and camera response.
Worked on a character-asset contract and validator for a future authored model.
Completed Cricket Mode’s portable v1: seven small commands, a shared behavior contract, adapters for Cursor, Codex, and Antigravity, and install/update tooling.
Started rebuilding an old college chatbot with Python and a local model through Ollama. That became TinyTalk.
Separated recent conversation from persistent memory, added identity instructions and explicit facts, and started tracking changing facts through a knowledge graph and fact history.
Broke / learned
A visually convincing first result still leaves plenty to figure out. With the wizard, I kept running into the connections between parts of the scene. With TinyTalk, I started asking what should happen after the conversation disappears from the model’s context.
“Remember this” sounds simple until the thing being remembered changes.
Taking Cricket Mode between tools
I kept wanting the same things from coding agents: narrow the scope, challenge the result, show me the evidence, and help me understand what just got built. Cricket Mode turns those habits into small commands I can carry between tools.
The shared contract defines what a command means. The adapters handle how each tool finds and invokes it. That let me keep one set of behaviors without pretending Cursor, Codex, and Antigravity work exactly the same way.
The checks also needed clear labels. Installation and conformance scripts check the files and bundled reply cases; they don’t open the host applications or prove every model will follow the instructions. Separate live sessions exercised explicit pitch and prove-it calls in all three hosts. One Antigravity pitch run also surfaced challenge, which went into the documented limitations.
An optional orchestration policy came next, but it doesn’t route models, run commands, or retry tasks on its own. Keeping that boundary clear was part of the work.
The wizard had groundwork for a future character asset, with integration still ahead. TinyTalk was turning into an experiment about what to save, what to retrieve, and how to tell an old fact from a current one.
The demo has to survive leaving my computer. This week was a lot of work on the parts that aren’t obvious in a screenshot: deployment, data labels, and what happens between a model asking for something and the application returning it.
Built / improved
Fixed Commit City’s deployed contribution endpoint and added live-demo media and sharing metadata.
Expanded Mutation Microscope’s provenance into six evidence classes across the data, validation, interface, and documentation. Added user-flow coverage and public-showcase improvements.
Clarified that Mutation Microscope’s public app serves committed data. Optional live AlphaGenome enrichment belongs to the developer pipeline; it isn’t inference happening in the visitor’s browser.
Returned to LoreMasterBot, a chatbot I started in April, and repaired malformed tool-call history so results stayed linked to their originating calls.
Improved multi-tool handling, API timeouts, caching, argument validation, and regression coverage. Reduced repeated code, moved authentication out of import time, and added configurable model providers.
Broke / learned
Commit City’s local setup and deployed server didn’t resolve imports the same way. The serverless function needed explicit module extensions. A passing local build was useful evidence, but I still had to check the deployed API.
Mutation Microscope raised a different question: what exactly is this number? A published value, a calculation, a reconstructed figure, and an illustrative example shouldn’t look equally authoritative just because they share a polished interface.
LoreMasterBot made me look more closely at the conversation plumbing. A tool result needs the right call ID, usable arguments, and a history format the provider can accept. Asking the model to behave better doesn’t repair those connections.
Where that left things
The projects had clearer boundaries and more specific checks. Commit City needed to work at its public URL, Mutation Microscope needed to explain its evidence, and LoreMasterBot needed to handle the tool cycle dependably before I kept expanding it.
A page falling apart, a city taking shape, and a microscope for data. The starting questions were different, but each project gave me something I could interact with, inspect, and improve.
Built / tested
Built Gravity Well: a quiet webpage whose own interface elements become textured Three.js bodies, move through a gravitational field, and collapse toward a singularity.
Added lensing and staged collapse, then improved background capture, activation timing, and the legibility of the distortion.
Started Commit City, turning a GitHub contribution calendar into a 3D skyline. Added parsing tests and refined the camera and loading experience.
Added scenery around the contribution city, then pulled it back and restored contrast when it competed with the towers.
Built Mutation Microscope’s initial interactive observatory and began improving provenance, testing, and accessibility around the displayed data.
Broke / learned
Gravity Well needed to capture the normal page before the transition styles changed it. It also needed protection against overlapping activations. Making the page fall apart was fun; making it recover and release its resources was part of making the experiment usable.
Commit City’s scenery could get more dramatic, but the contribution history still had to be the subject. More visual detail didn’t automatically make it a better visualization.
Mutation Microscope brought that same question back to the data: an interesting display still needs a clear explanation of where its values came from.
Where that left things
I had working foundations to keep refining. The next work was making them easier to trust and share: check the deployed behavior, clarify the data, and fix the parts that only show up when somebody actually uses the thing.