TinyTalk: What happens when AI memory isn’t one pile of text
TinyTalk started as a small local chatbot so I could revisit some older under-the-hood work. The question that took over was more practical: what happens if you stop treating everything the model has “seen” as one growing pile of text and instead keep identity, recent conversation, explicit facts, historical facts, and structured relationships as separate systems?
That question matters in real workplaces. When someone asks an AI tool a question about a process, a customer, or a past decision, they usually need to know three things: what the tool currently believes, what used to be true, and where the answer is coming from. A single undifferentiated transcript makes those distinctions hard.
The separation I tested
TinyTalk keeps several distinct pieces:
- Soul — persistent identity and behavior instructions (loaded every turn, not retrieved by similarity).
- Short-term context — the last 10 turns only. Older turns leave the live window on purpose.
- Explicit facts — things a person deliberately asked it to remember. These are treated as stronger evidence than something recovered from an old conversation.
- Fact history — previous values of the same fact, so “what used to be true” stays distinct from “what is current.”
- Saved conversations — searchable only on request, not on every ordinary question.
- Knowledge graph — simple structured relationships (for example, a single-value field like preferred name or a test spaceship name).
A few concrete behaviors followed from that split:
- Some profile questions (“What is my name?”) are answered from the current profile without calling the model at all.
- Fact replacement is staged and recoverable. A new value is not called “current” until the update finishes. The status line after a “
Remember this:” command is the one to trust, not the model’s earlier reply. - Similarity search has a floor. A memory that is technically related is not automatically useful.
/sourceslists what was actually supplied with the last answer (type, how it was selected, status, short preview). It does not claim the model used every item, and it does not claim the answer came only from model knowledge.
What broke and what that showed
Small models do not name relationships consistently. “test_spaceship” and “spaceship_name” needed normalization so an imaginary spaceship would not overwrite the real one. Broad questions like “What do you remember about me?” are not close enough to any single fact for similarity search to work well, so TinyTalk reads the current fact list directly for those cases.
Fact updates are not a single transaction. Until the new value is moved into current facts, the system continues to treat the previous finished value as current and labels the new one unfinished. That was deliberate. Calling something current before the records agree creates exactly the kind of silent inconsistency that is hard to catch later.
There is also a character budget per request. Oldest turns and optional retrieved records are dropped first. The current profile and any pending update are not optional. If even those plus the soul and the new message will not fit, the turn is refused and nothing is saved.
Why this is useful for enablement work
The interesting part is not the code. It is the set of distinctions that turned out to matter:
- Explicit facts are stronger than inferred or recovered ones. People need a clear way to say “remember this” and a clear signal that it was stored.
- Current and historical values must stay separate. An old fact that used to be true should not compete equally with the present one.
- Provenance and status signals reduce over-trust.
/sourcesand theMemory:status line make the difference between “the model said so” and “this was supplied and marked current.” - Some questions are better answered by a direct lookup than by generation. That is a reliability feature, not a limitation.
- Unfinished or conflicting records should stay visible as unfinished or conflicting. Silent promotion is worse than an honest “not settled.”
In a real adoption setting these map to practical coaching points: when to treat an AI suggestion as provisional, how to check whether a fact is still current, what “it remembered” actually means, and where a human still needs to confirm before acting. The same questions come up with Copilot agents, Now Assist, or any tool that mixes recent conversation, retrieved documents, and saved preferences.
Constraints and limits
This is a small personal experiment. Memory writes are not transactional. History tracking is narrow. Similarity thresholds and profile fields were tuned by watching specific failures. Nothing here is a production memory system or a general claim about how any particular commercial tool works. The value is the set of failure modes and design choices that became visible once memory was no longer one undifferentiated pile of text.