Mutiny: building a local AI console with clear boundaries

Mutiny is a single-user AI console built around local Ollama models, conversation history on disk, and explicit research actions. I wanted control over where the work happens, what gets saved, and when a request is allowed to reach beyond my machine.

Mutiny’s local console showing a fictional repair-workshop conversation, a closed research worksheet with its question, queries, answer, gaps, and saved-fact source, and the Memory panel.

The build centers on a Python FastAPI backend and a static JavaScript interface. Threads, messages, settings, saved facts, tool runs, and source records live in SQLite. The interface includes a conversation list, model and personality settings, a Memory panel, a research worksheet, and controls for local scheduled jobs.

Local chat and research needed separate paths. A normal Send submits the thread history to the configured loopback Ollama endpoint with no tools attached. It does not silently search the web or turn a message into a research run. Saving a fact and asking over saved notes are explicit Memory actions.

The current composer also has a Web search switch. Turning it on changes the action from Send to Search and routes the question to a web research run. The switch starts off and resets on page load. That makes the choice visible at the point where the request changes behavior.

ActionWhere it gets informationWho writes the response
Normal SendThe current conversation contextA local Ollama model, with no tools attached
Research notesSaved facts, available local memory, and files in the configured research folderA local Ollama model given the retrieved excerpts
Explicit Web searchWeb snippets returned through the local SearxNG serviceA local Ollama model given those snippets

There are two conditions for web research: outbound research must be enabled in configuration, and the run must request web mode. The configuration flag defaults to enabled. Setting MUTINY_OUTBOUND_ENABLED=0 hides and disables web research while leaving closed research available.

A search service running on localhost still sends queries to the public web. Mutiny keeps the writer local, but an explicit web search is an outbound action. The web path retrieves snippets through SearxNG and does not crawl full pages or combine the search with private local facts.

Research needed more than a useful-looking answer. Each run stores its question, mode, queries, model, and reported gaps alongside separate source records. The worksheet lets me inspect the excerpts and export the write-up instead of losing its origin inside a chat bubble.

Closed research reads saved facts and local documents. Optional MemPalace indexing is used only when the embedding files are already cached; otherwise recall falls back to SQLite facts. The application does not download those files to make the feature appear ready.

If retrieval finds no usable sources, the research pipeline reports the gap and skips the answer-writing call. If it finds sources, the local writer is instructed to answer from the numbered excerpts. Output filtering removes external URLs in closed mode. Web mode retains only URLs attached to that run, and both paths remove out-of-range numbered citations.

That is a useful boundary, but it does not establish that every generated sentence follows from its citation. A model can still misread a retrieved excerpt or add an unsupported claim. Source inspection and answer-quality evaluation remain part of using the tool.

The privacy work also reached below the interface. The server rejects non-loopback binding, and inference validates the Ollama endpoint and rejects provider-qualified or known cloud-backed model identifiers. The API uses a process-bound HttpOnly session cookie, checks the request origin, and requires a mutation header before changing state. Browser assets are bundled locally, and external documentation pages are disabled.

Dependency behavior mattered too. The pinned inference library can fetch a remote cost map during import unless its local-map setting is applied first. Mutiny’s privacy bootstrap sets that and the offline settings before loading the library. The tests check fresh imports for unexpected non-loopback lookups rather than relying on a reassuring configuration name.

The console cannot reconfigure an already running Ollama daemon. Its OLLAMA_NO_CLOUD=1 setting belongs in the daemon’s own environment. These controls describe a personal localhost application; they are not an account system or a design for public hosting.

Reliability needed attention alongside privacy. Completed chat requests can return their stored result when retried with the same request ID. Failed turns remain marked as failed. Resetting model context keeps the visible transcript, while clearing history is a separate action that leaves saved facts alone.

Local automation stays narrow. The Jobs drawer can schedule the local morning briefing, pause or resume it, and inspect its saved runs. The scheduler allows that specific tool rather than arbitrary shell commands or web research. Older database upgrades create a backup before migration, and eligible legacy schedules are imported paused for review.

I checked the current repository snapshot on Oct 6, 2026. All 143 tests passed. The suite covers request guards, local model restrictions, chat persistence, memory fallback, migrations, scheduler behavior, and research source handling. Its seeded research evaluation includes twenty hit-and-miss cases with test doubles. Those checks establish behavior under the fixtures; they are not a benchmark of a live model’s answer quality or a complete security audit.

The main lesson was that “local” needs to be visible in the execution paths. A saved note, a model response, a scheduled job, and a web request have different boundaries. Keeping those differences explicit made the product easier to reason about and gave the tests something concrete to check.

The next useful evaluation would use real local-model responses against a fixed set of notes, with separate scores for retrieval, supported claims, missing information, and citation accuracy. That would test the part the current offline fixtures cannot establish.

Explore the repository ↗ · Research run contract ↗ · Tests at the checked snapshot ↗