LoreMasterBot: a WoW companion that looks things up

LoreMasterBot is a conversational World of Warcraft companion that runs in the terminal. It pairs a friendly campfire voice with tools for Blizzard’s official Game Data API. The interesting part of the build became what happens between asking a question and getting an answer.

How a LoreMasterBot lookup becomes an answer A WoW question goes to a model that chooses a tool. Python runs the handler and requests Blizzard data, or uses a cached record. The linked tool result enters conversation history, then the model writes a reply. Greetings skip the lookup. Ask about Azeroth A name, an item ID, or a question The model chooses a tool Tool name + arguments Python gets the official record Blizzard API or an in-process cache The result enters history Linked to the request by its call ID The model writes the reply

A model can sound very comfortable talking about Azeroth. That doesn’t tell me whether it found an official record, remembered something correctly, or filled in a gap. LoreMasterBot puts a lookup between the question and the spoken answer.

The Python chat loop makes two model calls for a tool-assisted turn. First, the model selects a tool and supplies arguments. Python runs the matching handler, gets Blizzard data, and adds the result to conversation history. The second call writes the reply with that result available. The model chooses the request; the application executes it.

The tools cover items, mounts, quests, achievements, spells, dungeon-journal records, playable races and classes, specializations, professions, reputation, and other structured records. That breadth creates a routing problem. A playable race, a creature record, and a dungeon-journal boss need different lookups even when the question begins with “Tell me about…”

QuestionIntended lookup
Tell me about Thunderfury.Search for an item by name.
Look up item 19019.Fetch an item by numeric ID.
Tell me about the Human race.Search playable races.
Who is Professor Putricide?Search dungeon-journal encounters.
Hello!Reply conversationally without a lookup.

One early mistake was treating short prompts as casual conversation. “Invincible” is one word, but it’s also a mount name. The current router recognizes a small list of exact conversational phrases and requests a tool for everything else. Item names and numeric IDs also have separate schemas and handlers, so “Thunderfury” doesn’t become an invalid item-ID request.

The less visible work was keeping the tool exchange intact. Some local models put tool-call JSON in their response text instead of returning structured calls. The fallback parser gives those calls unique IDs, and argument normalization keeps the saved history serializable. Invalid arguments produce an error result rather than being passed to a handler.

Multiple calls belong in one assistant message, followed by a result for each call. History trimming keeps complete turns instead of cutting through those groups. A handler exception is isolated so another valid call in the same turn can still finish. Those details matter because the next model request has to understand which result belongs to which question.

Blizzard access has its own practical limits. Requests use a ten-second timeout. Authentication retries and refreshes an expiring access token. Name searches and static records are cached for the running process; the changing WoW Token gold price has a separate sixty-second cache. A failed refresh of that price returns no result rather than presenting the old cached value as current.

Ollama is the default model provider, with Gemini and custom OpenAI-compatible endpoints available through configuration. They share the same chat loop and Blizzard handlers. Local inference still makes outbound requests to Blizzard for data; selecting a cloud model also sends the conversation and tool results to that model provider.

Compatibility needed more than changing a URL. Gemini tool calls can carry extra metadata, including a thought signature. The history builder preserves those provider fields for the follow-up request while normalizing IDs and arguments. Tests cover that round trip alongside ordinary Ollama-style calls.

The biggest boundary is in the name. Blizzard’s structured data can describe a creature, an item, or an encounter. It does not necessarily provide a character’s biography, a complete raid story, or a farming guide. The prompt explicitly tells the companion to admit when the returned data cannot answer the question and to avoid adding details from memory.

That is a design rule, not an accuracy guarantee. There is no separate checker that verifies every sentence against the tool result. Name search can also fall back to its first result when an exact match is absent. Getting official data into the conversation is one accomplishment; choosing the right record and staying within it still need evaluation.

Checked Oct 7, 2026: python -m pytest -q passed all 188 tests against this repository snapshot. These were automated tests with simulated model and API responses, not a live model comparison.

The suite checks routing mechanics, argument handling, history structure, provider configuration, API paths, cache behavior, and failure cases. Its mocked model responses cannot establish how often a real model selects the right tool or writes a faithful final answer.

The repository’s build log starts the next step deliberately: measure behavior before changing it. Five written evaluation cases cover mount and item lookups, a journal boss, a playable race, and a greeting. They describe expected behavior; the repository does not yet include scored live results for them.

I’d build that baseline next, recording the selected tool, returned data, final answer, and a human assessment of whether the answer stayed within the record. Missing biographies and ambiguous names deserve cases of their own. LoreMasterBot has made the plumbing testable. The next useful evidence is what the companion actually says.

Explore the repository ↗ · Build log ↗ · Evaluation cases ↗