TinyDesk: building an AI assistant for a help desk
TinyDesk is a small help desk demo built around fictional Ashmere Press tickets. I wanted to explore something familiar from support work: understanding what has already happened on a ticket, spotting what’s missing, and preparing a useful response.
A ticket rarely ends with its original description. Someone adds troubleshooting notes, a repair appears to work, and the requester comes back days later with the same problem. An assistant needs to follow that history without treating an earlier success as the current outcome.
I built TinyDesk with AI coding tools, using a lightweight JavaScript frontend and Python backend. The interface has a navy-and-green design, a searchable ticket queue, activity history, internal notes, and a reply composer. TinyTalk Assist adds summaries, category suggestions, clarification questions, and editable reply drafts.
For AI enablement, the useful questions are practical: where does a suggestion fit in the workflow, what must the person check, and how do we judge whether the answer helps? TinyDesk gives me a place to explore those decisions with fictional data. My workplace agent and coaching notes describe separate experience; the evaluation below belongs to this demo.
The technician keeps control. Applying a suggested category changes the form; saving it is another action. Using a draft fills the composer, where it can be edited before saving. Replies are stored in the demo rather than emailed.
The assistant supports local models through Ollama and optional cloud analysis through the Grok API. The interface identifies which provider is configured and explains when ticket information goes to xAI. The API key stays in backend configuration.
One backend problem needed attention early: long ticket histories could exceed the local model’s available context. TinyDesk now creates a bounded copy for analysis while preserving the complete saved ticket. Shortened text is marked, and omitted older activity is counted. That budget is calibrated for the default Qwen model’s observed 4,096-token runtime context.
Then came answer quality.
Early checks suggested Qwen handled activity history and verified resolutions better than Llama. It still made some frustrating mistakes, including calling a named folder unspecified and asking the requester to supply its name again.
That became a useful lesson: valid JSON tells you the response has the expected shape. You still have to read what it says.
The later comparison used frozen fictional tickets and written expectations. We reviewed summaries, questions, and drafts for factual contradictions, invented details, repeated questions, and unsupported commitments. Grok handled the targeted mistakes better while still asking for clarification when information was genuinely missing.
I also compared Grok’s high and low reasoning settings:
| Twelve-ticket evaluation | High reasoning | Low reasoning |
|---|---|---|
| Median response time | 31.34 seconds | 12.96 seconds |
| Total API cost | $0.210136 | $0.108856 |
Low reasoning produced ten drafts rated usable as written and two needing edits, with none rejected. Each configuration ran once per case, so these results describe a small evaluation rather than a general accuracy guarantee.
Low reasoning became the default, with human review still required. The next feature is suggested next steps, followed by knowledge-base lookup. Duplicate-ticket suggestions and related-ticket patterns are also on the roadmap, with separate evaluations planned for those capabilities.