65% of support tickets now resolved end to end without a human, with satisfaction scores holding.
A support agent trained on the company's own resolved tickets and product docs, answering in the company's own voice.
- Fixed scope · 6 weeks
- 2025
- SaaS
- AI agent, Integration

The situation
The support team was answering the same questions every day. Ticket volume grew with the customer base, the team did not, and the hard tickets — the ones worth a human — sat behind a queue of password resets and billing questions.
They had tried an off-the-shelf chatbot. It answered confidently and wrongly, customers noticed, and it was switched off within a month. That failure mattered: the bar was no longer "deflect tickets", it was "never damage a customer relationship to deflect a ticket."
What we built
An agent trained on their own material — years of resolved tickets, the product documentation, and the tone guidelines their team already wrote to.
Before it answered anyone, it ran against a held-out set of real past tickets, and we compared its answers to what the team had actually replied. That evaluation pipeline stayed in place after launch, so a model change is measured rather than hoped about.
Every answer cites the documentation it came from, so a human agent can check it in one click rather than re-deriving it.
What we deliberately left out
No answering when it is not sure. The agent hands off to a human rather than guessing. Coverage went down; trust stayed intact, which was the actual requirement.
No account actions. It answers questions. It does not issue refunds, cancel plans or change account state — those routes stayed with the team.
No replacing the helpdesk. We integrated into the tool they already used instead of asking a working team to move.
The result
Roughly two thirds of incoming tickets now close without a person touching them, and satisfaction scores did not move — which, given how the previous attempt went, was the number the team was watching. The support team spends its day on the third of tickets that need judgement.
Built with
- Python
- FastAPI
- OpenAI fine-tuning
- Pinecone
- Next.js
Have something like this?
Tell me what the problem is on a 30-minute call. If it's a fit, discovery starts with a written scope and a fixed number — not an estimate that moves.