AI Agents & Assistants
Agents that work inside your software.
Copilots for your users, agents for your operations — built into your product, your permissions, and your data model. Not a chatbot bolted on the side.
What we build
Product copilots
Assistants inside your UI that act on your data model: drafting, summarizing, answering, executing — with your permissions enforced.
Operations agents
Autonomous workflows that classify tickets, route requests, reconcile records, and escalate to humans when confidence drops.
Voice and conversational interfaces
When hands-free or natural language is the right UX.
Supervision built in
Every agent ships with evaluation, logging, and human-review gates. Autonomy is earned, not assumed.
Proof
Amgen
Frankie, a voice assistant for Amgen’s field sales reps — in production since 2019, before the LLM era.
Read the storyOur own hiring
An agent that turns interview transcripts into client-ready profiles has run our hiring since early 2025 — a year without a dedicated operator; two to six hours per profile became about ten minutes.
- Frankie
- In production since 2019
- Hiring agent
- Since early 2025
- 2–6 h → ~10 min
- Per profile
Two kinds of agents
A copilot inside your product
Lives where your users already are. It drafts the reply, fills the form from the document, answers from the customer’s own records with the source shown, and suggests the next step — inside your UI, under your permissions, on your data model. The user stays in control; the copilot saves the first ninety seconds of every task, a few hundred times a day.
Read the storyAn agent in your operations
Works where nobody is looking: it watches a queue or an inbox, reads what arrived, classifies it, updates the record, opens the ticket, reconciles the statement — and hands the exception to a person. It runs on a schedule or on an event, inside the tools your team already uses, and it is measured in hours that are no longer spent and errors that no longer happen.
Most companies need one of each, and they are built on the same foundation: grounding on your data, actions limited to what the agent is allowed to touch, and a gate that opens only as the evidence supports it.
When teams come to us
- Support volume growing faster than headcount
- Users asking “why can’t your product just do this for me”
- An internal pilot that works in demos but nobody trusts in production
How it starts
The Discovery Sprint maps which agent creates measurable value first — or you bring a defined use case straight to implementation.
How we build it
Map the workflow
where an agent helps, and where a human must stay (the Sprint covers this)
Ground it
your permissions, your data model, the tools it may touch
Ship narrow, supervised
one workflow, review gates, evaluation before launch
Earn autonomy
scope expands as the evidence supports it
What “supervised” means
Every agent we ship starts with a person approving what it proposes. That is not a limitation to apologise for; it is how the agent learns the edge cases and how your team learns to trust it. Each correction becomes a rule. When corrections approach zero for one kind of action, the gate opens for that kind — and stays closed for the ones that still surprise you.
Grounding
the agent retrieves from your data before it acts, and its outputs point to where they came from
Scope
an explicit list of the systems and actions it may touch, enforced by your permissions, not by a prompt
Evaluation
a test set of real cases, run every time a prompt, a model, or the data changes
Audit trail
every action traceable to the input, the model, and the rule that allowed it
An agent without these is a demo with a schedule.
Sophilabs presented us with a product that we were proud to show people.
What an engagement looks like
It starts narrow on purpose: one workflow, one team, one metric. The AI Discovery Sprint finds which workflow that is and what your data can support; if you already know, you bring the use case straight to implementation. From there, a senior team scopes the agent, builds it inside your environment, runs it behind the gate with your reviewers, and hands it over with the evaluation set, the runbook, and the rules it has learned — yours to run, with us or without us. A first agent is weeks of work, not quarters; what comes after it is a decision made on evidence, not on a demo.
Works well with
FAQ
Will it hallucinate in front of our users?
Grounding, output constraints, and review gates — and we measure failure modes before launch, not after.
Can it act inside our existing tools?
That’s the whole point: your permissions, your systems. An agent beside your stack is a demo.
How much autonomy should we give it?
Less than vendors promise, more over time. Autonomy is earned with evidence, not granted at kickoff.
Which model does it use?
Whichever fits the job: a commercial model through its API for most copilots, an open-weight model on your own infrastructure when data can’t leave your environment or volume makes per-call pricing untenable, classic machine learning when the input is a table. The model is a component behind an interface your product owns; swapping it is a configuration change, and the evaluation set is what makes that safe.
Does our data leave our environment?
It doesn’t have to. Agents run inside the access you grant — your cloud, your permissions, your audit log — and commercial providers offer terms that exclude your data from training. Decide it in the architecture; we’ll say it plainly in the design.
Who maintains it after launch?
Your team, if you want: everything ships with the evaluation set, the runbook, and the learned rules. Many clients keep building with the same engineers on a monthly basis, priced on what ships, not hours; either way, the hand-off is written to work without us.