Testing Framework with Simulation & Evaluation Agents
A harness that replays real conversations against an agent and has a second model score each transcript against the client's own criteria.
- Role
- AI Engineer Intern
- Organization
- Atria
- Location
- Pittsburgh, USA
- Industry
- Artificial Intelligence / SaaS
- Timeline
- Jun – Aug 2025
- Stack
- OpenRouter
- OpenAI
- JavaScript
- AWS Lambda
- MongoDB
| Open | Date | greeted | tone_score | resolution |
|---|---|---|---|---|
| open the run | Sep 24 | true | 8.6 | a wrapped note |
| open the run | Sep 24 | true | 7.1 | a wrapped note |
| open the run | Sep 23 | false | 4.2 | a wrapped note |
| open the run | Sep 23 | true | 9.0 | a wrapped note |
| All runs | 68% true |
Problem
An agent that talks to customers cannot be tested the way ordinary software is. There is no single correct output, and what breaks is usually tone, a missed step or a compliance rule rather than a crash, so the failure is a judgement rather than an exception. Teams fall back on reading transcripts by hand, which does not scale and never catches drift. Atria was an early stage startup where this was still being worked out, and quality was checked by manual review while the platform itself was being defined.
Solution
The design separates the system under test from the system that judges it. A scheduler decides which test sets are due and dispatches them without waiting for a result. A simulation replays a real customer thread against the agent under test, turn by turn, until the agent signals it is finished. A second model, never part of the conversation, then reads the transcript and scores it against a typed set of criteria the client defines. Results are stored with the schema that types them, so the table a team reads is its own standard rather than a generic one.
- moves the work along
- holds the sets, criteria and results
- runs and judges
- shows
- the work moving along
- stored data, read and written
- the transcript under judgement
What made this possible
Two models, kept apart
The agent under test and the judge are separate systems, so the thing being measured never has a say in its own score. Collapsing them would make every result suspect.
Criteria as typed data
A client's standard is a named, typed list rather than a paragraph of guidance. That is what makes runs comparable, and what lets the results table be built from the standard itself.
An explicit end to a conversation
The agent signals that it is finished with a tool call, which hands the harness a definite transcript. Without that boundary there is nothing stable to judge.
Impact
Checking for regressions became something that runs on a schedule rather than something a person remembers to do. Each client's definition of a good conversation is written down as criteria and scored the same way on every run, so a shift in tone or a broken compliance rule shows up as a score that moved rather than as a complaint.