← Projects

Testing Framework with Simulation & Evaluation Agents

A harness that replays real conversations against an agent and has a second model score each transcript against the client's own criteria.

Role
AI Engineer Intern
Organization
Atria
Location
Pittsburgh, USA
Industry
Artificial Intelligence / SaaS
Timeline
Jun – Aug 2025
Stack
  • OpenRouter
  • OpenAI
  • JavaScript
  • AWS Lambda
  • MongoDB
Agent testing · last runscoredby criteria
Agent testing
3What came out
Columns ▾
Last 14 days ▾
OpenDategreeted tone_scoreresolution
open the runSep 24true8.6a wrapped note
open the runSep 24true7.1a wrapped note
open the runSep 23false4.2a wrapped note
open the runSep 23true9.0a wrapped note
All runs68% true
Every column after Date is one criterion the company defined, in its stored order.
The three screens the system has, in the order they matter. Sets is what replays and how often, Criteria is the client's own definition of a good conversation, and Results is what came out, scored against that definition rather than a generic one. Pick a tab.

Problem

An agent that talks to customers cannot be tested the way ordinary software is. There is no single correct output, and what breaks is usually tone, a missed step or a compliance rule rather than a crash, so the failure is a judgement rather than an exception. Teams fall back on reading transcripts by hand, which does not scale and never catches drift. Atria was an early stage startup where this was still being worked out, and quality was checked by manual review while the platform itself was being defined.

Solution

The design separates the system under test from the system that judges it. A scheduler decides which test sets are due and dispatches them without waiting for a result. A simulation replays a real customer thread against the agent under test, turn by turn, until the agent signals it is finished. A second model, never part of the conversation, then reads the transcript and scores it against a typed set of criteria the client defines. Results are stored with the schema that types them, so the table a team reads is its own standard rather than a generic one.

ScheduleScheduleron a timerReads every test set config,keeps the ones enabled and duefrequencies: 24 h, 7 days, 14 daysTest setsOne config per set: the persona,the frequency, the conversationsto replay and its documents3 personas · up to 5 documentsevery test set configSimulateSimulation runOne call per conversation, allin parallel; the scheduler neverwaits for an answer8 minute ceiling per runfire and forget batchthe set to runAssistantunder testThe live assistant, driven turnby turn with its own tools untilit calls the exit toolup to 5 tool calls per turnCustomer threadsA tool pulls a real thread, sothe run replays a conversationthat actually happenedread during a run, never writtenone run per conversationone real threadJudgeJudgeseparate modelA second assistant reads thetranscript and scores it, andnever joins the conversationpolled up to 3 minutes per runCriteriaNamed criteria the admin paneldefines and orders, each typedstring, boolean or number3 value types · one set per companytranscript on the exit toolthe criteria schemaSeeResultsOne record per run, its scoreper criterion written back whenthe judge returnsread back by date rangeAdmin APIServes the results with theschema that types them; its Runbutton fires the same batchREST · test sets, results, criteriaone score per criterionresults by date rangeResults tableA column per criterion, in thestored order: checks, a heatmapper number, text per notethe table follows the schemaone row per run
Two models, kept apart on purpose. The assistant under test answers turn by turn until it calls the exit tool; a second one then reads the transcript and scores it against the company's own criteria, and never joins the conversation.

What made this possible

  • Two models, kept apart

    The agent under test and the judge are separate systems, so the thing being measured never has a say in its own score. Collapsing them would make every result suspect.

  • Criteria as typed data

    A client's standard is a named, typed list rather than a paragraph of guidance. That is what makes runs comparable, and what lets the results table be built from the standard itself.

  • An explicit end to a conversation

    The agent signals that it is finished with a tool call, which hands the harness a definite transcript. Without that boundary there is nothing stable to judge.

Impact

Checking for regressions became something that runs on a schedule rather than something a person remembers to do. Each client's definition of a good conversation is written down as criteria and scored the same way on every run, so a shift in tone or a broken compliance rule shows up as a score that moved rather than as a complaint.

← Back to projects