Evals for voice agents in your cloud

Simulate, score, and calibrate any agent.

  • On-prem deployable
  • Zero-egress air-gapped mode
  • Customer-owned data plane

The primitives and automation behind both eval loops.

Loop 01

Improve your agent.

01 Version

Connect the agent you are changing.

02 Exercise

Run real traffic and lifelike simulations.

03 Evaluate

Apply the same Markers to every transcript.

04 Improve

Fix the prompt, tools, model, or workflow.

Loop 02

Improve how quality is measured.

01 Review

Route the right machine evals to humans.

02 Label

Capture the answer a human stands behind.

03 Align

Measure judge agreement on the same coordinate.

04 Refine

Strengthen judges, simulations, and monitors.

Choose where Marker runs.

T0 01

Hosted

Marker operated

Start without infrastructure work.

Hosted

Marker operates the full stack and handles upgrades. Start evaluating agents the same day, then move modes later without changing product.

Book Demo
T1 02

Your cloud

Customer VPC

Keep the data plane in your account.

Your cloud

Deploy the same image set into your own VPC. Transcripts, audio, and evidence stay inside your account and your access controls.

Book Demo
T2 03

On-prem

Customer network

Run beside private systems and models.

On-prem

Run the data plane inside your own network, next to the private systems, telephony, and models your agents already depend on.

Book Demo
T3 04

Air-gapped

Zero internet egress

Operate without a phone-home path.

Air-gapped

The same images run with outbound egress removed entirely. No phone-home path, no cloud dependency, no separate build.

Book Demo
01

See the conversation your customer actually had.

  • Live transcript ingest
  • Audio and timing signals
  • Tool-call and trace context
Transcript / 9:13AM Call to Help Desk processed
I need to change the address before this ships.
I can update it. Let me confirm the new address first.
tool order.update_address 409
tool order.update_address 204
Deployment
prod / us-east-2
Commit
3f9c2ba
Trace
1 error
trace 8c41d0 · 18.2s · 9 spans
call.session 18.2s
stt.transcribe 5.7s
agent.turn 3.2s
llm.completion 2.6s
tool order.lookup 0.4s
agent.turn 4.3s
tool order.update_address 409
tool order.update_address 0.7s
tts.synthesize 1.4s
02

Compare agent versions before release.

  • Voice and chat scenarios
  • Version-pinned batches
  • Repeatable regression evidence

Automated alerting

A failed scenario alerts the channels your team already watches.

Teams Slack GitHub

Regression test

checkout-agent v25 · 8d21f4c

running

Scenarios

48

Simulated Users

6

Simulations

288

Refund after renewal Calm caller pass
Card declined twice Impatient caller warn
Cancel during upgrade Confused caller fail
03

Run the same evals in simulation and production.

  • Boolean, numeric, and category outputs
  • Pass, warn, and fail assertions
  • Versioned rubrics with explanations

Did the agent verify identity before disclosure?

rev 08 · c41f9ae
Evaluator LLM judge
Output Boolean
Assertion false → fail

Judge explanation

ChatGPT GPT-4o mini

The agent asked for the account postcode before revealing the billing address.

Value true Pass
04

Calibrate automated judges with human Labels.

  • In-context review
  • Machine-human agreement
  • Auditable correction history

Alignment

Machine Marks vs. Human Labels

93.5% agreement
n = 322 Human pass Human fail Judge pass 214agree 9review Judge fail 12review 87agree

21 disagreements are the next judge-improvement set.

01 Audio

LoudnessSilenceTalk time

02 Conversation

InterruptionsLatencyPacing

03 Language

OutcomePolicyTone

04 Execution

Tool callsErrorsTrace path

One verdict

Every signal, same evidence

Stop testing by hand. Let’s close the gap.

Book Demo
90% agent reliability
10%
edge cases · tool errors · interruptions · silent regressions