Self-driving agent development

Sandbox, score, and calibrate any voice or chat agent.

  • Cloud, on-prem, or air-gapped
  • Run any models
  • Zero data egress

Autonomously improve your agent and your evals

Improve your agent.

Connect

Version + tools

Exercise

Production + simulation

Evaluate

Shared Markers

Improve

Prompt + workflow

Improve how quality is measured.

Review

Machine evals

Label

Human truth

Align

Judge agreement

Refine

Better judges

Deploy Marker where your data lives.

The same image set in every mode. Choose where it runs and what it can reach.

Hosted

Marker’s cloud

Your call audio and transcripts live in Marker’s cloud; we run the stack and install upgrades. Move to your own infrastructure when your data policy requires it.

Talk to us

Your cloud

Any cloud account

Install Marker into your own cloud account, under your access controls. Audio and transcripts stay inside it; you allow-list every endpoint it reaches.

Talk to us

On-prem

Your datacenter

Connect the data plane directly to models on your network, reaching only what you allow-list. Your team owns the upgrade schedule.

Talk to us

Air-gapped

Isolated network

Egress is removed, not restricted: no phone-home, no telemetry, no online license check. Signed updates cross the boundary on media you carry in.

Talk to us

Observe agent behavior.

  • Live transcript ingest
  • Audio and timing signals
  • Tool-call and trace context
Transcript / 9:13AM Call to Help Desk processed
I need to change the address before this ships.
I can update it. Let me confirm the new address first.
tool order.update_address 409
tool order.update_address 204
Deployment
prod / us-east-2
Commit
3f9c2ba
Trace
1 error
trace 8c41d0 · 18.2s · 9 spans
call.session 18.2s
stt.transcribe 5.7s
agent.turn 3.2s
llm.completion 2.6s
tool order.lookup 0.4s
agent.turn 4.3s
tool order.update_address 409
tool order.update_address 0.7s
tts.synthesize 1.4s

Compare agent versions before release.

  • Voice and chat scenarios
  • Repeatable regression evidence

Automated alerting

A failed scenario alerts the channels your team already watches.

Teams Slack GitHub

Regression test

checkout-agent v25 · 8d21f4c

running

Scenarios

48

Simulated Users

6

Simulations

288

Refund after renewal Calm caller pass
Card declined twice Impatient caller warn
Cancel during upgrade Confused caller fail

Run the same evals in simulation and production.

  • Pass, warn, and fail assertions
  • Versioned rubrics with explanations

Did the agent verify identity before disclosure?

rev 08 · c41f9ae

LLM Judge Explanation

ChatGPT GPT 5 mini

The agent asked for the account postcode before revealing the billing address.

Value true Pass

Calibrate LLM judges with human Labels.

  • Scale QA to every call
  • Machine-human agreement
  • Auditable correction history

Machine Marks vs. Human Labels

93.5% agreement
n = 322 Human pass Human fail Judge pass 214agree 9review Judge fail 12review 87agree

21 disagreements are the next judge-improvement set.

Audio
Transcript
Tool calls
Traces
One verdict

Every signal, same evidence

Stop testing by hand. Let’s close the gap.

Talk to us
90% agent reliability
10%
edge cases · tool errors · interruptions · silent regressions