Human judgment. Measured results.
Manual Evaluation
A shared corpus. Two engines. One clear comparison.
Current batch
Loading the latest batch…
Now evaluating
No active recording.
Playback is optional and does not affect evaluation.
Call volume
Labeled recordings in the current batch
Verdict mix
What each engine decided
Verified accuracy
Against your human or machine judgment
Detection latency
Median audio-to-decision time · p95 and sample count
Recording corpus
Loading recording metadata…
| Recording | Transcript | Judgment | Current batch |
|---|
Open a recording to listen, review its timed transcript, and label it. Only labeled recordings are eligible to run.
Add a recording
Upload a WAV to the shared corpus. Evaluation starts only when you run a batch.
BlandAMD is not applicable to uploads without an original provider call.
Streaming tournament
Experiments
Compare the registered shadow-only lineup on one frozen corpus and transcript profile.
Current experiment
No tournament run yet.
Observed evidence
Current leader — provisional
No eligible score yet.
| Candidate | State | Completed | Failed | Reason |
|---|---|---|---|---|
| No candidate execution state yet. | ||||
Tournament evidence
Winners
Observed rankings, matched denominators, timing evidence and unresolved limitations.
Primary evidence
No eligible leader yet
—No balanced five-second score reported.
No matched denominator reported.
Complete horizon
Best accuracy at 15 seconds
—No measured 15-second score.
Comparable timing
Fastest eligible method
—No comparable timing reported.
Leading composition
The backend-declared leader and its registered decision process.
No eligible composition yet.
No decision-process explanation has been supplied.
Accuracy versus time
Balanced correct-by-deadline on matched denominators. Missing values stay missing.
Per-class error comparison
Leading method against current Kooya and available Bland classification evidence.
| Method | Matched | Human → machine | Machine → human | Unknown | Errors | Decision-only p95 |
|---|
Full leaderboard
Completed, pending, failed and blocked candidates remain in the same table.
| Matched | Human recall | Machine recall | Unknown | Errors | Decision-only p95 |
|---|
Paired uncertainty
Paired uncertainty unavailable for the leading method versus K.
Exploratory retrospective 95% intervals on paired groups. Holm adjusts for testing multiple methods. Selection on this corpus requires a separate independent confirmation corpus.
Recording results
Current attempts, ten results per page. Run gold is frozen; review shows the current editable judgment. Refresh to reload results.
| Case | Method | Run gold | Engine outcome | Comparison | Review |
|---|