Here’s how we measure agreement, and when you’ll see the numbers.

Burna is in ongoing internal testing demonstrating strong agreement with expert clinicians. Here is exactly what that means and how we measure it.

What the literature already measured.

Published inter-rater agreement on CTCAE grading sits between kappa 0.59 and 0.68. Structured calibration pushes expert agreement to 0.85 and above. This baseline is what gives any claim about AI grading accuracy its meaning.

expert grade
G0G1G2G3G4
G0G1G2G3G4

engine grade

Illustrative rendering

published inter-rater baseline
structured calibration
00.20.40.60.81.0
kappa, inter-rater agreementIllustrative rendering

Figure 1. Grade agreement and kappa bands, rendered the way the literature renders them. Illustrative rendering; methodology below.

Sources: Hillman 2010; Le-Rademacher 2017; Hong 2020.

How we're measuring it.

Blind, adjudicated comparison against expert clinician panels on synthetic and de-identified cases, scored with weighted kappa and per-grade agreement, including the Grade 3+ boundary where the clinical stakes concentrate.

CasesBlind expert panelAdjudicationWeighted kappa + per-grade agreement

Figure 2. The comparison protocol.

The concordance study runs across three campuses.

Burna AI is a participant in Mayo Clinic Platform_Accelerate. A concordance study is running across three campuses against an adjudicated two-clinician reference panel, on 1,600 cases drawn from a de-identified dataset of 3 million patients, with Andrea Pirzkall as co-principal investigator. Results read out in December 2026.

A de-identified dataset of

3,000,000 patients

Cases sampled for grading

1,600

Campus

One

Campus

Two

Campus

Three

Reference standard

Two clinicians grade every case independently

Every disagreement is

adjudicated by the panel

Readout

December 2026

Figure 3. The concordance study as designed. Andrea Pirzkall, co-principal investigator. No result is shown here because none has read out.

“Rather than just declaring it good, because then they'll be skeptical and say, oh, we don't need this, you need compelling data. And the best way to do that is to do a formal randomized trial.”

Charles Balch, MD · former EVP and CEO, ASCO

Where the numbers will appear

Accuracy claims belong in peer-reviewed venues. When our results publish, that is where you will read them first. Until then, this page carries the full methodology so your own statisticians can judge the design.

We publish the ledger.

We keep a working ledger of the hard cases:

  • 01ambiguous negations
  • 02syndrome-level grading
  • 03grades where the evidence ran thin

Each one is tracked, characterized, and worked down.

Bring your hardest questions. The methodology is built for them.