Here’s how we measure agreement, and when you’ll see the numbers.
Burna is in ongoing internal testing demonstrating strong agreement with expert clinicians. Here is exactly what that means and how we measure it.
What the literature already measured.
Published inter-rater agreement on CTCAE grading sits between kappa 0.59 and 0.68. Structured calibration pushes expert agreement to 0.85 and above. This baseline is what gives any claim about AI grading accuracy its meaning.
engine grade
Illustrative rendering
Figure 1. Grade agreement and kappa bands, rendered the way the literature renders them. Illustrative rendering; methodology below.
Sources: Hillman 2010; Le-Rademacher 2017; Hong 2020.
How we're measuring it.
Blind, adjudicated comparison against expert clinician panels on synthetic and de-identified cases, scored with weighted kappa and per-grade agreement, including the Grade 3+ boundary where the clinical stakes concentrate.
Figure 2. The comparison protocol.
The concordance study runs across three campuses.
Burna AI is a participant in Mayo Clinic Platform_Accelerate. A concordance study is running across three campuses against an adjudicated two-clinician reference panel, on 1,600 cases drawn from a de-identified dataset of 3 million patients, with Andrea Pirzkall as co-principal investigator. Results read out in December 2026.
A de-identified dataset of
3,000,000 patients
Cases sampled for grading
1,600
Campus
One
Campus
Two
Campus
Three
Reference standard
Two clinicians grade every case independently
Every disagreement is
adjudicated by the panel
Readout
December 2026
Figure 3. The concordance study as designed. Andrea Pirzkall, co-principal investigator. No result is shown here because none has read out.
“Rather than just declaring it good, because then they'll be skeptical and say, oh, we don't need this, you need compelling data. And the best way to do that is to do a formal randomized trial.”
Where the numbers will appear
Accuracy claims belong in peer-reviewed venues. When our results publish, that is where you will read them first. Until then, this page carries the full methodology so your own statisticians can judge the design.
We publish the ledger.
We keep a working ledger of the hard cases:
- 01ambiguous negations
- 02syndrome-level grading
- 03grades where the evidence ran thin
Each one is tracked, characterized, and worked down.