Evidence · Study
Evidence study: Ustadi Demo Group — what AI-assisted interviews reveal about human judgment
Anonymized, aggregated snapshot drawn from live platform records: panels, flags and hours.
12
AI-assisted interviews
scorecards captured with the panel
0.82
Panel calibration (κ)
agreement beyond chance across panelists
8%
Bias flags raised
leading-question and talk-time checks
94%
Feedback loop closure
15 of 16 raised issues reported back
340
Hours returned
from the hours-returned metric only
What we measured
Twelve AI-assisted interviews in Ustadi Demo Group, each scorecard drafted by the model and signed off by a named human. Panel talk-time and question style were checked as the interviews went.
What the numbers say
Panels agreed beyond chance (κ = 0.82), which means structured scoring landed the same judgment across interviewers. Bias flags fired on 8% of sessions — every one logged for a human to weigh, never applied by the model. On the voice side, 15 of 16 raised issues were acted on and reported back to the people who raised them.
What it means
Consistent panels, visible flags and closed loops are the pattern of a team whose tooling amplifies judgment instead of standing in for it. The human signs every scorecard, every approval, every closure — and 340 hours returned to that work.
Live aggregates — seeded cohort snapshot · Reviewed and published by Tara Njeri on 2026-10-03 · Aggregated and anonymized — no individual is identified.