Ustadi

Evidence · Study

Evidence study: Ustadi Demo Group — what AI-assisted interviews reveal about human judgment

Anonymized, aggregated snapshot drawn from live platform records: panels, flags and hours.

12

AI-assisted interviews

scorecards captured with the panel

0.82

Panel calibration (κ)

agreement beyond chance across panelists

8%

Bias flags raised

leading-question and talk-time checks

94%

Feedback loop closure

15 of 16 raised issues reported back

340

Hours returned

from the hours-returned metric only

What we measured

Twelve AI-assisted interviews in Ustadi Demo Group, each scorecard drafted by the model and signed off by a named human. Panel talk-time and question style were checked as the interviews went.

What the numbers say

Panels agreed beyond chance (κ = 0.82), which means structured scoring landed the same judgment across interviewers. Bias flags fired on 8% of sessions — every one logged for a human to weigh, never applied by the model. On the voice side, 15 of 16 raised issues were acted on and reported back to the people who raised them.

What it means

Consistent panels, visible flags and closed loops are the pattern of a team whose tooling amplifies judgment instead of standing in for it. The human signs every scorecard, every approval, every closure — and 340 hours returned to that work.

Live aggregates — seeded cohort snapshot · Reviewed and published by Tara Njeri on 2026-10-03 · Aggregated and anonymized — no individual is identified.