Mira Analytics
Rater Oversight

Your primary endpoint, scored independently.

In CNS trials the primary endpoint is typically a rater-administered scale — MADRS, HAM-A, HAM-D. Mira’s model scores the same interview audio item by item, independently of the site rater, so the gap between the two becomes visible measurement error on the result the study is judged by.

Request a demo

Automated scale ratings

From endpoint capture to an independent score.

Raters administer assessments as they always have. Mira ingests the questionnaire audio locally, locates each scale item within the interview, and scores it with a model trained on expert-rated interviews — independently of the rater’s own scoring.

Where the model and the rater disagree, that gap is the oversight signal: measurement error on the endpoint, visible while the study is still running rather than at database lock.

One engine works with any interviewer-rated scale or questionnaire, and travels across languages, sites and accents.

Interview waveform flowing into a Mira engine and a bar chart of scale ratings.

Explainability

Every score carries its evidence.

Mira’s scoring traces back to specific moments in the interview. Reviewers see which parts of the dialogue drove the model’s judgement — which is what makes a review fast, consistent, and worth trusting.

Insomnia Severe MADRS · Visit 4
RaterHow have you been sleeping lately?
ParticipantNot terrible, but not great either.
RaterAnd on the bad nights?
ParticipantI’m awake until three or four almost every night, and when I do sleep I’m up again after an hour.
Mira model output Short, fragmented sleep The participant reports frequent nights with delayed sleep onset and repeated awakenings — a pattern typical for severe insomnia.
RaterHow do these nights affect your day?
ParticipantI can’t focus in meetings and I’m forgetting simple things. I feel wiped out most days.
Mira model output Daytime cognitive impact Reduced concentration and mistakes with routine tasks — evidence that the insomnia is functionally impairing.
RaterHow’s your mood and appetite otherwise?
ParticipantMy mood’s okay, and I still go for walks. Just wish I wasn’t so tired.

Rater analytics

From single predictions to rater performance.

Individual scores aggregate into rater behaviour: drift over the study, over- and under-scoring tendencies, and agreement on exclusion decisions. Viewed per assessment, or longitudinally as they evolve.

Every rater is weighed against the others and against the study, so retraining goes to the people who need it — while the study is still running.

Rater scores tracked over time to reveal drift and scoring tendencies.

Our results

Measured against central raters, in a live Phase 3.

The founders have a long track record of building internal AI tools for the pharma industry — including the oversight platform used at Definium Therapeutics (formerly MindMed) to help monitor key endpoints (HAM-A and MADRS) in their Phase 3 clinical programs.

  • 95.2% accuracy on central rater training.
  • 1.57 (± 1.39) point average difference vs. central raters in Phase 2b.
  • Deployed in ongoing Phase 3 to monitor HAM-A & MADRS.

“Using Large Language Models for Endpoint Oversight”, ISCTM 2025 — Distinguished Poster Award — poster · abstract

Scatter plot of Mira scores against central rater scores.

Score your own interviews.

We run Mira against your blinded audio and show you the item-level scores beside your raters’.

Request a demo