← Back to blog
· 7 min read

AI Clinical Note Accuracy: How to Test It Before You Buy

By The Kitt Team

TL;DR: AI clinical note accuracy is not about whether the note reads well. Test four things: capture, context, treatment detail, and downstream workflow. The tool that produces the cleanest prose may still lose clinically important information.

Why a clean note is not an accurate note

The note reads beautifully. It is still wrong. Most trial evaluations default to readability because it is the easiest thing to assess. But sounding clinical and being accurate are different things.

Research shows this matters even for the best tools in the category. A 2025 randomised trial at UCLA of 238 physicians across roughly 72,000 encounters found that ambient AI scribes significantly improved burnout, task load and work exhaustion, but time savings were tool-dependent: one tool cut time spent in the note by 9.5 per cent, while another showed no significant reduction at all. Same category. Different results.

AI clinical note accuracy is the degree to which a generated note reflects what happened in the session, carries the client’s history forward, preserves clinical detail at the level you would have written it, and arrives intact in the record you actually work from.

The four dimensions

Testing AI clinical note accuracy means going beyond first impressions. Four dimensions determine whether a note is genuinely useful or just presentable.

DimensionWhat you are testingWhat a fail looks like
CaptureDid it hear the session correctlyWrong side, wrong joint, garbled test names, invented findings
ContextDoes it carry the client forwardEvery note reads like a first appointment
Treatment detailIs the clinical substance intactSets, reps, load, progression and reasoning are flattened to “exercises given”
WorkflowDoes the note land where you workYou copy and paste it into your practice management system by hand

Dimension 1: Capture

Test for side, joint, named orthopaedic tests, numbers, and anything the tool could plausibly invent. The failure mode that matters is confident invention, not a typo. A note that reads “positive Lachman’s on the right” when you said “left” is not minor.

Capture errors are easiest to catch immediately on review. They become much harder to catch a week later, when memory has faded and the note is filed.

Capture errors are the dimension most likely to affect whether your record supports your AHPRA record-keeping obligations. In clinical documentation for allied health, that means: the facts recorded must be the facts that occurred.

Dimension 2: Context

Run a second and third synthetic session for the same client and check whether the note references previous findings, goals, and progression. A tool with no memory of the prior visit produces unconnected snapshots rather than a clinical narrative.

Studies show that 40 to 80 per cent of medical information given in a consultation is forgotten immediately, and almost half of what is remembered is remembered incorrectly (Kessels, Journal of the Royal Society of Medicine, 2003). The record carries the client forward. Your memory does not.

A second-visit note that treats the client as a new presentation is not a progress note. It is a starting-over note. If you write progress notes by hand, you know exactly what the prior session contained. The tool needs to know it too.

Dimension 3: Treatment detail

Dictate a full exercise prescription with dosage, load, tempo, progression criteria, and a reason for each choice. Then check what survives in the note.

Allied health notes live or die on this detail. A note that records “home exercise program provided” has lost the part that makes it defensible next visit and useful for any clinician who sees the client after you.

SOAP notes for physiotherapy depend on precise treatment plans. Sets, reps, load and progression criteria are not decoration. They are the note.

Dimension 4: Downstream workflow

Test where the finished note goes and how many actions it takes to get there. Research shows physicians spend about two hours on electronic record and desk work for every hour of direct care, with only 27 per cent of the office day spent face to face (Sinsky and colleagues, Annals of Internal Medicine, 2016). That is category evidence from medicine, and it reflects the real cost of the work around the note.

A note you retype into your practice management system has not saved you the time you were sold. Count the clicks. That number is part of the evaluation.

Run the test yourself

This protocol takes about an hour and gives you a like-for-like comparison across every tool you trial.

  1. Write three synthetic cases on paper first: a new acute presentation, a review two weeks in, and a complex case with two comorbidities and a workers compensation angle. Write down what a good note must contain for each case before you touch the tool.
  2. Dictate all three into the tool exactly as you would in clinic, including the messy parts: interruptions, corrections, the client changing their story.
  3. Score each of the four dimensions out of five against what you wrote in step one, not against how the note reads.
  4. Run cases two and three a second time in the same client record and check whether context carried forward.
  5. Push a finished note through to your practice management system and count the clicks.

Keep the scores. Re-run the same three cases on any tool you trial and you have a direct comparison rather than an impression.

Where kitt fits

kitt drafts, you review and approve. Your clinical judgement is the last word. kitt Clinician generates structured, editable, Medicare-ready notes, letters and treatment plans from the session. Client Contextual Memory surfaces prior findings, progression and goals before each session, which is the context dimension. One-Click Treatment Plans and Exercise Prescriptions carry sets, reps, cues and progressions rather than flattening them, which is the treatment-detail dimension.

Notes push back to Cliniko and Nookal, which covers the workflow dimension. kitt Companion is included with every subscription: the plan goes out to your client, and their adherence and pain data come back to you between visits.

Key takeaways

  • AI clinical note accuracy is about capture, context, treatment detail and workflow, not readability.
  • The same category of tool can produce very different results: test, do not assume.
  • Your test cases should be written down before you touch the tool, so you score against clinical requirements rather than impressions.
  • The workflow dimension is real: count how many steps it takes to get the note into the system you actually use.

FAQ

How accurate are AI clinical notes? Accuracy varies significantly between tools, even tools described in the same category. A 2025 randomised trial found meaningful differences in time savings between tools marketed for the same purpose. Test the four dimensions against your own clinical cases before you decide.

What is the most common AI clinical note error? Two error types stand out: capture errors such as wrong side or garbled test names, which are easier to spot on review but harder to catch once the note is filed; and context loss, where a returning client is treated as a first visit, losing the thread of their history. Both need testing.

Do I still have to review an AI-generated note? Yes. An AI-generated note is a draft. You review it, correct it, and approve it before it becomes the record. This is how kitt works: kitt drafts, you decide. Your clinical judgement and your AHPRA record-keeping obligations do not transfer to the tool.

How long should it take to evaluate an AI note tool? The protocol takes about an hour. You score each dimension against what you wrote down in step one, before you touched the tool, not against how the note reads. That gives you a basis for direct comparison.

Try kitt free for one month, no credit card required.

Ready to stay a step ahead?

No integration required. No credit card.