Colorful Radiology quality program
Quality Improvement, Measured.
Every feature ships as a Plan-Do-Study-Act cycle with an outcome measure. Results are published either way, including no-effect and insufficient-data findings.
Public datasets load from the audited local exporter.
Program snapshot
Evidence before claims
Primary outcome
About the numbered points. Each marked point is an intervention — one deliberate change, made and then measured. They are numbered rather than described: the methods are our own engineering, but the results are published exactly as measured, including the cycles that changed nothing. A dated marker lets a reader judge whether the line moved when we said it would.
The markers are a selection, not the total. Only cycles expected to move THIS measure are annotated, because a chart carrying every one would be unreadable. The count in the snapshot above is every commit that stated a hypothesis and named a measure, taken from the commit history rather than any hand-kept list: every commit states a hypothesis and names the measure it targets, and the count includes each one that named a measure.
Hands-on reporting time
What this measures. The minutes a radiologist spends actively producing one report — dictating, correcting, and reviewing it — with idle time excluded. It is deliberately not wall-clock turnaround, which can look excellent simply because a study sat in a queue while nobody worked on it. Active time cannot flatter itself that way. Lower is better.
Which days count. Weekends are excluded, and so is any day with fewer than 100 exams — a median over a couple of dozen cases describes that morning rather than the practice. The rule reads only how many exams a day held, never how good the day looked, so it cannot quietly drop an inconvenient result.
Why it matters. Time spent wrestling with software is time not spent looking at images. Every minute returned here is a minute available for the next patient, or for a more careful second look at this one.
View data
| Date | Minutes | Rolling median (10) |
|---|---|---|
| Data accruing. | ||
Elapsed time
Minutes from one exam to the next
What this measures. Wall-clock minutes between the start of one exam and the start of the next, weekends and non-working days excluded. Unlike the hands-on chart above it counts everything in between, including interruptions, phone calls, and time not spent on this software at all.
Why it is published anyway. It is the honest picture of how long a case actually takes end to end, and a quality programme that only published its flattering measure would not be one. Read it beside the hands-on chart, never instead of it: a rise here with a flat line above means the radiologist was doing something else, not that reporting got slower.
| Date | Minutes | Exams |
|---|---|---|
| Data accruing. | ||
Case study 01
Voice Dictation Quality
What this measures. How often speech recognition has to be corrected by hand, counted as manual fixes per 100 spoken phrases. Using a rate rather than a raw count means the number reflects accuracy, not whether a day happened to be busy or quiet. Lower is better.
Why it matters. Radiology vocabulary is unforgiving: the difference between two similar-sounding words can be the difference between two diagnoses. Every correction is time spent fixing words instead of reading images, and a correction missed is worse than one made.
What this counts, and what it does not. A radiologist changes dictated text for two different reasons: the software misheard him, or he decided to write something else. Only the first is a defect. This chart plots the phrases he SAID AGAIN, which is what he does when the machine got it wrong; phrases he retyped are counted separately, because adding a sentence is authoring, not correction. An earlier version of this page combined them into one number roughly three times larger, which overstated the software's error rate by counting the radiologist's own writing against it. Lower is better.
How honest this chart is. It separates errors the software made from edits the radiologist chose to make, because conflating the two would flatter the result. It stays marked provisional until enough reports accumulate to support a claim.
View data
| Date | Exam | Corrections / 100 words | Rolling median (10) |
|---|---|---|---|
| Data accruing. | |||
Case study 02
Accuracy QI v2
Problem. Accuracy features applied to every utterance could not show whether an individual lever improved correction burden or only added latency.
Intervention. Deterministic control and treatment arms now stamp utterance and report records; low-confidence decodes can be rescored while latency is measured and capped.
Result. Each lever receives a predeclared verdict, and insufficient data is published rather than converted into a marketing claim.
The levers are lettered, not named. As with the numbered intervention markers, the methods are our own engineering; the verdicts are published exactly as measured, including the ones that changed nothing and the ones still short of data. What a reader can check is whether we tested before claiming, and whether we published the failures.
| Lever | Metric | Control A | Treatment B | Verdict |
|---|---|---|---|---|
Data accruing - verdicts publish automatically as both arms accumulate. | ||||
Active measures
Workflow reliability
View data
| Week | Needing correction (%) | n |
|---|---|---|
| Data accruing. | ||
View data
| Wait | Week | Mean seconds | n | Failures |
|---|---|---|---|---|
| Data accruing. | ||||
Methodology
ABR PQI framing, public by default
Plan, Do, Study, Act. Each intervention states a hypothesis, identifies an outcome and balancing measure, and waits for the post-change interval before claiming an effect.
PHI-free by design. The public exporter emits rounded aggregates, date-only observations, and curated intervention labels. Report text, audio, identifiers, and workstation details are excluded.
No-effect results are published too. Improving, worse, no-effect, and insufficient-data verdicts remain visible under the same rules.