One tool spots hidden heart disease from a routine ECG faster than any cardiologist alive. A different one, already running in NHS consulting rooms, dropped a single word and turned a normal test result into a multiple sclerosis diagnosis. Same technology category. Same year. Wildly different track records.
Medical AI produced two contrasting stories, one promising, one alarming. Researchers presenting at the European Society of Cardiology‘s annual congress in Munich unveiled what they’re calling a “superhuman” AI tool, capable of spotting hidden heart disease from a routine ECG in under two seconds. Days later, Healthwatch England, the statutory patient watchdog, warned that AI transcription tools already deployed across NHS consulting rooms are misrecording drug names and diagnoses — errors patients are catching, not clinicians. Meanwhile, at Macquarie University in Sydney, a Microsoft-built AI tutoring chatbot kept delivering 10% grade improvements, expanding well beyond its original pilot.
What’s Happening & Why It Matters
Two Seconds, 81% Accuracy: What the Heart Tool Does

The Munich breakthrough addresses a diagnostic gap. A standard electrocardiogram has recorded the heart’s electrical activity for a century, but it can’t detect structural disease — that requires an echocardiogram, an ultrasound scan patients often wait months to get. Dr Ahmed El-Medany, the BHF clinical research fellow at Imperial College London who led the analysis, described the new tool’s actual function: extracting diagnostic signals buried in routine ECG traces that the human eye can’t see.
The numbers are specific and independently reported. Trained on more than 1.6 million ECGs from Brazil and several million more from the US, the tool was validated against 67,000 patients, identifying up to 81% of people with heart failure and up to 90% of those with heart valve disease. El-Medany called it “superhuman” — not hyperbole in this specific, narrow context, since the tool is extracting information no human clinician can perceive from the same trace. The next challenge, he said, is building a handheld version cardiologists can use at the bedside.
Producing Opposite Outcomes in Consultation Rooms

Here’s where the story turns less reassuring. Healthwatch England documented specific, patient-discovered errors from AI scribes — ambient listening tools that transcribe doctor-patient consultations into medical records. In one case, a woman was told her scan showed “demyelination,” the nerve damage multiple sclerosis. The actual result read “null demyelination.” An AI scribe had dropped a single word that reversed the meaning. She caught the error herself, because she happened to work in the NHS and questioned the result.
That wasn’t an isolated incident. Other documented cases include a scribe swapping a prescribed drug for a different medication with a similar name, and another that dropped a consultant’s instruction for a repeat migraine prescription — an omission that could have left a patient without medication. Twenty-seven different AI scribes operate across English GP surgeries and hospitals. None of them is regulated as a medical device.
Two Tools, Varying Perceptions

The regulatory gap here isn’t accidental — it’s built into how the rules are written. The Medicines and Healthcare products Regulatory Agency published guidance in August clarifying the line: a system that only transcribes what was said isn’t classified as a medical device, while one that suggests a diagnosis or treatment is. That distinction creates an obvious incentive. A scribe marketed as a passive transcriber avoids the regulatory scrutiny a clinical-assistant tool would have to pass through — even though a transcription error that swaps one drug name for another produces identical patient harm to a diagnostic error.
Rachel Power of the Patients Association argued the technology needs better communication and partnership with patients, not blanket rejection. That’s a fair view given the cardiology results beside the transcription failures — this isn’t a story about AI being unsuitable for medicine. It’s about which specific applications have been rigorously validated, and which have been deployed into live patient care ahead of the oversight that should govern them.
A Quieter Success Story

Away from both the breakthrough and the warning, Macquarie University’s AI tutoring chatbot — built with Microsoft and known as Virtual Peer — continues expanding on the strength of results that have held up since its original pilot. Students who used the tool in the two weeks before an exam saw a 9.45% grade improvement in a controlled study of 1,400 psychology students, with the university extending the trial across multiple faculties. Head of AI Phil Laufenberg attributed the tool’s reliability to a deliberate constraint: Virtual Peer only answers from verified, grounded institutional data, with every response linked to a source document students can click and check themselves.
That “grounded data only” design is the specific safeguard the NHS scribes appear to lack. Virtual Peer isn’t generating information — it’s retrieving and citing it, with a built-in verification path for every claim. The AI scribes, by contrast, are transcribing live spoken conversation in real time, a harder technical problem with no equivalent citation trail once an error slips into the permanent medical record.
TF Summary: What’s Next
The heart disease detection tool is in the validation and handheld-device development stage, with no confirmed NHS rollout timeline yet. Healthwatch England’s warning has not yet produced a formal MHRA reclassification of AI scribes as medical devices, despite the watchdog’s concerns. Macquarie University’s Virtual Peer trial continues expanding toward its stated goal of covering hundreds of course units by 2026.
MY FORECAST: Expect the MHRA to face regulatory pressure to close the transcription-tool loophole within the next year, given how Healthwatch’s documented cases — a reversed MS diagnosis, a swapped medication — demonstrate patient harm occurring today. The vendor incentive to market scribes as “passive transcription” to dodge medical-device classification is harder to sustain once enough of the cases accumulate. The heart disease tool’s trajectory is more promising, because it went through peer-reviewed validation against real patient outcomes before any deployment claim was made — the exact sequence the NHS scribes skipped.
Related Stories
- Mayo Clinic Is Testing AI Across Diagnostics and Patient Care
- A Brain Implant Just Let a Paralysed Man Feed Himself for the First Time
- ChatGPT Health Arrives for All US Users

