Mining clinical notes with language models to surface medication harm
Every week in Australian hospitals, thousands of patients experience an unwanted reaction to a medicine they were given in good faith. Adverse drug events are estimated to contribute to around 230,000 hospital admissions nationally each year, and they remain a leading cause of preventable harm. The signals are usually there in the patient record, but they are not always easy to find with the tools most hospitals rely on.
Electronic health record systems in Australia capture vast amounts of clinical information, yet a surprising share of it lives in free text. Discharge summaries, progress notes, nursing entries, and outpatient letters are written in shorthand that varies from clinician to clinician. A single reaction to an antibiotic might appear as "rash on amoxicillin" in one note and "erythematous rash, possible hypersensitivity" in another, with no structured field to pull these together. Standard reporting to the Therapeutic Goods Administration captures only a fraction of this activity, leaving the rest buried in narrative.
A new generation of natural language processing tools is beginning to change that picture. By training algorithms to read clinical text the way a trained pharmacist or clinical reviewer would, research teams are starting to surface medication harm at a scale that manual chart review could never match. Brisbane Diamantina Health Partners sits at the heart of this work in Queensland, connecting hospitals, universities, and research institutes to test these tools in real-world settings and translate them into safer care.
Why coded fields miss most of the story
Hospital information systems in this country are designed primarily for billing and workflow. The fields that get coded reliably are those tied to funding under activity-based models, which means diagnosis codes, procedure codes, and a small set of allergy flags. Everything else ends up in the narrative portion of the record, where clinicians write quickly using abbreviations, local conventions, and personal shortcuts.
This is where automated language understanding earns its place. A patient in a Brisbane teaching hospital might have "metformin stopped, pt c/o diarrhoea" in a junior doctor's progress note, with the reason for stopping left to inference. Without natural language understanding, coded-field searches miss this entirely. The same goes for mentions of complementary products, dose changes in letters, and subtle temporal cues like "started two weeks ago".
Queensland Health data shows more than 80 per cent of documented adverse drug events live outside structured fields. The Australian context adds its own flavour: high rates of polypharmacy among older patients, frequent use of over-the-counter products from the chemist, and prescribing under the Pharmaceutical Benefits Scheme that is tightly governed but not always tightly documented.
How clinical language models read the record
The technical core of this work is a pipeline that turns messy clinical prose into structured findings. Tokenisation breaks a note into individual words and sub-words, named entity recognition identifies mentions of drugs, doses, symptoms, and patient states, and a negation module works out whether a clinician is describing something that happened or something that did not.
More advanced pipelines add temporal reasoning, so the system knows that "rash on day 5 of flucloxacillin" refers to an event five days into a course. Transformer-based models adapted for clinical text handle the heavy lifting. Australian teams have built systems that cope with local spellings, abbreviations such as "prn", "mane", and "nocte", and the bilingual cues that appear in communities with high migrant populations.
Validation is rigorous. Every flagged event is checked by a clinical pharmacist or senior clinician, and the system is tuned to balance sensitivity against false positives that can drown a ward team. The aim is not to replace human review but to triage the haystack so reviewers can spend time on the needles.
Queensland-led progress and national reach
Translation of these tools into routine care is uneven across the country, but Queensland has emerged as an early adopter. Work at the Princess Alexandra Hospital and the Royal Brisbane and Women's Hospital has paired NLP extraction with electronic medical record data to support local quality improvement. Outputs feed into ward-level dashboards and inform prescribing committees that review medication safety trends.
Nationally, the Therapeutic Goods Administration and the Australian Institute of Health and Welfare have signalled interest in text-mined signals alongside spontaneous reports, and the Australian Digital Health Agency continues to invest in infrastructure for large-scale record analysis. My Health Record, while primarily a patient-facing summary, also creates opportunities for shared learning across jurisdictions.
Partnership networks make the difference. Linking a data science team at a university with a hospital's electronic medical record team, a pharmacovigilance pharmacist, and an ethics and governance office takes more than a single grant. It takes the kind of sustained collaboration that metropolitan and regional Queensland are quietly getting better at.
From research to real-time clinical support
The longer-term goal is real-time decision support. Imagine a GP writing a letter in Medical Director or Best Practice and the system quietly flagging that a patient had a possible statin-related muscle reaction in a hospital note two years earlier. Or a pharmacist reviewing a new admission and seeing a warning that a prescribed anticoagulant interacts with a complementary product the patient is taking.
Complementary medicine is a rich area for text mining. Many patients do not record supplement use with their GP but mention it in hospital notes. A patient managing blood sugar with cinnamon, for example, may experience unexpected interactions with prescribed hypoglycaemics, and a well-tuned pipeline can flag this from a single sentence where the supplement is never formally coded.
Embedding these tools into the workflow of busy clinicians is the next challenge. The harder work is making sure alerts are useful, reach the right people, and do not add to the documentation burden that already weighs on clinical staff in public hospitals.
Governance, privacy, and the path ahead
Working with clinical text at scale raises questions no algorithm can answer. Patient consent, data sovereignty, and the safe handling of identifiable information sit at the centre of any serious deployment. In Australia, ethics committees, site-specific governance offices, and state data custodians each play a role, and approval timelines can stretch the work over years.
The national conversation around My Health Record has sharpened public awareness of how health data is stored and shared. Researchers need to be clear about what is being extracted, how it is de-identified, and where structured outputs live. Frameworks that support responsible research translation align these activities with the expectations of patients, clinicians, and regulators.
There is also the question of who benefits. Tools built for one hospital network may not transfer cleanly to another without local tuning, and rural and remote services can be left behind if deployment concentrates in metropolitan teaching hospitals. Building equity into the rollout is part of the work, not an afterthought.
If you work in pharmacovigilance, clinical informatics, or hospital medication safety, there are practical ways to get involved. Reach out to Brisbane Diamantina Health Partners to start a conversation about bringing text-based pharmacovigilance into the Australian health system at scale.