KSA-ZHAW Digital Health Lab Symposium 2026

When an AI reads a clinical note, how stable is the answer?

That was the question behind the work Martin Murin PhD presented at the Digital Health Lab Symposium at Kantonsspital Aarau.

Hospitals hold enormous amounts of useful information in discharge summaries and other free text. LLMs can help turn this into structured data.

But there is a catch.

The output can look perfectly clean and structured, while still depending on choices made behind the scenes: 1. which model, 2. which prompt and 3. how the output is defined.

Using MIMIC-IV, a population-scale dataset with more than 330,000 hospital discharge summaries, we looked at what happens when we change these choices one at a time.

What stood out:

  1. Rewriting the prompt changed the primary admission category in roughly 1 in 8 notes.

  2. Changing the model changed it in close to 1 in 2.

  3. and another interesting finding: much of the disagreement was not about whether something was actually present. It was about the difference between “no” and “not documented.”

That may sound subtle, but clinically those are two very different things.

For a hospital, the practical takeaway is quite concrete.

If an LLM becomes part of a clinical data pipeline, a model update, prompt change, or schema adjustment can change the structured data, even though the underlying clinical notes have not changed.

So, these components need to be treated like any other production system: versioned, validated before deployment, and checked again when something changes.

Full paper: here

Download poster: here

Next
Next

DryLabz starts clinical feasibility project with KSB and HIH Aargau