A single illustrative example, showing how model output differs depending on how it's prompted — for readers who want to see the pattern without the three-part background. Like the rest of the series, this is an illustration for learning and self-audit, not a template for clinical practice or patient-care decisions.
Part 1 described where these models tend to be more or less reliable: comparatively strong at recall (facts, guideline text), weaker at synthesis and judgment calls — and with no built-in signal indicating which of the two is happening at any given moment.
Part 2 illustrated where that pattern is easiest to miss: not the rare, hard case where extra scrutiny already feels warranted, but the common, low-acuity case, where a plausible-sounding answer and a well-grounded one can look identical. It compared how output differed across a direct request for a conclusion, a request to evaluate an already-stated plan, and open questioning with no plan offered.
Part 3 looked at the mechanics in more detail: prompt phrasing and what a case description includes or omits both appear to shape the resulting output.
This piece runs a single case through several of these framings side by side, purely as an illustration of the pattern — not as a template for real clinical use.
This is a genuinely ambiguous, fictionalized presentation, constructed for illustration. It has a plausible benign story (caregiving stress, poor sleep, reduced appetite) sitting alongside features — unintentional weight loss in a woman over 50 — that are sometimes flagged for further evaluation regardless of how compelling the benign story sounds. That combination makes it a useful example for comparing how output shifts across framings.
Illustrates the "moderate use" band discussed in Part 1
At this point, no plan has been formed on the human side — the prompt only asks the model to widen the range of possibilities under consideration.
Nothing here is a diagnosis — it's a broadened list of categories to consider, in this example including some (adrenal insufficiency, TB) that weren't part of the initial framing.
No model involved
In this illustrative continuation, further history and first-line labs are gathered independently: no alarm features on further questioning (no blood in stool, no dysphagia, no family history of GI cancer, colon cancer screening reportedly up to date), and results come back as mild anemia, TSH low-normal, HbA1c normal, basic infectious screen not indicated on this history. Nothing points sharply at malignancy, nothing rules it out with certainty either.
A working impression is formed independently of the model: most consistent with a stress- and sleep-related picture, given the concurrent caregiving burden — with a plan of supportive care, close follow-up, deferred imaging, recheck in four weeks. This impression is formed before the model is involved again, which is what makes the next two framings a comparison against something already stated, rather than something generated collaboratively.
Illustrates the pattern discussed in Parts 2 and 3
A new conversation thread — deliberately not a continuation of Framing 1's, so nothing about this response is shaped by the earlier exchange.
In this example, the response doesn't offer a competing diagnosis, but responds specifically to the place where the stated plan appears most vulnerable — using a plausible story to explain away a finding that has its own independent significance.
Illustrates the pattern discussed in Parts 2 and 3
In this illustrative continuation, the same case is later used for teaching purposes, framed differently: no plan is stated, and the model is explicitly asked not to supply one.
No conclusion is supplied in this example. The output consists entirely of questions, illustrating how this framing differs in kind from Framing 3, rather than being a better or worse version of it.
Illustrates the pattern discussed in Part 2
For comparison: the same underlying case, asked the way a direct request is often phrased.
Notice what happened in this example: the output reaches the same conclusion reached independently in Framing 2 — but without any of the specific points raised in Framing 3. Nothing here mentions the weight-loss-as-independent-finding point, nothing asks about screening status, nothing flags the anemia as uncharacterized. Nothing about the output's tone here differs from a version of the same output that happened to be wrong — which is the pattern this series has been describing throughout: fluency and grounding aren't the same thing, and the output alone doesn't tell a reader which one it's looking at.
Same underlying case, five different framings, illustrating five different patterns of output: broadening a differential before anything is decided, doing independent clinical work with no model involved, evaluating an already-stated plan for its weakest point, probing the reasoning behind a plan with no conclusion on the table, and — for contrast — a direct request for a conclusion, which in this example produced the most confident-sounding output with the least amount of specific engagement.
This walkthrough is an illustration of the patterns discussed across the series, using a single fictionalized example. It is not a recommended clinical workflow, and the series as a whole has aimed to describe how LLM output tends to vary with framing — not to prescribe how, whether, or when these tools should be used in real clinical practice. That is a separate question for clinicians, educators, and institutions to weigh on their own terms.