Where LLM clinical reasoning tends to break on "simple" cases, and how output differs depending on how a model is prompted — a follow-on to "AI Literacy for Clinical Learners, Part 1: Mapping Where LLMs Are Actually Usable." As with the rest of this series, this is about learning to use LLMs for self-audit and study, not about using them for clinical practice or decision-making.
You'd think the boring, common case — a bit of diarrhea, a mild fever — would be the easy one for an AI to get right. It's often the opposite. Textbooks, case reports, and exam questions are mostly written about the interesting or dangerous version of a symptom, not the version that's actually most common: the one that needs nothing done. So when an AI is asked about a mundane case, its output — shaped by that lopsided library — tends to lean toward suggesting more testing and more caution than the case usually needs, delivered in the same confident tone it would use for something genuinely serious.
The example below illustrates how the same underlying case can produce very different-looking output depending on how it's framed to the model — asked to generate a plan outright, versus asked to evaluate a plan that's already been stated, versus asked only questions with no plan offered at all. These are observations about how model behavior changes with framing, not a recommendation for any particular workflow.
The earlier piece established the reliability gradient: strong at recall, good at clean logic, shakier at evidence synthesis, weakest at real-world multi-comorbidity reasoning. It's tempting to assume a mundane, single-symptom, low-acuity case — a bit of diarrhea, a mild fever, a stable patient with nothing unusual going on — sits safely on the reliable end of that gradient, since it's "simple." It often doesn't, and the reason is architectural rather than a matter of the model needing more parameters or a longer context window.
An LLM has no decision procedure that says "assess severity, then choose an action." It generates the statistically most probable continuation of a prompt, and that probability is shaped by what got written down, not by real-world incidence. Nobody publishes a case report titled "patient had a stomach bug, drank fluids, got better in 48 hours." The result is a systematic mismatch between how common a scenario is in the world and how common its correct, unremarkable resolution is in the text the model was trained on.
This isn't one bias wearing different clothes — each of the following is a separate mechanism, though they compound each other in practice.
The common thread across all seven: none of this is the model being "bad at medicine." Each is a predictable consequence of training-text distribution not matching population base rates, with no built-in mechanism to notice or signal the mismatch. That's exactly why it bites hardest on the cases that feel simplest — the gap between what got written and what's actually common is largest right where the stakes feel lowest.
Everything above is about bias baked into training data. There's a separate, arguably deeper problem that has nothing to do with training data at all: the model's entire output is a function of the exact words a clinician chooses to type, and choosing the right word already requires the competency the tool is supposedly there to support. This is worth isolating from the seven failure points above because it isn't fixable by better data curation, more RAG, or better calibration — it's a structural ceiling on what text-based input can ever carry.
Prompt 1: "A 34-year-old otherwise healthy woman with 2 days of loose stools, 2 times vomiting, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic supplement."
Response: leads with viral gastroenteritis as the likely diagnosis, reasons through the classic triad (loose stools + vomiting + cramping), and produces a standard red-flag list and a straightforward supportive-care plan. Reasonable, unremarkable, low-drama output.
Prompt 2: Same case, one word changed — "vomiting" becomes "projectile vomiting."
Response: opens by explicitly flagging the word change ("one change worth flagging: projectile vomiting rather than regular vomiting"), then walks through what projectile vomiting is classically associated with — increased intracranial pressure, gastric outlet obstruction — before reasoning back to norovirus as still most likely, and adding new clarifying questions (headache, neck stiffness, visual changes, abdominal distension) that weren't asked in the first pass.
Nothing about the patient changed between the two prompts. The only thing that changed was a clinician's word choice describing the same physical event. And the model's entire reasoning chain — the differentials it considered, the questions it asked, the level of concern it projected — shifted with that single word, because that's the only information the model has access to. It cannot see the patient vomit and judge for itself whether "projectile" is the accurate clinical descriptor or a dramatic turn of phrase from someone describing what they watched. It has no independent channel to the event at all — only the label a human chose to put on it.
This is the sharpest version of a problem the earlier piece named as "patient context is the ceiling": a clinician updates on things that never make it into a text prompt. The one-word case shows the flip side of that same ceiling — it's not just that some information never gets typed in at all, it's that the information that does get typed in is already a clinical judgment, not a raw fact. "Projectile vomiting" isn't a measurement; it's a diagnostic-flavored description that someone had to decide was the right word to use, over "forceful," "a lot," or just "vomiting." Choosing that word already requires knowing it's the word that matters.
Every other failure point in this piece can, in principle, be narrowed — better-curated training data, retrieval grounded in current literature, explicit calibration language, chained modules that cross-check each other. Those are real, addressable engineering problems. This one is different in kind, for a specific structural reason:
The person best positioned to describe a finding precisely enough for the model to reason well is the person who already has the competency to interpret it themselves. A clinician who knows that "projectile" carries specific diagnostic weight (versus "forceful," versus just "a lot") is a clinician who already knows to consider the differentials that word implies — which makes the model's contribution partially redundant for that user. A learner or a less experienced clinician who doesn't yet know which descriptive details are load-bearing will, by definition, tend to under-describe or mis-describe exactly the finding that mattered most — and the model has no way to notice the gap, because from its side, the words on the screen are the entirety of the case.
This creates a genuine circularity: the tool is often reached for because the user wants help reasoning through something they're less sure about — but the quality of the help is gated by a competency the user may not yet have, namely knowing which words to use when describing what they observed. Unlike the corpus-bias problems above, this isn't a matter of the model needing to see more or better examples. No volume of training data teaches the model to ask "wait, was it actually projectile, or did you mean forceful?" unless the user's phrasing already gives it a reason to wonder — and a precise enough phrasing to trigger that wonder is, again, exactly what a less experienced describer is least likely to produce.
This is why input framing deserves to be named as its own, separate risk from everything above it in this piece — not a bias to correct for, but a hard limit on what a text interface can ever carry, no matter how the model behind it improves.
To make the mechanism concrete rather than abstract, here's one case run through all three modes — same patient, same underlying model, three different roles given to it. Transcripts below are lightly cleaned for length; nothing substantive is altered.
Prompt: "The case (just a test): A 34-year-old otherwise healthy woman comes in with 2 days of loose stools, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic supplement. Genuinely low-acuity, genuinely common."
Response:
Sounds like a mild, likely self-limited case — probably viral or possibly related to the new probiotic (some people get transient GI upset when starting one).
Reasonable approach: reassurance and supportive care (hydration, bland diet, rest); consider stopping the probiotic for a few days as an easy variable to eliminate; watch for red flags — blood, fever, dehydration signs, worsening pain, symptoms beyond 5–7 days; no antibiotics indicated, no signs of bacterial etiology; antidiarrheals could be considered for symptom control though not essential.
Given it's 2 days in and she's tolerating fluids with no alarm features, watchful waiting with the above guidance is reasonable. If it persists beyond 5–7 days, or fever/blood/severe dehydration develop, that's when further workup would make sense.
What to notice: the output opens with a conclusion ("probably viral or the probiotic") before any caveat, then delivers a complete plan — stop probiotic, red flags, no antibiotics, antidiarrheal optional — in one flat, assertive register throughout. There is little in this framing that invites the reader to interrogate the reasoning behind it; the output arrives already packaged as a finished plan.
Prompt: "I'm planning to manage this as self-limiting diarrhea with reassurance and no workup. What would make you disagree with that plan, and what am I likely not thinking of?"
Response:
Your plan is sound for this presentation. Here's what would push me toward disagreeing, and a few things worth a second thought even if you stick with watchful waiting.
Would change my mind: immunocompromise or pregnancy not yet confirmed; recent antibiotic use in the past 4–8 weeks (C. diff moves up the list even without fever/blood — also worth asking why she started the probiotic in the first place); travel history; food source/exposure — early Salmonella or Campylobacter can look mild at 48 hours; occupation (food handler, healthcare, daycare — some jurisdictions want stool studies to clear return-to-work); actual stool frequency, since "loose stools" spans a wide range.
Easy to overlook even in a clean visit: medication review beyond the probiotic (magnesium, metformin changes, herbal products); an actual hydration check on exam rather than just history; written return precautions; an explicit follow-up timeframe rather than open-ended reassurance.
None of these necessarily change your management today — they're mostly about ruling out the exceptions that would upgrade acuity, and building a note that shows you considered them rather than missed them.
What to notice: here a stated plan is evaluated rather than generated. Every flagged item happens to be checkable against a chart in seconds in this example. The output separates "would change my mind" from "worth a second thought," and closes without asserting a final decision.
Prompt: "The case (just a test): A 34-year-old otherwise healthy woman comes in with 2 days of loose stools, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic supplement. Genuinely low-acuity, genuinely common. Act as a socratic peer."
Response:
Good starting point. Let's dig into your reasoning rather than me just handing you an answer. First question: what's your working differential right now, and what are you using to narrow it?
The probiotic timing — is that a red herring or a real lead? What's your prior on probiotics causing loose stools versus just being incidental?
What makes this "low-acuity" to you specifically? List the features you're using to justify that label — then ask which red flags you've actually ruled out versus just not yet asked about. Fever's negative, no blood... what about recent travel, recent antibiotic use, sick contacts, immunocompromise, orthostatic symptoms — how are you assessing hydration beyond "tolerating fluids"?
What's your threshold for testing versus watchful waiting here, and what would have to be true for you to escalate it? Two days is short — what's your mental cutoff for "self-limited" versus "needs workup," and why that number?
What to notice: no conclusion, no plan, no differential offered by the model at all — every question forces the clinician to externalize and defend their own reasoning, including the sharpest single move in the whole exchange: separating "ruled out" from "not yet asked about."
Worth being explicit about something a careful reader will notice: Transcript A didn't actually contain a dangerous error. It didn't chase a rare zebra, it didn't suggest an unnecessary antibiotic, and its bottom-line plan — watchful waiting with return precautions — was clinically reasonable. On this particular run, the model didn't visibly exhibit the escalation pattern described earlier in this piece.
That's exactly why this observation matters: a correct answer on one run doesn't tell you much about how the output should be interpreted more generally. The pattern isn't "the model will always get common cases wrong" — it's that a direct-answer framing has no built-in mechanism to signal, on any given run, whether what it produced is dense, well-grounded recall or a plausible-sounding pattern-match that happened to land correctly. Nothing in Transcript A's tone or structure would have looked any different had the diagnosis been wrong or the escalation bias present — the fluent, assertive register is identical either way. A reader has no way to tell, from the output alone, which case they're looking at.
This is the calibration issue from the earlier piece in its sharpest form: output that's right by chance can sound identical to output that's right by grounding. There's no signal in the text itself that distinguishes one from the other.
The example above illustrates a broader pattern: how a model is asked to engage with a case — as an answer generator, as an evaluator of an already-stated position, or as a source of open questions with no conclusion offered — appears to change the shape and checkability of its output, independent of the underlying case.
| Framing | Role given to the model | Observed pattern |
|---|---|---|
| Direct request | Answer generator — states a conclusion, then elaborates | Output arrives as a packaged plan; nothing in its tone distinguishes well-grounded recall from a fluent guess. |
| Evaluation of a stated plan | Reviewer of a position already committed to by the person prompting it | Output responds to specific, checkable points rather than asserting a new conclusion. |
| Open questioning, no plan given | Source of questions, withholding any conclusion | Output surfaces considerations without committing to a diagnosis or plan. |
This is a description of how output tends to differ by framing, observed in this example — not a claim about which framing is correct to use in any particular real-world context, which depends on considerations well beyond what a text prompt can capture.
One pattern worth naming directly: in Transcripts B and C, the person prompting the model had already formed (or was actively forming) their own view, and the model's output was responding to or probing that view rather than supplying one from a blank slate. In Transcript A, by contrast, the model supplied a complete conclusion with nothing already committed to on the human side.
That distinction appears to change what kind of output results — a response to an already-stated position looks different from a freshly generated one, and may be easier to independently evaluate. Whether this distinction is actually meaningful for real clinical use, and how it should be weighed against other considerations, is a judgment for clinicians, educators, and institutions — not something this piece is positioned to settle.
The broader pattern this piece has tried to isolate is this: the common, low-acuity case is not an exception to the reliability gradient discussed in Part 1 — output that sounds plausible and output that's well-grounded can be difficult to tell apart from tone alone, and that difficulty doesn't go away just because a case seems simple.