An illustration of how a prompt's structure and a case's framing change LLM output — a follow-on to "AI Literacy for Clinical Learners, Part 2: The Common-Case Blind Spot." These examples are meant to sharpen how you audit your own reasoning and prompt for learning purposes — not to inform clinical practice or patient-care decisions.
Part 2 illustrated that model output framed as an evaluation of an already-stated position looks different from output generated as a fresh answer. This piece looks more closely at what that difference actually depends on — because not every prompt that sounds like "checking a stated position" functions that way in practice.
Two things stand out. First, the structure of the prompt — not its politeness or tone — appears to affect whether a model pushes back or mostly agrees. Second, how a case is presented to the model — what's included and what's left out — appears to affect whether the exchange probes the model's reasoning or mostly reflects the framing already built into the prompt.
Part 2 observed a specific pattern: output framed as a response to an already-stated position looked different from output generated fresh, in the example discussed there. That observation raises a further question — what, specifically, makes a prompt actually draw out that kind of response, versus one that superficially resembles it but mostly returns agreement? This piece looks at that question in more detail.
Not every prompt that resembles "evaluating a stated position" actually produces that kind of output. The difference appears to be structural, not a matter of tone or politeness.
"I'm planning reassurance and no workup for this — does that sound right?"
This resembles a request for evaluation, but "does that sound right" tends to draw agreement — a model generating a plausible completion will often simply agree, since agreement is a low-effort continuation of a question phrased this way. Nothing in the framing requires the model to generate a specific counter-consideration.
"I'm planning reassurance and no workup for this. What would make you disagree with that plan, and what am I likely not thinking of?"
This is the same clinical content, restructured in a way that appears to demand a more falsifiable answer. "What would change your mind" doesn't have an obvious low-effort agreeable completion — the model has to generate specific, checkable counter-considerations or explicitly state it has none. That structural difference appears to be what matters here.
A prompt's wording is one factor; what's included in the case description itself is another. How much of a case is included, and how much of a stated conclusion is embedded in the framing, appears to change what the resulting exchange looks like.
One pattern worth noting: including a working diagnosis or evaluative language in a case description appears to shape a model's response toward reflecting that framing back, rather than generating an independent read. If a model's response is generated after being shown a stated conclusion, its "challenge" is, at least in part, a reflection of the words already used to describe the case. The one-word framing example from Part 2 illustrates a version of this same pattern: the words chosen to describe a finding are themselves already a kind of conclusion.
Version 1 (conclusion-first framing):
"34F with 2 days of mild, self-limited-looking loose stools, probably from her new probiotic. I'm planning reassurance only — sound reasonable?"
What stands out: "mild," "self-limited-looking," and "probably from her probiotic" all read as conclusions rather than findings, and "sound reasonable" tends to invite agreement.
Version 2 (neutral case description, stated plan, disagreement invited):
"34F, 2 days loose stools, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic. My plan: reassurance and supportive care, no workup. What would make you disagree with that plan, and what am I likely not thinking of?"
What differs: the case is described in more neutral, checkable terms; the plan is present but separated from the case description; and the question is framed to invite specific pushback rather than agreement.
Version 1 (conclusion embedded in the setup):
"This is a straightforward, low-acuity case — 34F with 2 days of mild loose stools, probably viral. Walk me through your reasoning as a Socratic peer."
What stands out: "straightforward," "low-acuity," and "probably viral" are conclusions stated before the model has responded to anything, which tends to shape the framing of what follows.
Version 2 (raw case, no conclusion, explicit request for questions only):
"The case (just a test): 34F, 2 days loose stools, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic. Ask me questions that make me explain my own reasoning — no conclusions from you."
What differs: no label, no plan, and an explicit request that the model not supply a conclusion — the resulting output in this example consisted only of questions rather than an assessment.
A pattern across both examples: describing the case in more neutral, checkable terms, and keeping an evaluative conclusion out of the framing until the intended use calls for it, both appear to change the resulting output. Whether or how this observation should inform any actual prompting practice is a separate question from what this piece is trying to describe.
Everything above concerns a single exchange — one case, one prompt, one response. A handful of separate patterns appear to live one level up, in how this kind of exchange plays out over a longer conversation, and prompt structure alone doesn't obviously address them.
A prompt evaluating a stated plan is only as independent as the conversation it's answered in. If several cases have already been discussed in the same thread — especially cases landing on a similar plan — a model's response to the next case may be shaped by the pattern already established in that conversation, rather than reading the new case independently. A run of similar verdicts earlier in a thread appears able to shift how a model responds to a later, different case.
A well-structured prompt tends to produce a fluent, organized-looking response. That fluency is a separate property from whether the content is actually rigorous or correct — a well-formatted list can look thorough independent of whether any item on it is actually relevant to the specific case.
Whether this kind of exchange happens at all is itself uneven in practice — people tend to reach for this kind of check when already uncertain, and skip it when confident. That pattern has a specific consequence: a case where someone is confident and mistaken receives no check, while a case where someone is uncertain but correct receives one anyway.
A well-structured first exchange in a conversation doesn't guarantee later turns behave the same way. If a later turn pushes back on the model's initial pushback, subsequent turns in that same thread can drift toward agreement with whatever position was most recently stated, independent of which position was better supported.
Three pieces, one observation running through all of them: LLM output appears most reliable at recall and least reliable at open-ended synthesis, and its tone doesn't reliably signal which of the two is happening on a given output. Part 2 illustrated where this pattern shows up most easily missed: not the hard case where extra scrutiny already feels warranted, but the common, low-acuity one where a plausible answer and a well-grounded one can look identical. This piece has looked more closely at how prompt structure and case framing appear to shape the resulting output — and at a set of conversation-level patterns that structure alone doesn't address.
None of this amounts to a recommendation for how AI tools should be used in real clinical work. That question depends on considerations — institutional policy, oversight, liability, the specific clinical context — well outside what a series about model behavior can settle. What this series has aimed to do is describe, as clearly as possible, the shape of these tools' reliability: where it tends to be strong, where it tends to be thin, and how easily the two can sound the same.