AI Literacy for Clinical Learners series — Index · Part 3 of 4

How Prompt Structure Shapes Model Behavior

An illustration of how a prompt's structure and a case's framing change LLM output — a follow-on to "AI Literacy for Clinical Learners, Part 2: The Common-Case Blind Spot." These examples are meant to sharpen how you audit your own reasoning and prompt for learning purposes — not to inform clinical practice or patient-care decisions.

The first two pieces described a pattern: where LLM output tends to sit on a reliability gradient, and why the common, low-acuity case is a blind spot rather than a safe zone. Part 2 also observed that output framed as a response to an already-stated position looks different from output generated from a blank slate. This piece illustrates that observation further: what tends to separate a prompt that draws out genuine pushback from a model from one that mostly returns agreement, how case framing changes what a model's response looks like, and worked before/after examples. These are observations about model behavior under different prompt structures, offered for AI literacy — not instructions for how to use these tools in clinical practice.

In plain terms

Part 2 illustrated that model output framed as an evaluation of an already-stated position looks different from output generated as a fresh answer. This piece looks more closely at what that difference actually depends on — because not every prompt that sounds like "checking a stated position" functions that way in practice.

Two things stand out. First, the structure of the prompt — not its politeness or tone — appears to affect whether a model pushes back or mostly agrees. Second, how a case is presented to the model — what's included and what's left out — appears to affect whether the exchange probes the model's reasoning or mostly reflects the framing already built into the prompt.

Key points in this piece

Background

Part 2 observed a specific pattern: output framed as a response to an already-stated position looked different from output generated fresh, in the example discussed there. That observation raises a further question — what, specifically, makes a prompt actually draw out that kind of response, versus one that superficially resembles it but mostly returns agreement? This piece looks at that question in more detail.

How Prompt Structure Appears to Affect Output

Not every prompt that resembles "evaluating a stated position" actually produces that kind of output. The difference appears to be structural, not a matter of tone or politeness.

An example that tends to draw agreement

"I'm planning reassurance and no workup for this — does that sound right?"

This resembles a request for evaluation, but "does that sound right" tends to draw agreement — a model generating a plausible completion will often simply agree, since agreement is a low-effort continuation of a question phrased this way. Nothing in the framing requires the model to generate a specific counter-consideration.

An example that tends to draw more specific pushback

"I'm planning reassurance and no workup for this. What would make you disagree with that plan, and what am I likely not thinking of?"

This is the same clinical content, restructured in a way that appears to demand a more falsifiable answer. "What would change your mind" doesn't have an obvious low-effort agreeable completion — the model has to generate specific, checkable counter-considerations or explicitly state it has none. That structural difference appears to be what matters here.

Three properties that seem to distinguish this kind of prompt

  1. The plan appears before the question is asked. In this example, the plan is present in the prompt, fully formed, rather than something the model is asked to help build.
  2. The question asks for disagreement, not confirmation. "What would change your mind" and "what am I not considering" seem to structurally invite specific pushback; "does this look right" seems to structurally invite agreement.
  3. The question asks what's checkable, not what's true. In this example, a specific, falsifiable question tends to produce checkable items rather than a broad, harder-to-evaluate differential.
An observation worth noting: when the most likely low-effort response to a prompt is agreement, that may be a sign the prompt functions more as a request for validation than as a probe for disagreement.

How Case Presentation Appears to Affect Output

A prompt's wording is one factor; what's included in the case description itself is another. How much of a case is included, and how much of a stated conclusion is embedded in the framing, appears to change what the resulting exchange looks like.

One pattern worth noting: including a working diagnosis or evaluative language in a case description appears to shape a model's response toward reflecting that framing back, rather than generating an independent read. If a model's response is generated after being shown a stated conclusion, its "challenge" is, at least in part, a reflection of the words already used to describe the case. The one-word framing example from Part 2 illustrates a version of this same pattern: the words chosen to describe a finding are themselves already a kind of conclusion.

Observed in the case description

Observed to shift the output when present

An observation worth noting: if a prompt's wording alone makes a working conclusion easy to infer, the resulting model output tends to trend toward reflecting that conclusion rather than independently testing it.

Illustrative Example: Evaluating a Stated Plan

Version 1 (conclusion-first framing):

"34F with 2 days of mild, self-limited-looking loose stools, probably from her new probiotic. I'm planning reassurance only — sound reasonable?"

What stands out: "mild," "self-limited-looking," and "probably from her probiotic" all read as conclusions rather than findings, and "sound reasonable" tends to invite agreement.

Version 2 (neutral case description, stated plan, disagreement invited):

"34F, 2 days loose stools, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic. My plan: reassurance and supportive care, no workup. What would make you disagree with that plan, and what am I likely not thinking of?"

What differs: the case is described in more neutral, checkable terms; the plan is present but separated from the case description; and the question is framed to invite specific pushback rather than agreement.

Illustrative Example: Open Questioning With No Plan Stated

Version 1 (conclusion embedded in the setup):

"This is a straightforward, low-acuity case — 34F with 2 days of mild loose stools, probably viral. Walk me through your reasoning as a Socratic peer."

What stands out: "straightforward," "low-acuity," and "probably viral" are conclusions stated before the model has responded to anything, which tends to shape the framing of what follows.

Version 2 (raw case, no conclusion, explicit request for questions only):

"The case (just a test): 34F, 2 days loose stools, mild cramping, no blood, no fever, tolerating fluids, recently started a new probiotic. Ask me questions that make me explain my own reasoning — no conclusions from you."

What differs: no label, no plan, and an explicit request that the model not supply a conclusion — the resulting output in this example consisted only of questions rather than an assessment.

A pattern across both examples: describing the case in more neutral, checkable terms, and keeping an evaluative conclusion out of the framing until the intended use calls for it, both appear to change the resulting output. Whether or how this observation should inform any actual prompting practice is a separate question from what this piece is trying to describe.

Patterns That Prompt Structure Alone Doesn't Address

Everything above concerns a single exchange — one case, one prompt, one response. A handful of separate patterns appear to live one level up, in how this kind of exchange plays out over a longer conversation, and prompt structure alone doesn't obviously address them.

Recency contamination across a conversation

A prompt evaluating a stated plan is only as independent as the conversation it's answered in. If several cases have already been discussed in the same thread — especially cases landing on a similar plan — a model's response to the next case may be shaped by the pattern already established in that conversation, rather than reading the new case independently. A run of similar verdicts earlier in a thread appears able to shift how a model responds to a later, different case.

Fluency and apparent rigor

A well-structured prompt tends to produce a fluent, organized-looking response. That fluency is a separate property from whether the content is actually rigorous or correct — a well-formatted list can look thorough independent of whether any item on it is actually relevant to the specific case.

Which cases get this kind of scrutiny at all

Whether this kind of exchange happens at all is itself uneven in practice — people tend to reach for this kind of check when already uncertain, and skip it when confident. That pattern has a specific consequence: a case where someone is confident and mistaken receives no check, while a case where someone is uncertain but correct receives one anyway.

Drift across follow-up turns

A well-structured first exchange in a conversation doesn't guarantee later turns behave the same way. If a later turn pushes back on the model's initial pushback, subsequent turns in that same thread can drift toward agreement with whatever position was most recently stated, independent of which position was better supported.

A pattern across all four: none of them are addressed by a single well-structured prompt, because none of them live inside a single prompt — they concern the conversation's history, how consistently this kind of exchange gets applied, and what happens in the turns that follow the first one.

Closing

Three pieces, one observation running through all of them: LLM output appears most reliable at recall and least reliable at open-ended synthesis, and its tone doesn't reliably signal which of the two is happening on a given output. Part 2 illustrated where this pattern shows up most easily missed: not the hard case where extra scrutiny already feels warranted, but the common, low-acuity one where a plausible answer and a well-grounded one can look identical. This piece has looked more closely at how prompt structure and case framing appear to shape the resulting output — and at a set of conversation-level patterns that structure alone doesn't address.

None of this amounts to a recommendation for how AI tools should be used in real clinical work. That question depends on considerations — institutional policy, oversight, liability, the specific clinical context — well outside what a series about model behavior can settle. What this series has aimed to do is describe, as clearly as possible, the shape of these tools' reliability: where it tends to be strong, where it tends to be thin, and how easily the two can sound the same.