VibeRounds Prompt Modules · 2026

The Ultima Thule
Socratic Enrichment Suite

Two companion prompt modules — S3 and S4 — that walk a raw clinical narrative across a 100-level domain hierarchy before any diagnosis is discussed, so the blind spots surface before the differential does.

Modules S3 & S4 Domain source: Ultima Thule of Healthcare, Dr. Avinash Kumar Gupta 5 worked case examples
Clinical disclaimer & independent verification required Every output from these modules — domain selections, questions, tier labels, yield calls — is an educational reasoning-enrichment exercise, not a diagnosis, triage decision, or care plan. Nothing here substitutes for clinical judgment or institutional protocol, and no output should enter a patient record without independent review by a licensed clinician.
Naming the third system

What "System 3" is actually being asked to do

Kahneman's dual-process model, which the module cites as its theoretical basis, has two systems: System 1, fast and intuitive; System 2, slow and effortful. There is no System 3 in that framework.

"System 3" here is VibeRounds' own coinage, not a third mode of cognition the literature has identified. It names a specific engineering goal instead: force a taxonomy-guided, domain-exhaustive scan to run before System 2's differential-diagnosis reasoning gets to narrow the field. S4 brands itself "Forced System 3" for exactly this reason — it's the version built to make that forcing real rather than assumed.

That's a legitimate, well-motivated design goal even though the "3" borrows Kahneman's numbering rather than his evidence. The question worth answering isn't whether System 3 exists as a cognitive category — it doesn't need to, for the tool to be useful — it's whether the forcing mechanism does what it's built to do: make coverage real and auditable instead of assumed. That's what the rest of this page checks.

Domain source

The 100-level hierarchy both modules scan against

Both S3 and S4 walk the same raw case narrative across the same taxonomy: the Ultima Thule of Healthcare — 100 Levels of Clinical Thinking, by Dr. Avinash Kumar Gupta. It's the reference domain map that gives S3 its "100 domains" and S4 its "20 five-level bands" — the fixed grid a case is checked against before ranking or yield-sorting begins.

Because the hierarchy is fixed and external to any single case, it's what makes a "silent domain" a checkable claim rather than a vibe: a domain is either addressed by the case narrative, partially addressed, or untouched, and that verdict can be read against the same 100 levels every time — which is exactly the property S4 exploits by forcing all 20 bands into visible output.

Why this exercise matters

What's actually being tested

This isn't an exercise in whether a language model can produce plausible-sounding clinical questions — it can, easily. It's an exercise in whether a "coverage" claim survives being checked. S3 claims to scan 100 domains silently; that claim is exactly as trustworthy as the model's word for it, since the scan itself never appears in the output. S4 exists to close that gap: it turns an unverifiable internal step into a printed, band-by-band verdict that a reader can audit line by line.

The significance, then, is less about the specific 100-level taxonomy and more about a general pattern in these tools: forcing structure changes what a model is willing to produce. A silent scan tends to converge on the obvious differential; a forced sweep surfaces the quieter domains that a case narrative left untouched, whether or not they turn out to matter. Comparing S3 and S4 side by side on the same cases is a way of measuring that difference directly, rather than assuming it.

Practical payoff

What you actually get out of running this

Beyond the theory, the modules are built to change what happens in a real case review. A few concrete benefits:

Catches the quiet finding

Surfaces the domain a case narrative never touched — the one an obvious differential tends to crowd out — before it becomes a missed diagnosis instead of a flagged gap.

Makes coverage auditable

Turns "I scanned everything" from an unverifiable claim into a printed, band-by-band verdict someone else can actually check, line by line.

Fast when you need fast

S3's ranked top-10 gives a quick widen-the-aperture pass on a case you mostly trust, without paying for a full 20-band audit every time.

Thorough when you need thorough

S4's forced sweep is there when the stakes justify it — a full, nothing-dropped pass before you let a ranking or a differential narrow the field.

Teaches the habit, not just the answer

Walking the same 100-level hierarchy repeatedly builds the reflex of checking quieter domains even without the tool — the exercise trains a scanning habit, not just a one-off output.

Chains into existing workflows

Both modules hand off downstream to differential-diagnosis reasoning, so the benefit compounds rather than replacing the judgment a clinician already applies.

The two modules

One scans and ranks. The other sweeps and shows its work.

Both modules run the same raw case narrative against the same 100-level Ultima Thule domain hierarchy. They differ in what they're optimized for — and in what they're willing to trade away to get it.

Module S3

Socratic Enrichment Engine

Silent scan → ranked top 10

Scans all 100 levels internally, then hands back the ten highest-impact Socratic questions the case narrative left unasked — ranked, scored, and ready to act on.

  • Fast: one table, ten rows, an impact score on each
  • Good for a quick widen-the-aperture pass on a case you mostly trust
  • The domain scan itself is not shown — you see the ranking, not the sweep behind it
Best when you want a short, actionable list — not a full audit trail.
Module S4

Exhaustive Domain Sweep

Forced 20-band sweep → yield-sorted, nothing dropped

Removes the silent step. Every one of 20 five-level bands gets a visible verdict — Touched, Partial, or Silent — before anything is judged useful. Only then does it sort into high-yield vs. rough-work.

  • Thorough: 20–40+ rows is normal on a dense case
  • Every band appears in the final output, even the boring ones
  • Built specifically to check whether S3's "silent scan" claim is actually true
Best when you suspect a case's obvious differential is crowding out quieter domains.

Why S4 exists

S3's "silently scan 100 domains" step is unverifiable — a model can generate its favorite differential first and retrofit tier labels onto it. S4's fix is structural: print a verdict for every band before any ranking is allowed.

Module S3
Module S4
Touched / ranked & visible Partial Silent
Illustrative only — the actual band count is 20 for S4; S3 doesn't expose a band-by-band view at all, which is precisely the gap S4 was built to check.
How a run unfolds

Both modules share the same shape

Orient the model with a contract, hand over the de-identified case, let it work in visible stages, then debrief and hand off to a downstream reasoning module.

01 · Initiation
Set the contract, paste the case
02 · Sweep
Scan (S3) or full 20-band pass (S4)
03 · Generation
Case-specific questions, one per gap
04 · Sort
Rank (S3) or yield-sort (S4)
05 · Debrief
Reflect, then hand off downstream
Usability

Who each module actually suits

Reach for S3 when

You have a case narrative and want a fast, ranked list of the highest-value questions it left unasked — a widen-the-aperture pass, not a full audit.

Reach for S4 when

An S3 run felt too diagnosis-driven, or you specifically want to check whether "obvious" findings are crowding out quieter domains before you trust a ranking.

Chain them together

Run S4 first to build a verified-complete sweep, then feed that sweep into S3 so its ranking sits on an inspectable base rather than a silent one.

Worked examples

Five case runs, side by side

Each link below is a full module output on a real de-identified narrative — the actual tables, tier tags, and gap notes a run produces, not a mock-up.