Vibe Rounds Cognitive Analytics Series · Companion Analysis

Weighting Ground-Truth Layers by Case Complexity and the Limits of System-Level Grounding

On when Promption, Provocation, knowledge graphs, institutional pathways, and appraised evidence each earn their weight — and the uncertainty none of them can close.

The base Vibe Rounds framework runs two cognitive modes — Promption (scaffolding) and Provocation (stress-testing) — over case narratives. A companion piece then anchors those modes to structured ground truth: a knowledge graph, an institutional pathway, and appraised evidence. This analysis asks the next question neither piece fully answers: given finite effort, which of these layers deserves the most weight, for which case, and what remains entirely outside all of their reach?

Architectural note: this is a learning-stack companion analysis for studying clinical reasoning frameworks — not a clinical decision tool, and not a substitute for primary literature, institutional protocol, or bedside judgment.

1. The final-stage problem: balancing methods, not ranking them

Each method in the stack targets a distinct failure mode rather than competing for the same job. The correct "final stage" is not a vote between methods but a map from output to the specific error each output is built to catch.

MethodWhat it's good forWhat it misses or risks
PromptionOrganizing messy narrative into structure, fastFluent but not fact-checked; can scaffold a wrong idea as smoothly as a right one
ProvocationCatching anchoring and premature closureCan manufacture doubt even where the original hypothesis was sound
Knowledge graphMakes facts checkable and falsifiableA correct fact can still be misapplied to an atypical or comorbid case
Institutional pathwayGets timing and sequencing rightCarries borrowed "per protocol" authority even when wrong
Appraised evidenceGrades the strength of the justification itselfThe most convincing hallucination surface if not backed by real retrieval
Governing principle: disagreement between layers is a usable audit signal, not noise to be smoothed over. A well-designed synthesis stage fails loudly — an explicit "no strong match" flag — rather than silently inheriting the base model's overconfidence.

2. Weighting by disease rarity: three tiers

The value of stacking ground-truth layers is not uniform. It rises with case rarity up to a point, then falls again once ground truth itself runs out.

Tier 1 · Common presentation, standard workup

The base model is already close to ground truth here. Adding a graph, pathway, and evidence layer mostly adds seams where independently-sourced components can quietly contradict each other, for a small accuracy gain. A random case series of common diseases adds little — other systems are already powerful for this tier.

Weight: Promption alone; light EBM search/screening if needed

Tier 2 · Rare-but-known presentation

This is the highest-leverage zone. The base model tends to over-fit an unusual case to the nearest common script — exactly the failure Provocation's anchor-extraction and the graph's differential field are built to catch. Ground truth exists here; it just has to be actively invoked. A small, well-chosen case series can meaningfully strengthen calibration, because it stress-tests whether the pipeline catches known-but-uncommon patterns reliably, not whether new pathophysiology can be discovered.

Pros
  • Directly tests Provocation + graph-based differential against real premature-closure cases
  • Builds a calibration ledger: confidence tracked against documented criteria across cases
  • High payoff relative to cost — grounding is possible, unlike Tier 3
Cons
  • Small samples risk overfitting the evaluation to the specific presentations chosen
  • Graph/pathway coverage is uneven even within "rare but known"
Weight: all three ground-truth layers highest here

Tier 3 · Ultra-rare or novel presentation

Gains shrink again, but for the opposite reason from Tier 1 — not because the base model is already good, but because there is no ground truth to anchor to. A structure-first approach can force a sparse or absent graph entry into the nearest available match, producing confident wrongness that is harder to detect because it wears the appearance of verified rigor.

Sharpest risk: evidence-appraisal-shaped language is the most convincing hallucination surface in the entire stack — it is only safe when backed by real, traceable retrieval, and retrieval is thinnest exactly where it matters most.
Weight: downweight graph/pathway confidence; upweight the coverage/confidence flag itself as the primary deliverable

3. Complexity axis, restated as scenario × task

The EBM Prompt Library adds a second axis: task type within evidence appraisal, each with its own verification burden. Crossed against the rarity tiers above, weighting becomes sharper.

ScenarioDominant cognitive toolsDominant EBM tasksVerification burden
A · CommonPromption onlyPICO framing, search string, abstract screeningLow — mechanical, easy to spot-check
B · Complex / rare-but-knownProvocation + graph + pathwayMethodology extraction, ARR/RRR/NNT arithmetic, stress-test of conclusionHigh — appraising primary literature against an atypical case, not a guideline default
C · Complex + ultra-rare/novelCoverage-flag over confident synthesisMethodology extraction and adversarial stress-test at maximum intensityMaximal — every number and citation substitutes for missing structured ground truth
The One Hard Rule, generalized: the LLM drafts, the human verifies every number and citation against the source — every time, not just when something looks off. As complexity rises from A to C, the proportion of the workflow that must be human-verified rises with it, because the tools' fluency stays constant while the ground truth backing that fluency shrinks.

4. The uncertainty that none of the layers cover

Every layer examined — Promption, Provocation, the knowledge graph, the institutional pathway, appraised evidence, even the EBM prompt tasks — operates on a documented representation of the case. Each trades individual specificity for checkability: a graph can only be a graph because it generalizes across patients; a pathway is a pathway because it is population-level; a PICO population is a category, never this one patient.

This is not a coverage gap a fourth layer could close by being more thorough. It is a category mismatch. The richest source of uncertainty — how sick the patient looks walking through the door versus what the numbers say, family and functional context, the trajectory of the last hour, what was already tried at bedside and failed — is exactly the information that never survives compression into a structured artifact.

What the layers can reduce
  • Uncertainty about whether documented facts are real and correctly applied
  • Uncertainty about timing and sequencing against a known protocol
  • Uncertainty about how strong the cited justification actually is
What only the human orchestrator can reduce
  • Whether the documented case is a faithful compression of the real patient
  • Fidelity of the narrative capture itself — deidentification, completeness, absence of leading framing
  • The gap between the note and the bedside, which no layer ever inspects
A perfect graph match on a poorly-captured case is confidently wrong for a reason no downstream layer will ever surface — the layer only ever sees the text, never the gap between the text and the bedside.

This adds a fourth axis to the diagnostic / sequencing / contextual correctness split proposed for the ground-truth layers: input curation correctness — owned entirely by the human orchestrator, and structurally unfixable by any layer downstream of the documentation step.

5. Conclusion: what the current project actually is, and what remains ahead

5.1 What the present toolset genuinely achieves

Taken together, the main article's Promption/Provocation architecture and the ground-truth layers analysis constitute a real, working demonstration at the reasoning-hygiene level: they show, on one worked case, that scaffolding messy narrative and then adversarially stress-testing it surfaces a genuine diagnostic discrepancy (the HbA1c/DKA-severity mismatch) that a single-pass answer would likely have missed. That is a legitimate and non-trivial contribution — it operationalizes System 2 friction in a way that is teachable and repeatable.

5.2 What it has not yet done

The current status is a single-case proof of concept, not a validated pipeline. It has not yet: run a matched case series across rarity tiers to test the accuracy/detectability/calibration metrics proposed in the ground-truth layers piece; integrated the knowledge graph, pathway, and evidence layers into one live pipeline rather than describing them as three separate reference resources; or connected the EBM prompt library's verification discipline into the same case-walkthrough interface used for Promption/Provocation. Each piece currently exists as a strong standalone artifact, not yet as one orchestrated system.

5.3 Against the full possible plan ahead

Relative to the complete plan implied across all three source documents — three ground-truth layers, three rarity tiers with differential weighting, an EBM verification discipline layered on top, and an explicit coverage/confidence output at every stage — the present project sits closer to the architecture and demonstration stage than the validation stage. The proposed empirical study (matched common/rare cases, measured for diagnostic accuracy, failure detectability, and learner calibration) has not been run. The input-curation problem identified in Section 4 has not been addressed by any tooling at all — it remains, correctly, an unassigned human responsibility. In proportion terms: the cognitive-mode architecture is largely built; the ground-truth grounding is designed but not yet wired in as a live system; the tiered weighting logic and cross-task verification discipline are, at this point, analytical recommendations rather than implemented features.

6. Would a NotebookLM-style import prototype this plan?

Importing a knowledge graph export, the acute pancreatitis pathway, relevant guidelines, meta-analyses, and an EBM-prompt-derived appraised-evidence summary into a NotebookLM-style tool, then analyzing a case against that source set, would replicate a meaningful slice of this plan — but not the whole of it, and the gap is informative.

What it would replicate well
  • Grounding answers in checkable source material rather than free-floating model assertion — the core purpose of all three ground-truth layers
  • Citation-anchored responses, which directly satisfies the EBM library's demand for quotable, locatable claims
  • A crude version of cross-layer disagreement, since a notebook-style tool can be asked to compare what the graph, the pathway, and the evidence summary each imply and flag where they diverge
Where it would fall short
  • No native Promption/Provocation pipeline sequencing — a notebook tool answers questions against sources, it does not natively enact a scaffold-then-stress-test cognitive workflow
  • No built-in tiering logic — it will not automatically know that a rare-but-known case should weight the graph and pathway higher, or that an ultra-rare case should downweight its own confidence; a human still has to impose that judgment turn by turn
  • No coverage/confidence flag as a first-class output — it can be prompted to produce one, but nothing forces it to fail loudly rather than fluently answering from thin source material
  • Cannot address input curation correctness — it grounds against the sources it is given, with no way to assess whether the case narrative fed into it was itself a faithful capture of the real patient
Net assessment: such a setup would function as a capable manual prototype of the grounding layer specifically — closer to Section 1 and 2 of the ground-truth layers analysis than to the full plan. It would still depend entirely on the human orchestrator to sequence Promption before Provocation, to decide how much weight the sources deserve given the case's rarity tier, to apply the EBM verification discipline number-by-number, and to judge whether the imported case narrative itself was trustworthy. It is a strong grounding substrate, not a replacement for the orchestration layer this whole analysis keeps returning to.