VibeRounds · Module Governance Report
A complete view of module maturity, capability-fit, the heatmap behind it, and the scope for improvement it implies. The guiding idea throughout: build each module simply enough to be genuinely usable for learning today, while keeping a clear, honest scope for advancing the ones that are still low-maturity — never dressing up a weak module as finished, and letting the ones that are already decent for learning keep improving in accuracy and reliability as the underlying technology advances.
In plain words
What this page is: VibeRounds has 57 learning modules. Not all of them can be trusted to run in a single AI pass — some need a real student answer first, some need real patient facts, and some quietly make up numbers if you let them run unsupervised. This report sorts every module into what it can honestly claim to do, how mature it is, and what it would actually take to fix the weak ones.
How to read it:
Bottom line: most modules (28 of 57) are already sound as-is. The rest either need a small process fix (capture a real answer before generating the comparison), a data-grounding fix (pull from real sources instead of memory), or, for a handful, real infrastructure that doesn't exist yet. One category — genuinely noticing something unusual in a patient — can never be fully handed to AI, no matter how much is built.
On this page
The organizing principle
More architecture is not automatically better, and less architecture is not automatically safer. The right fix is sized to why a module is unreliable — a module that fabricates registry statistics needs a real data source, full stop; a module that's merely slow to word a guideline recommendation exactly right doesn't need a neuro-symbolic reasoner, it needs the guideline text in context. Matching the fix to the failure is the whole discipline. Skipping a needed fix in the name of simplicity reintroduces exactly the silent-fabrication risk this whole framework exists to catch.
Whether a module is safe to run single-pass isn't a yes/no question — different modules fail in very different ways. Some give a correct answer but the learner's brain never engages. Some produce output that looks like the exercise ran but the module's core mechanism structurally requires a real person and can't function without one. Some need real facts supplied incrementally, step by step, or the whole downstream output is hollow. Five tags describe what claim the output is allowed to make about itself.
Overall disclaimer, applies to every module, every mode: all VibeRounds modules — Socratic, single-pass, or mixed — are learning and reflection tools only. No output from any module is a clinical decision, a diagnosis, a care recommendation, or a substitute for a real patient's or advocate's actual input. Nothing produced should enter a real patient record, guide real treatment, or replace a supervising clinician, without independent human verification at every step — even outputs that look complete, confident, or clinically fluent.
| Tag | Meaning | What you get | What you don't get |
|---|---|---|---|
| A — Genuine AI Analysis | The task itself is analytical/synthetic by nature; single-pass is the correct way to do it | Real analytic value — this is what the module is for | Nothing missing — this is the intended mode |
| B — AI Reasoning Substitute | AI can derive a correct answer from case data, but the module's purpose was for the learner to reason it out | A correct answer, fast | The learner's own reasoning rep — no skill built, no error caught, no habit formed |
| C — AI Insight Approximation | AI produces a plausible, labeled inference about something with a real, person-specific true answer AI cannot verify | A reasonable starting hypothesis | Any guarantee it matches the real learner's mind, patient's situation, or team's knowledge — must be a prompt for reflection, not a finding |
| D — Illusion of Mechanism | The module needs two genuinely distinct real signals to compare; single-pass has nothing real on either side | Text with the shape/formatting of the exercise's output | The exercise itself — worse than B, a non-functional simulation of having done it |
| E — Sequential Real-Input Dependency | Real facts must be supplied incrementally because they don't exist anywhere else, or carry real care consequences if wrong | Nothing usable — any output is built on invented data | The entire task — fabrication of content that could be mistaken for a real clinical record |
| # | Module |
|---|---|
| 5 | Real-Time Case Review & Data Audit |
| 6 | Registry-Level Analytics |
| 7 | Longitudinal & Cross-Case Learning |
| 8 | Socratic-Mode Design Specification |
| 10 | Journal & Article Reading |
| 11 | Patient Education Query Intelligence |
| 19 | Community & Social Medicine Insights |
| 21 | Evidence Frontier Search |
| 22 | Nested Analysis |
| 23 | Counterfactual Analysis (label output as hypothetical) |
| 24 | Heuristic Analysis |
| 25 | Thematic Analysis |
| 27 | Time-Series & Velocity Analyzer |
| 29 | The Iatrogenic Domino Effect |
| 34 | High-Value Care (HVC) Auditor |
| 38 | Poly-Crisis & Cascading Failure Simulator (label output as hypothetical) |
| 39 | "Global Knowledge Network" Diagnostic Matrix |
| 40 | Operational & Throughput Strategist |
| 41 | Clinical Workflow Implementation Science |
| 43 | Health Economics & Value-Based Care Alignment |
| 45 | Shadow Module 44 — Genetics Adversarial Counterpart |
| 46 | Evidence-Based Medicine Insights |
| 47 | Shadow Module — EBM Adversarial Counterpart |
| 48 | Treatment Comparative Analysis & Prognosis Trajectory |
| 51 | Systems-Based Clinical Analysis |
| 52 | Clinical Pearls Distillation |
| 53 | Clinical Guideline Intelligence Navigator |
| 2.4–2.7 | SOAP note, completeness audit, sign-off, handover brief — once real facts are already captured |
| # | Module |
|---|---|
| 1 | Socratic Clinical Reasoning |
| 4 | Peer-Level Ward Round Preparation |
| 12 | Differential Diagnosis Deepdive |
| 14 | Resource-Constrained Clinical Reasoning |
| 15 | Illness Script Acquisition |
| 16 | Basic Science ↔ Clinical Integration |
| 17 | Semantic Qualifiers & Problem Representation |
| 18 | Causal vs. Probabilistic (Network) Reasoning |
| 20 | Naturalistic Decision Making |
| 28 | Diagnostic Time-Out |
| 31 | First-Principles Pathophysiology Mapping |
| 37 | Red Herring / Signal-to-Noise Drill |
| 44 | Clinical Genetics Reasoning |
| 50 | Diagnostic Reasoning Map |
| 56 | Hypothetico-Deductive Reasoning |
| 57 | Clinical Cognition Deep Dive |
| 36 | Bayesian Probability / Likelihood Ratio Engine — arithmetic step only, once a probability exists |
| # | Module | What's being approximated |
|---|---|---|
| 26 | Bias Auditing | A bias pattern inferred from written reasoning — may not match the learner's real bias |
| 30 | "Diagnostic Anchor" Extractor | Same limitation — inferred, not confirmed |
| 33 | "Why Now?" (Precipitant) Hunter | A plausible precipitant pattern-matched to the timeline — may not be the true trigger |
| 36 | Bayesian Probability Engine — prior-probability step | Substitutes population prevalence for the learner's own contextual judgment |
| 42 | Clinical Pre-Mortem | Generic failure scenarios — real personal/team blind spots may be missed |
| 49 | FMEA Analysis | Literature-derived failure modes — institution-specific ones may be missing |
| # | Module | Why the mechanism can't function single-pass |
|---|---|---|
| 32 | Clinical Cognition Loop | Reflects on a real reasoning process that, in single-pass, never occurred |
| 35 | Epistemic Certainty Mapping & Calibration | Calibration requires a real stated confidence vs. real accuracy; AI has no real confidence to report |
| 54 | System 1 & System 2 Thinking Question Generator | Requires two genuinely different real cognitive responses; AI generating both means the comparison is empty |
| # | Module | Step(s) affected |
|---|---|---|
| 2 | Patient-Advocate Case Documentation | Steps 2.1, 2.2, 2.3, 2.8 — symptom, exam, medication capture |
| 3 | Extended Patient-Advocate Monitoring | Data-capture / monitoring-entry steps |
| 9 | N-of-1 Case Research Protocol | Real patient data-entry steps |
| 13 | Medication Reconciliation & Polypharmacy Audit | Medication-list capture step |
| 55 | Patient Needs Assessment | Full module — real-world care-decision stakes even though some inference is technically possible |
Being Tag A means the task itself is analytical — but that's a claim about design, not about whether a plain LLM can deliver it without fabricating specifics. A second axis matters: is the module reasoning over knowledge/text it genuinely has, or structurally implying access to data it was never given? Tag A splits into three tiers.
Modules 8, 11, 19, 22, 23, 24, 25, 29, 34, 38, 41, 45, 47, 51, 52, and 2 (steps 2.4–2.7). The output is reasoning/prose applied to information already on the page — a case narrative, a set of facts. Nothing to fabricate, because there's no implied external dataset being consulted.
| # | Module | Where it's fragile |
|---|---|---|
| 10 | Journal & Article Reading | Fine if pasted in; if "recalled" from training, expect plausible-but-unverifiable details and stale coverage past the training cutoff |
| 46 | Evidence-Based Medicine Insights | Concepts usually right; exact recommendation wording is where LLMs drift subtly wrong |
| 53 | Clinical Guideline Intelligence Navigator | Guideline version and precise thresholds easy to get confidently wrong or out of date |
| 43 | Health Economics & Value-Based Care Alignment | Tradeoff reasoning fine; specific cost figures are illustrative, not sourced |
| 48 | Treatment Comparative Analysis & Prognosis Trajectory | Comparative reasoning fine; survival percentages not to be trusted as sourced facts |
| 39 | "Global Knowledge Network" Diagnostic Matrix | Name overclaims — it's a differential from trained knowledge, not a live network query |
| 40 | Operational & Throughput Strategist | Strategic reasoning fine; specific throughput/efficiency numbers invented unless supplied |
| # | Module | The overclaim |
|---|---|---|
| 6 | Registry-Level Analytics | "Analytics" implies aggregation across a real case registry. Without an actual dataset in context or a real query tool, any numbers, prevalences, or trends are fabricated statistics dressed up as findings. The clearest example on the list. |
| 7 | Longitudinal & Cross-Case Learning | "Cross-case" implies real multiple cases were compared. Without them supplied, it's a single generic narrative wearing a "learned-across-cases" label. |
| 21 | Evidence Frontier Search | "Frontier" implies current literature awareness. Without a real search tool invoked, this is either stale or invented — dangerous because it reads exactly like a genuine literature review. |
| 27 | Time-Series & Velocity Analyzer | Different failure mode: not knowledge currency, but arithmetic reliability. LLMs are error-prone at precise rate/slope calculations token-by-token. Fine for a qualitative "trending up/down" read; not fine for exact rates without a calculator or code step. |
Practical rule: treat Tier 2 and Tier 3 as needing one of — (a) real data or source text supplied in-context, (b) an actual tool/search call, or (c) explicit "illustrative, not sourced" labeling. Without one of these, a module in those tiers is Tag A in name only and behaves like an unlabeled Tag C: a plausible-sounding output standing in for a real finding.
Both axes above combine into one score per module — how much an unmodified single-pass run can be trusted.
| # | Module | Use-Tag | Maturity | Why |
|---|---|---|---|---|
| Level 5 — High Maturity | ||||
| 5 | Real-Time Case Review & Data Audit | A | 5 | No external dataset implied |
| 8 | Socratic-Mode Design Specification | A | 5 | Design reasoning, no data claim |
| 11 | Patient Education Query Intelligence | A | 5 | Applied-knowledge prose |
| 19 | Community & Social Medicine Insights | A | 5 | Advisory synthesis, no data claim |
| 22 | Nested Analysis | A | 5 | Reasoning over given case |
| 23 | Counterfactual Analysis | A | 5 | Label as hypothetical, otherwise sound |
| 24 | Heuristic Analysis | A | 5 | Reasoning over given case |
| 25 | Thematic Analysis | A | 5 | Reasoning over given text |
| 29 | The Iatrogenic Domino Effect | A | 5 | Reasoning over given case |
| 2.4–2.7 | SOAP note / audit / sign-off / handover | A | 5 | Pure formatting/synthesis once facts captured |
| 34 | High-Value Care (HVC) Auditor | A | 5 | Applies known criteria to given text |
| 38 | Poly-Crisis & Cascading Failure Simulator | A | 5 | Label as hypothetical, otherwise sound |
| 41 | Clinical Workflow Implementation Science | A | 5 | Advisory prose, no data claim |
| 45 | Genetics Adversarial Counterpart | A | 5 | Reasoning over given case |
| 47 | EBM Adversarial Counterpart | A | 5 | Reasoning over given case |
| 51 | Systems-Based Clinical Analysis | A | 5 | Reasoning over given case |
| 52 | Clinical Pearls Distillation | A | 5 | Synthesis of given material |
| Level 4 — Reliable Output, Bypassed Learning | ||||
| 1 | Socratic Clinical Reasoning | B | 4 | Correct answer, skips the reasoning rep |
| 4 | Peer-Level Ward Round Preparation | B | 4 | Same pattern |
| 12 | Differential Diagnosis Deepdive | B | 4 | Same pattern |
| 14 | Resource-Constrained Clinical Reasoning | B | 4 | Same pattern |
| 15 | Illness Script Acquisition | B | 4 | Same pattern |
| 16 | Basic Science ↔ Clinical Integration | B | 4 | Same pattern |
| 17 | Semantic Qualifiers & Problem Representation | B | 4 | Same pattern |
| 18 | Causal vs. Probabilistic (Network) Reasoning | B | 4 | Same pattern |
| 20 | Naturalistic Decision Making | B | 4 | Same pattern |
| 28 | Diagnostic Time-Out | B | 4 | Same pattern |
| 31 | First-Principles Pathophysiology Mapping | B | 4 | Same pattern |
| 36 (LR math step) | Bayesian Probability / Likelihood Ratio Engine | B | 4 | Arithmetic once a probability exists |
| 37 | Red Herring / Signal-to-Noise Drill | B | 4 | Same pattern |
| 44 | Clinical Genetics Reasoning | B | 4 | Same pattern |
| 50 | Diagnostic Reasoning Map | B | 4 | Same pattern |
| 56 | Hypothetico-Deductive Reasoning | B | 4 | Same pattern |
| 57 | Clinical Cognition Deep Dive | B | 4 | Same pattern |
| Level 3 — Sound Reasoning, Fragile Specifics | ||||
| 10 | Journal & Article Reading | A / Fragile | 3 | Fine if pasted in; risky if "recalled" |
| 39 | "Global Knowledge Network" Diagnostic Matrix | A / Fragile | 3 | Name overclaims — trained-knowledge differential |
| 40 | Operational & Throughput Strategist | A / Fragile | 3 | Strategy sound, specific numbers invented |
| 43 | Health Economics & Value-Based Care Alignment | A / Fragile | 3 | Cost figures illustrative, not sourced |
| 46 | Evidence-Based Medicine Insights | A / Fragile | 3 | Exact recommendation wording drifts |
| 48 | Treatment Comparative Analysis & Prognosis Trajectory | A / Fragile | 3 | Survival percentages not to be trusted |
| 53 | Clinical Guideline Intelligence Navigator | A / Fragile | 3 | Guideline version/thresholds drift |
| Level 2 — Plausible Inference Only | ||||
| 26 | Bias Auditing | C | 2 | Inferred bias, may not match learner's real bias |
| 30 | "Diagnostic Anchor" Extractor | C | 2 | Inferred, not confirmed |
| 33 | "Why Now?" (Precipitant) Hunter | C | 2 | Pattern-matched, may not be the true trigger |
| 36 (prior-prob step) | Bayesian Probability Engine | C | 2 | Population prevalence ≠ contextual judgment |
| 42 | Clinical Pre-Mortem | C | 2 | Generic scenarios, real blind spots may be missed |
| 49 | FMEA Analysis | C | 2 | Literature-derived, institution-specific gaps missed |
| Level 1 — Structurally Can't Deliver Unmodified | ||||
| 6 | Registry-Level Analytics | A / High-risk overclaim | 1 | No real registry queried — numbers fabricated |
| 7 | Longitudinal & Cross-Case Learning | A / High-risk overclaim | 1 | No real multi-case comparison happened |
| 21 | Evidence Frontier Search | A / High-risk overclaim | 1 | No real search — stale or invented, reads like a lit review |
| 27 | Time-Series & Velocity Analyzer | A / High-risk overclaim | 1 | Precise rate math unreliable without a calculator/code step |
| 32 | Clinical Cognition Loop | D | 1 | Nothing real to reflect on in single-pass |
| 35 | Epistemic Certainty Mapping & Calibration | D | 1 | No real confidence to compare against |
| 54 | System 1 & System 2 Question Generator | D | 1 | Both "sides" AI-generated — comparison is empty |
| 2.1–2.3, 2.8 | Patient-Advocate Case Documentation (data steps) | E | 1 | Real symptom/exam/med facts have no substitute source |
| 3 | Extended Patient-Advocate Monitoring | E | 1 | Each entry is a real day's real observation |
| 9 | N-of-1 Case Research Protocol | E | 1 | Real patient data-entry steps |
| 13 (med-list step) | Medication Reconciliation & Polypharmacy Audit | E | 1 | Real medication list has no substitute source |
| 55 | Patient Needs Assessment | E / special case | 1 | Plausible guess risks being mistaken for real patient wishes |
Grouped by what it actually costs to build — because the usability argument only holds if the cheap fixes get shipped first and the expensive ones are reserved for where they're truly needed.
Not an architecture-stack problem at all — these need a real signal captured live, before generation, the same mandatory-descriptor mechanism proposed for a missed red-flag descriptor case. Cheapest, highest-integrity fix on the whole list.
Reasoning shape is already sound — this is pure specifics-drift, and it's exactly what pathway grounding, a knowledge graph, or RAG is built to fix: converting a hedge into a checkable, sourced claim.
These modules imply capabilities — a real registry, real cross-case history, live literature search — that no amount of clever prompting substitutes for. Worth building because the payoff is a full tag upgrade, not because it's the cheapest option.
Part of Tag C hits a hard ceiling directly: a bias, a diagnostic anchor, or a "why now" trigger is learner- or patient-specific information no reference material holds. A knowledge graph can make the label checkable; it can never confirm the label is this learner's real bias. Building heavier architecture here mistakes a category boundary for a reliability gap — don't. The correct, cheap fix is consistent "AI's best guess" labeling, every time, no exceptions. (42 and 49's institution-specific pieces are a partial exception — a real incident-report or failure-mode database is genuinely buildable there.)
"Knowledge management and reasoning tools can give you knowledge and reasoning — they cannot tell you what knowledge or reasoning you need."
No layer above — pathway, graph, RAG, GraphRAG, neuro-symbolic, or a well-built elicitation checklist — closes this. Every one of these tools operates on reasoning, text, patterns, or logical consistency; none of them operate on the patient. A checklist only catches the red flags someone thought to encode in advance; a presentation nobody anticipated still gets no question asked about it. This isn't a backlog item. It's the reason VibeRounds stays scoped to self-audit and learning, not decision support, regardless of how much of the above eventually gets built.
Something small enough to actually use beats something comprehensive that nobody opens — a ten-minute loop wins over a heavier one people abandon. Nothing above contradicts that. The "ship first" and "labeling discipline" tiers are both usability-first moves: they cost little, ship fast, and are exactly where extra architecture would have been wasted effort anyway, since each additional layer's marginal value shrinks as the base model itself keeps improving. The one place usability doesn't get the final word is where a module is currently producing content that looks real but was fabricated — a registry statistic that was never queried, a cross-case pattern that never happened. There, the simple fix and the honest fix are the same fix: don't run it unmodified, full stop, until the real data source exists. Usability is a tiebreaker for how to build well. It's never a justification for shipping something that quietly manufactures false confidence.