VibeRounds · Module Governance Report

Right-Sized Rigor: Where the Stack Earns Its Keep, and Where It Doesn't

A complete view of module maturity, capability-fit, the heatmap behind it, and the scope for improvement it implies. The guiding idea throughout: build each module simply enough to be genuinely usable for learning today, while keeping a clear, honest scope for advancing the ones that are still low-maturity — never dressing up a weak module as finished, and letting the ones that are already decent for learning keep improving in accuracy and reliability as the underlying technology advances.

In plain words

What this page is: VibeRounds has 57 learning modules. Not all of them can be trusted to run in a single AI pass — some need a real student answer first, some need real patient facts, and some quietly make up numbers if you let them run unsupervised. This report sorts every module into what it can honestly claim to do, how mature it is, and what it would actually take to fix the weak ones.

How to read it:

Bottom line: most modules (28 of 57) are already sound as-is. The rest either need a small process fix (capture a real answer before generating the comparison), a data-grounding fix (pull from real sources instead of memory), or, for a handful, real infrastructure that doesn't exist yet. One category — genuinely noticing something unusual in a patient — can never be fully handed to AI, no matter how much is built.

On this page

  1. The five-tag use-classification framework
  2. Full module classification, all five tags
  3. Capability-fit within Tag A
  4. Module maturity heatmap, full table
  5. Scope for improvement, grouped by build cost
  6. Why usability still wins the tie

The organizing principle

More architecture is not automatically better, and less architecture is not automatically safer. The right fix is sized to why a module is unreliable — a module that fabricates registry statistics needs a real data source, full stop; a module that's merely slow to word a guideline recommendation exactly right doesn't need a neuro-symbolic reasoner, it needs the guideline text in context. Matching the fix to the failure is the whole discipline. Skipping a needed fix in the name of simplicity reintroduces exactly the silent-fabrication risk this whole framework exists to catch.

1 · The five-tag use-classification framework

Whether a module is safe to run single-pass isn't a yes/no question — different modules fail in very different ways. Some give a correct answer but the learner's brain never engages. Some produce output that looks like the exercise ran but the module's core mechanism structurally requires a real person and can't function without one. Some need real facts supplied incrementally, step by step, or the whole downstream output is hollow. Five tags describe what claim the output is allowed to make about itself.

Overall disclaimer, applies to every module, every mode: all VibeRounds modules — Socratic, single-pass, or mixed — are learning and reflection tools only. No output from any module is a clinical decision, a diagnosis, a care recommendation, or a substitute for a real patient's or advocate's actual input. Nothing produced should enter a real patient record, guide real treatment, or replace a supervising clinician, without independent human verification at every step — even outputs that look complete, confident, or clinically fluent.

TagMeaningWhat you getWhat you don't get
A — Genuine AI AnalysisThe task itself is analytical/synthetic by nature; single-pass is the correct way to do itReal analytic value — this is what the module is forNothing missing — this is the intended mode
B — AI Reasoning SubstituteAI can derive a correct answer from case data, but the module's purpose was for the learner to reason it outA correct answer, fastThe learner's own reasoning rep — no skill built, no error caught, no habit formed
C — AI Insight ApproximationAI produces a plausible, labeled inference about something with a real, person-specific true answer AI cannot verifyA reasonable starting hypothesisAny guarantee it matches the real learner's mind, patient's situation, or team's knowledge — must be a prompt for reflection, not a finding
D — Illusion of MechanismThe module needs two genuinely distinct real signals to compare; single-pass has nothing real on either sideText with the shape/formatting of the exercise's outputThe exercise itself — worse than B, a non-functional simulation of having done it
E — Sequential Real-Input DependencyReal facts must be supplied incrementally because they don't exist anywhere else, or carry real care consequences if wrongNothing usable — any output is built on invented dataThe entire task — fabrication of content that could be mistaken for a real clinical record

2 · Full module classification

Tag A — Genuine AI Analysis (28 modules; single-pass is correct)

#Module
5Real-Time Case Review & Data Audit
6Registry-Level Analytics
7Longitudinal & Cross-Case Learning
8Socratic-Mode Design Specification
10Journal & Article Reading
11Patient Education Query Intelligence
19Community & Social Medicine Insights
21Evidence Frontier Search
22Nested Analysis
23Counterfactual Analysis (label output as hypothetical)
24Heuristic Analysis
25Thematic Analysis
27Time-Series & Velocity Analyzer
29The Iatrogenic Domino Effect
34High-Value Care (HVC) Auditor
38Poly-Crisis & Cascading Failure Simulator (label output as hypothetical)
39"Global Knowledge Network" Diagnostic Matrix
40Operational & Throughput Strategist
41Clinical Workflow Implementation Science
43Health Economics & Value-Based Care Alignment
45Shadow Module 44 — Genetics Adversarial Counterpart
46Evidence-Based Medicine Insights
47Shadow Module — EBM Adversarial Counterpart
48Treatment Comparative Analysis & Prognosis Trajectory
51Systems-Based Clinical Analysis
52Clinical Pearls Distillation
53Clinical Guideline Intelligence Navigator
2.4–2.7SOAP note, completeness audit, sign-off, handover brief — once real facts are already captured

Tag B — AI Reasoning Substitute (17 modules; correct answer, learner's brain bypassed)

#Module
1Socratic Clinical Reasoning
4Peer-Level Ward Round Preparation
12Differential Diagnosis Deepdive
14Resource-Constrained Clinical Reasoning
15Illness Script Acquisition
16Basic Science ↔ Clinical Integration
17Semantic Qualifiers & Problem Representation
18Causal vs. Probabilistic (Network) Reasoning
20Naturalistic Decision Making
28Diagnostic Time-Out
31First-Principles Pathophysiology Mapping
37Red Herring / Signal-to-Noise Drill
44Clinical Genetics Reasoning
50Diagnostic Reasoning Map
56Hypothetico-Deductive Reasoning
57Clinical Cognition Deep Dive
36Bayesian Probability / Likelihood Ratio Engine — arithmetic step only, once a probability exists

Tag C — AI Insight Approximation (6; plausible but unverifiable — flag as hypothesis)

#ModuleWhat's being approximated
26Bias AuditingA bias pattern inferred from written reasoning — may not match the learner's real bias
30"Diagnostic Anchor" ExtractorSame limitation — inferred, not confirmed
33"Why Now?" (Precipitant) HunterA plausible precipitant pattern-matched to the timeline — may not be the true trigger
36Bayesian Probability Engine — prior-probability stepSubstitutes population prevalence for the learner's own contextual judgment
42Clinical Pre-MortemGeneric failure scenarios — real personal/team blind spots may be missed
49FMEA AnalysisLiterature-derived failure modes — institution-specific ones may be missing

Tag D — Illusion of Mechanism (3; output looks complete, exercise didn't happen)

#ModuleWhy the mechanism can't function single-pass
32Clinical Cognition LoopReflects on a real reasoning process that, in single-pass, never occurred
35Epistemic Certainty Mapping & CalibrationCalibration requires a real stated confidence vs. real accuracy; AI has no real confidence to report
54System 1 & System 2 Thinking Question GeneratorRequires two genuinely different real cognitive responses; AI generating both means the comparison is empty

Tag E — Sequential Real-Input Dependency (5; single-pass fabricates content with no real basis)

#ModuleStep(s) affected
2Patient-Advocate Case DocumentationSteps 2.1, 2.2, 2.3, 2.8 — symptom, exam, medication capture
3Extended Patient-Advocate MonitoringData-capture / monitoring-entry steps
9N-of-1 Case Research ProtocolReal patient data-entry steps
13Medication Reconciliation & Polypharmacy AuditMedication-list capture step
55Patient Needs AssessmentFull module — real-world care-decision stakes even though some inference is technically possible

3 · Capability-fit within Tag A

Being Tag A means the task itself is analytical — but that's a claim about design, not about whether a plain LLM can deliver it without fabricating specifics. A second axis matters: is the module reasoning over knowledge/text it genuinely has, or structurally implying access to data it was never given? Tag A splits into three tiers.

Tier 1 — Solid: real text/knowledge reasoning, low fabrication risk

Modules 8, 11, 19, 22, 23, 24, 25, 29, 34, 38, 41, 45, 47, 51, 52, and 2 (steps 2.4–2.7). The output is reasoning/prose applied to information already on the page — a case narrative, a set of facts. Nothing to fabricate, because there's no implied external dataset being consulted.

Tier 2 — Fragile: the reasoning shape is legitimate, precise specifics are where it breaks

#ModuleWhere it's fragile
10Journal & Article ReadingFine if pasted in; if "recalled" from training, expect plausible-but-unverifiable details and stale coverage past the training cutoff
46Evidence-Based Medicine InsightsConcepts usually right; exact recommendation wording is where LLMs drift subtly wrong
53Clinical Guideline Intelligence NavigatorGuideline version and precise thresholds easy to get confidently wrong or out of date
43Health Economics & Value-Based Care AlignmentTradeoff reasoning fine; specific cost figures are illustrative, not sourced
48Treatment Comparative Analysis & Prognosis TrajectoryComparative reasoning fine; survival percentages not to be trusted as sourced facts
39"Global Knowledge Network" Diagnostic MatrixName overclaims — it's a differential from trained knowledge, not a live network query
40Operational & Throughput StrategistStrategic reasoning fine; specific throughput/efficiency numbers invented unless supplied

Tier 3 — High-risk overclaim: the premise requires a capability a plain LLM structurally doesn't have

#ModuleThe overclaim
6Registry-Level Analytics"Analytics" implies aggregation across a real case registry. Without an actual dataset in context or a real query tool, any numbers, prevalences, or trends are fabricated statistics dressed up as findings. The clearest example on the list.
7Longitudinal & Cross-Case Learning"Cross-case" implies real multiple cases were compared. Without them supplied, it's a single generic narrative wearing a "learned-across-cases" label.
21Evidence Frontier Search"Frontier" implies current literature awareness. Without a real search tool invoked, this is either stale or invented — dangerous because it reads exactly like a genuine literature review.
27Time-Series & Velocity AnalyzerDifferent failure mode: not knowledge currency, but arithmetic reliability. LLMs are error-prone at precise rate/slope calculations token-by-token. Fine for a qualitative "trending up/down" read; not fine for exact rates without a calculator or code step.

Practical rule: treat Tier 2 and Tier 3 as needing one of — (a) real data or source text supplied in-context, (b) an actual tool/search call, or (c) explicit "illustrative, not sourced" labeling. Without one of these, a module in those tiers is Tag A in name only and behaves like an unlabeled Tag C: a plausible-sounding output standing in for a real finding.

4 · Module maturity heatmap, full table

Both axes above combine into one score per module — how much an unmodified single-pass run can be trusted.

5 — High maturity: reason over given/known material, trustworthy as-is
4 — Reliable output, but bypasses the learner's own reasoning
3 — Reasoning sound, precise specifics need verification
2 — Plausible inference only, must be labeled a hypothesis
1 — Structurally can't deliver unmodified
#ModuleUse-TagMaturityWhy
Level 5 — High Maturity
5Real-Time Case Review & Data AuditA5No external dataset implied
8Socratic-Mode Design SpecificationA5Design reasoning, no data claim
11Patient Education Query IntelligenceA5Applied-knowledge prose
19Community & Social Medicine InsightsA5Advisory synthesis, no data claim
22Nested AnalysisA5Reasoning over given case
23Counterfactual AnalysisA5Label as hypothetical, otherwise sound
24Heuristic AnalysisA5Reasoning over given case
25Thematic AnalysisA5Reasoning over given text
29The Iatrogenic Domino EffectA5Reasoning over given case
2.4–2.7SOAP note / audit / sign-off / handoverA5Pure formatting/synthesis once facts captured
34High-Value Care (HVC) AuditorA5Applies known criteria to given text
38Poly-Crisis & Cascading Failure SimulatorA5Label as hypothetical, otherwise sound
41Clinical Workflow Implementation ScienceA5Advisory prose, no data claim
45Genetics Adversarial CounterpartA5Reasoning over given case
47EBM Adversarial CounterpartA5Reasoning over given case
51Systems-Based Clinical AnalysisA5Reasoning over given case
52Clinical Pearls DistillationA5Synthesis of given material
Level 4 — Reliable Output, Bypassed Learning
1Socratic Clinical ReasoningB4Correct answer, skips the reasoning rep
4Peer-Level Ward Round PreparationB4Same pattern
12Differential Diagnosis DeepdiveB4Same pattern
14Resource-Constrained Clinical ReasoningB4Same pattern
15Illness Script AcquisitionB4Same pattern
16Basic Science ↔ Clinical IntegrationB4Same pattern
17Semantic Qualifiers & Problem RepresentationB4Same pattern
18Causal vs. Probabilistic (Network) ReasoningB4Same pattern
20Naturalistic Decision MakingB4Same pattern
28Diagnostic Time-OutB4Same pattern
31First-Principles Pathophysiology MappingB4Same pattern
36 (LR math step)Bayesian Probability / Likelihood Ratio EngineB4Arithmetic once a probability exists
37Red Herring / Signal-to-Noise DrillB4Same pattern
44Clinical Genetics ReasoningB4Same pattern
50Diagnostic Reasoning MapB4Same pattern
56Hypothetico-Deductive ReasoningB4Same pattern
57Clinical Cognition Deep DiveB4Same pattern
Level 3 — Sound Reasoning, Fragile Specifics
10Journal & Article ReadingA / Fragile3Fine if pasted in; risky if "recalled"
39"Global Knowledge Network" Diagnostic MatrixA / Fragile3Name overclaims — trained-knowledge differential
40Operational & Throughput StrategistA / Fragile3Strategy sound, specific numbers invented
43Health Economics & Value-Based Care AlignmentA / Fragile3Cost figures illustrative, not sourced
46Evidence-Based Medicine InsightsA / Fragile3Exact recommendation wording drifts
48Treatment Comparative Analysis & Prognosis TrajectoryA / Fragile3Survival percentages not to be trusted
53Clinical Guideline Intelligence NavigatorA / Fragile3Guideline version/thresholds drift
Level 2 — Plausible Inference Only
26Bias AuditingC2Inferred bias, may not match learner's real bias
30"Diagnostic Anchor" ExtractorC2Inferred, not confirmed
33"Why Now?" (Precipitant) HunterC2Pattern-matched, may not be the true trigger
36 (prior-prob step)Bayesian Probability EngineC2Population prevalence ≠ contextual judgment
42Clinical Pre-MortemC2Generic scenarios, real blind spots may be missed
49FMEA AnalysisC2Literature-derived, institution-specific gaps missed
Level 1 — Structurally Can't Deliver Unmodified
6Registry-Level AnalyticsA / High-risk overclaim1No real registry queried — numbers fabricated
7Longitudinal & Cross-Case LearningA / High-risk overclaim1No real multi-case comparison happened
21Evidence Frontier SearchA / High-risk overclaim1No real search — stale or invented, reads like a lit review
27Time-Series & Velocity AnalyzerA / High-risk overclaim1Precise rate math unreliable without a calculator/code step
32Clinical Cognition LoopD1Nothing real to reflect on in single-pass
35Epistemic Certainty Mapping & CalibrationD1No real confidence to compare against
54System 1 & System 2 Question GeneratorD1Both "sides" AI-generated — comparison is empty
2.1–2.3, 2.8Patient-Advocate Case Documentation (data steps)E1Real symptom/exam/med facts have no substitute source
3Extended Patient-Advocate MonitoringE1Each entry is a real day's real observation
9N-of-1 Case Research ProtocolE1Real patient data-entry steps
13 (med-list step)Medication Reconciliation & Polypharmacy AuditE1Real medication list has no substitute source
55Patient Needs AssessmentE / special case1Plausible guess risks being mistaken for real patient wishes

5 · Scope for improvement, grouped by build cost

Grouped by what it actually costs to build — because the usability argument only holds if the cheap fixes get shipped first and the expensive ones are reserved for where they're truly needed.

SHIP FIRST — NO HEAVY INFRA Elicitation-first fixes

Not an architecture-stack problem at all — these need a real signal captured live, before generation, the same mandatory-descriptor mechanism proposed for a missed red-flag descriptor case. Cheapest, highest-integrity fix on the whole list.

LIGHT GROUNDING Pathway / knowledge-graph / RAG on already-sound modules

Reasoning shape is already sound — this is pure specifics-drift, and it's exactly what pathway grounding, a knowledge graph, or RAG is built to fix: converting a hedge into a checkable, sourced claim.

REAL INFRASTRUCTURE Where the claim genuinely requires a data source

These modules imply capabilities — a real registry, real cross-case history, live literature search — that no amount of clever prompting substitutes for. Worth building because the payoff is a full tag upgrade, not because it's the cheapest option.

LABELING DISCIPLINE — DON'T OVER-BUILD Modules 26, 30, 33, 42, 49

Part of Tag C hits a hard ceiling directly: a bias, a diagnostic anchor, or a "why now" trigger is learner- or patient-specific information no reference material holds. A knowledge graph can make the label checkable; it can never confirm the label is this learner's real bias. Building heavier architecture here mistakes a category boundary for a reliability gap — don't. The correct, cheap fix is consistent "AI's best guess" labeling, every time, no exceptions. (42 and 49's institution-specific pieces are a partial exception — a real incident-report or failure-mode database is genuinely buildable there.)

PERMANENTLY OUT OF SCOPE The noticing gap

"Knowledge management and reasoning tools can give you knowledge and reasoning — they cannot tell you what knowledge or reasoning you need."

No layer above — pathway, graph, RAG, GraphRAG, neuro-symbolic, or a well-built elicitation checklist — closes this. Every one of these tools operates on reasoning, text, patterns, or logical consistency; none of them operate on the patient. A checklist only catches the red flags someone thought to encode in advance; a presentation nobody anticipated still gets no question asked about it. This isn't a backlog item. It's the reason VibeRounds stays scoped to self-audit and learning, not decision support, regardless of how much of the above eventually gets built.

6 · Why usability still wins the tie

Something small enough to actually use beats something comprehensive that nobody opens — a ten-minute loop wins over a heavier one people abandon. Nothing above contradicts that. The "ship first" and "labeling discipline" tiers are both usability-first moves: they cost little, ship fast, and are exactly where extra architecture would have been wasted effort anyway, since each additional layer's marginal value shrinks as the base model itself keeps improving. The one place usability doesn't get the final word is where a module is currently producing content that looks real but was fabricated — a registry statistic that was never queried, a cross-case pattern that never happened. There, the simple fix and the honest fix are the same fix: don't run it unmodified, full stop, until the real data source exists. Usability is a tiebreaker for how to build well. It's never a justification for shipping something that quietly manufactures false confidence.