Vibe Rounds Cognitive Analytics Series · Part 5
Author - Dr. Avinash Kumar Gupta
Get started →It started as a single Socratic prompt written for clinical clerkship students — an attempt to make an AI ask questions rather than hand over answers, so a learner's own reasoning wasn't short-circuited before it started.
A single prompt isn't a system. What followed was less about ideas and more about plumbing: tracking how the prompt changed across iterations, and adjusting difficulty based on how a learner actually performed, instead of shipping one fixed script and hoping it fit everyone.
The next change was conceptual rather than technical. Rather than add another standalone tool, the existing modules were woven together under one shared instructional layer, built from four meta-skills:
Humanistic persona — specific affirmation of the reasoning move made, before any challenge, because challenge without affirmation triggers defensive thinking rather than open ones.
Fink's taxonomy — six non-hierarchical learning dimensions applied at closure, so a session produces durable insight, not just a correct answer.
Bloom's revised taxonomy — six cognitive levels mapped onto clinical reasoning tasks, used to calibrate difficulty between sessions.
Critical awareness — a standing closing prompt naming the biases the framework itself is prone to: automation bias, anchoring, hallucination risk, rare-diagnosis overweighting. The protocol auditing itself, by design.
With the layer defined, it needed a working instance. Module 01 — Socratic Clinical Reasoning — was built with explicit initiation, execution, and closure steps, and became the template every later module would follow.
Once the pattern held, it repeated: patient advocate documentation, ward-round preparation, and a widening set of distinct clinical cognition skills, each following the same shape rather than reinventing structure each time.
At some point the modules stopped being a loose collection of prompts and became components of an actual builder — what would later be called the Clinical Cognition Operating System (CCOS).
The next experiment applied a knowledge-graph concept: a structured knowledgebase of clinical concepts with explicit interrelations. It produced a working tutor tool and a few related demos. They were functional — but not especially useful or accurate in practice. That result is worth stating plainly rather than quietly dropping: it was a real attempt, it did not clearly outperform the simpler approach already working, and that was itself the signal to redirect effort rather than keep layering complexity onto a structure that wasn't earning its cost. The same judgment call underlies Part 3's argument that added structure has to justify its own verification burden — this is where that lesson was first learned in practice, not in the abstract.
Redirected effort converged the modules and the CCOS architecture into one consolidated tool set — the "OS" home most of the current tools now launch from.
The system was eventually written up as a design-rationale paper describing the CCOS architecture: step-numbered micro-modules, structural protocol lifecycles, and the cross-cutting pedagogical layers described above. Getting there wasn't a straight line. Several earlier AI-generated write-ups of the same work were produced first, then deleted — for being inaccurate, or overloaded with jargon that obscured more than it explained — before the writing was tightened into something worth keeping. That editing process is not a footnote; it's the same critical-appraisal habit this whole series argues for, applied to the AI's own account of itself rather than to a patient case.
The last piece kept from this stretch of work is deliberately minimal: a tool built specifically to collect data, meant to turn the system from something built into something that can be studied. That is where this series and the underlying system currently stand — Part 3 named the proposed empirical study; this is the infrastructure waiting to run it.
Part 4 argued for restraint — the smallest usable version of the stack. This part is the evidence for why that restraint was earned rather than assumed: it came from watching a more elaborate approach (the knowledge-graph layer) fail to justify itself, and from discarding write-ups of the work that didn't survive their own scrutiny. The framework in Parts 1–4 is easier to trust once it's clear it was built by cutting things, not just adding them.