STATUS: DRAFT — REQUIRES HUMAN REVIEW
Illustrative sample — a fictional program, AI-drafted and stamped DRAFT for human review. This is a narrative hook; the full executive brief and comprehensive technical plan accompany every run.
The General Surgery Core Clerkship is an eight-week rotation in which students train across an academic medical center, community hospitals, and ambulatory surgery centers, assessed across five competency domains.
Faculty have flagged a measurement-quality concern: clinical-evaluation scores vary, but it isn't clear how much of that reflects real differences in student performance versus different graders, different patient volumes, and different operative opportunity from site to site. As it stands, the program can't yet say whether students reach comparable competence across sites — there are no defined performance thresholds, no site-adjusted comparison, and the technical-skills assessment leans on a checklist faculty themselves consider inadequate.
From the program description alone, the run isolated the handful of issues that actually decide whether this evaluation can stand up — not another wall of indicators.
No site-adjusted comparison of operative case volume or hands-on role has been built, so "students receive equivalent surgical training across sites" is currently an assumption — in either direction. EvalSmart identified the comparison (effect sizes across academic, community, and ambulatory sites) that would settle it.
Equity / Comparability keystone · Needs setup — no site-adjusted comparison built yet
No performance thresholds are defined for any instrument, so a competency-attainment claim cannot yet be interpreted defensibly. Defining thresholds is the prerequisite that unlocks the competency-attainment measures.
Prerequisite: define performance thresholds — blocks the competency-attainment measures
Until rater and site identifiers are confirmed as captured and linkable, the variance-decomposition analysis that separates student, rater, and site effects can't run — so lower site-level scores could reflect stricter grading rather than a weaker educational experience.
Variance decomposition · Blocked — rater/site identifiers not yet confirmed
The current simulation-lab checklist plus subjective impressions can't anchor a defensible competence claim. Two validated alternatives are surfaced for the Clerkship Director to weigh — options to verify, not a prescription.
Validated alternatives to weigh (e.g. OSATS, O-SCORE) — inferred options to verify
Before data collection, four prerequisites and two inputs need a human decision:
These are judgment and governance calls. EvalSmart flags each one and explains what it unlocks — it does not make them for you. Standards (e.g. LCME, JCSEE) are cited as organizing context and must be institution-verified, never as a compliance verdict.
Stripped to its spine, each concern runs into the gap blocking it and the human decision that clears the way:
That's the skeleton. The full technical plan — every indicator and qualitative question, each gap, each tagged stated / inferred / gap and carrying its ID — is the comprehensive report that accompanies the run. This page is the narrative; the comprehensive plan is the source of truth.
Generated by an EvalSmart run and reframed as a case study. The full executive brief and comprehensive technical plan (logic model, indicator matrix, qualitative protocol, full gap memo) accompany every run — every item tagged stated, inferred, or gap, and reviewed by a human.