EvalSmart

General Surgery Core Clerkship — Evaluation case study

STATUS: DRAFT — REQUIRES HUMAN REVIEW

Illustrative sample — a fictional program, AI-drafted and stamped DRAFT for human review. This is a narrative hook; the full executive brief and comprehensive technical plan accompany every run.

The problem

The General Surgery Core Clerkship is an eight-week rotation in which students train across an academic medical center, community hospitals, and ambulatory surgery centers, assessed across five competency domains.

Faculty have flagged a measurement-quality concern: clinical-evaluation scores vary, but it isn't clear how much of that reflects real differences in student performance versus different graders, different patient volumes, and different operative opportunity from site to site. As it stands, the program can't yet say whether students reach comparable competence across sites — there are no defined performance thresholds, no site-adjusted comparison, and the technical-skills assessment leans on a checklist faculty themselves consider inadequate.

What EvalSmart surfaced

From the program description alone, the run isolated the handful of issues that actually decide whether this evaluation can stand up — not another wall of indicators.

The program cannot yet demonstrate comparable training opportunity across sites.

No site-adjusted comparison of operative case volume or hands-on role has been built, so "students receive equivalent surgical training across sites" is currently an assumption — in either direction. EvalSmart identified the comparison (effect sizes across academic, community, and ambulatory sites) that would settle it.

Equity / Comparability keystone · Needs setup — no site-adjusted comparison built yet

The program can't yet substantiate that students achieve competency.

No performance thresholds are defined for any instrument, so a competency-attainment claim cannot yet be interpreted defensibly. Defining thresholds is the prerequisite that unlocks the competency-attainment measures.

Prerequisite: define performance thresholds — blocks the competency-attainment measures

Site score differences can't yet be interpreted — real gap, or just different graders?

Until rater and site identifiers are confirmed as captured and linkable, the variance-decomposition analysis that separates student, rater, and site effects can't run — so lower site-level scores could reflect stricter grading rather than a weaker educational experience.

Variance decomposition · Blocked — rater/site identifiers not yet confirmed

Operative competence is being judged with an instrument faculty already call inadequate.

The current simulation-lab checklist plus subjective impressions can't anchor a defensible competence claim. Two validated alternatives are surfaced for the Clerkship Director to weigh — options to verify, not a prescription.

Validated alternatives to weigh (e.g. OSATS, O-SCORE) — inferred options to verify

The decisions only you can make

EvalSmart names them · the program owns them

Before data collection, four prerequisites and two inputs need a human decision:

  1. Define passing / competency-achievement thresholds for every instrument — the prerequisite to any attainment claim.
  2. Confirm rater and site identifiers are captured and linkable in the clinical-evaluation dataset — the prerequisite to variance-components and rater-leniency analysis.
  3. Secure IRB approval or an exemption determination before any qualitative data collection begins.
  4. Confirm cross-site data-linkage feasibility and data-use agreements before any multi-instrument, student-level analysis.
  5. Provide cohort numbers — students per year, rotation blocks, and students per site per block — so power calculations and cross-site subgroup analyses are feasible.
  6. Share the clinical-evaluation form and confirm its scale, anchors, and competency mapping.

These are judgment and governance calls. EvalSmart flags each one and explains what it unlocks — it does not make them for you. Standards (e.g. LCME, JCSEE) are cited as organizing context and must be institution-verified, never as a compliance verdict.

From question to decision

Stripped to its spine, each concern runs into the gap blocking it and the human decision that clears the way:

That's the skeleton. The full technical plan — every indicator and qualitative question, each gap, each tagged stated / inferred / gap and carrying its ID — is the comprehensive report that accompanies the run. This page is the narrative; the comprehensive plan is the source of truth.


Generated by an EvalSmart run and reframed as a case study. The full executive brief and comprehensive technical plan (logic model, indicator matrix, qualitative protocol, full gap memo) accompany every run — every item tagged stated, inferred, or gap, and reviewed by a human.