EvalSmart — General Surgery Core Clerkship — Sample Evaluation Brief
STATUS: DRAFT — REQUIRES HUMAN REVIEW
Planning artifact — not an authorization. This plan does not authorize new data collection, participant-level linkage, cross-border data transfer, analysis of identifiable or minor-participant data, or any compliance certification. The prerequisites it names must be resolved before execution.
Illustrative sample — a fictional program, AI-drafted and stamped DRAFT for human review.
The questions this evaluation answers
The General Surgery Core Clerkship evaluation plan addresses four leadership priorities — competency attainment, cross-site assessment consistency, equity of hands-on opportunity, and assessment-system strengthening — through a 16-indicator quantitative strand and an 8-question explanatory sequential qualitative strand. Three prerequisites must be resolved before any substantive analysis can begin: (1) performance standards must be defined for all assessment instruments; (2) IRB and data-governance clearance must be obtained; and (3) the clinical evaluation instrument must be confirmed as consistent across all sites.
q1 — Are students achieving the five intended competencies — including basic technical and procedural skills — by the end of the clerkship? — The program cannot yet show that students are meeting any of the five competency standards because no pass/fail thresholds or performance standards have been defined for any assessment instrument, and the assessment blueprint has not been confirmed to cover all five competencies with scored instruments.
q2 — Is clinical-evaluation scoring consistent across rotation sites, and does observed variability reflect real differences in student performance or differences in rater/service context? — The program cannot yet distinguish true student performance differences from rater- or site-effect artifacts because it is unknown whether all sites use the same clinical evaluation instrument, rater training and calibration practices are undescribed, and individual rater IDs may not be captured in evaluation records.
q3 — Is the student experience equitable across sites — specifically, do students at different sites receive comparable hands-on operative and procedural opportunity? — The program cannot yet show that operative and procedural opportunity is equitable across sites because case-log accuracy is unverified, minimum case requirements are undefined, implementation fidelity of curricular components across sites is unknown, and pre-clerkship student characteristics by site are not available to rule out selection effects.
q4 — Where can the assessment system be strengthened, particularly for technical and procedural skills currently assessed mainly by a brief simulation checklist and subjective faculty impressions? — The program cannot yet characterize the reliability or validity of its technical-skills assessment because no double-scored simulation sessions are confirmed to exist for inter-rater reliability analysis, direct-observation assessment of OR and bedside skills beyond the checklist is not systematically documented, and the assessment blueprint has not been formally mapped to all five competencies.
q5 — Are there systematic differences in outcomes or opportunity by student demographic characteristics (e.g., gender, race/ethnicity, prior experience) that warrant equity-focused attention? — The program cannot yet show whether demographic equity exists because student demographic data have not been confirmed as available or linkable to assessment records under current data-governance policies, and IRB clearance for this analysis has not been established.
The decisions only you can make
EvalSmart names these; your team owns them. A few must be resolved before any data collection begins — the full checklist is in the executive brief.
1: Establish performance standards (pass/fail thresholds) for all assessment instruments before computing M-1 through M-4
2: Obtain IRB and data-governance clearance for multi-source data linkage and qualitative data collection before beginning any data collection
3: Confirm that all sites use the same clinical evaluation instrument and rating scale before conducting any cross-site score comparisons
As the first prerequisite step, convene clerkship leadership to determine whether performance standards already exist in unpublished form; if not, initiate a standard-setting process (e.g., modified Angoff, contrasting groups, or expert panel) for each scored instrument before computing proportion-at-standard indicators M-1 through M-4
Consult the institution's IRB and data-governance office before finalizing the evaluation design; determine whether a formal review, exemption determination, or data-use agreement is required, and build the resulting timeline into the evaluation plan
Obtain and compare the clinical evaluation forms currently in use at each site; if forms differ, cross-site score comparisons are invalid until a common instrument is adopted or a crosswalk is established
Priority measures — and what blocks them
Status — Ready: computable now; Needs setup: you supply a value or data first (a baseline, threshold, identifier, or linkage); Blocked: a decision or approval must happen first (IRB / privacy / consent, a scope or design decision, or instrument mapping).
Outcome Evidence
*LCME 8.4 Evaluation of Educational Program Outcomes; LCME 6.1 Program and Learning Objectives* Evaluation quality (JCSEE): Accuracy
Measure (quant): M-4 — Proportion of students meeting the program-defined competency-attainment standard on end-of-rotation clinical evaluations, overall and by competency domain (status: Needs setup)
Explain (qual): QQ7 — How do students and faculty understand and enact the five intended competencies in daily clinical work — and do these enacted understandings align with the program's formal competency definitions?
M-4 (proportion meeting the competency standard on clinical evaluations, overall and by domain) is the broadest competency-attainment indicator spanning all five competencies; QQ7 probes whether students and faculty enact the five competencies in daily work in ways that align with formal program definitions, explaining what the numbers mean in practice.
Clinical Performance
*LCME 9.4 Assessment System; LCME 6.1 Program and Learning Objectives* Evaluation quality (JCSEE): Utility
Measure (quant): M-8 — Mean number and type of operative cases logged per student by site (primary surgeon role vs. assistant vs. observer), and proportion of students meeting program-defined minimum case requirements (status: Needs setup)
Explain (qual): QQ2 — What site-level contextual factors — case mix, service culture, team hierarchy, physical environment — do students and faculty perceive as shaping hands-on operative and procedural opportunity?
M-8 (mean operative cases logged per student by site and role) is the primary quantitative measure of hands-on clinical performance opportunity; QQ2 explores the site-level contextual factors — case mix, service culture, team hierarchy — that faculty and students perceive as shaping that opportunity, providing the explanatory layer.
Measure (quant): M-5 — Decomposition of clinical-evaluation score variance into student, rater, and site components (G-study or ICC by site) (status: Blocked)
Explain (qual): QQ1 — How do attending surgeons and residents interpret and apply the clinical evaluation instrument's performance anchors when rating students — and do these interpretations differ systematically across sites or rater roles?
M-5 (variance-components decomposition of clinical evaluation scores into student, rater, and site effects) directly quantifies the measurement quality problem leadership has already identified; QQ1 investigates how raters interpret and apply performance anchors across sites and roles, explaining the sources of variance the numbers reveal.
Equity / Comparability
*LCME 8.7 Comparability of Education/Assessment; LCME 8.6 Monitoring of Completion of Required Clinical Experiences* Evaluation quality (JCSEE): Propriety
Measure (quant): M-14 — Differential outcomes by student demographic group: mean shelf score, mean clinical evaluation score, and mean operative case count, disaggregated by gender and race/ethnicity (status: Blocked)
Explain (qual): QQ3 — How do students experience the equity of their learning environment across sites — including access to procedures, quality of supervision, and inclusion within the surgical team — and do experiences differ by student demographic characteristics?
M-14 (differential outcomes by student demographic group across shelf scores, clinical evaluation scores, and operative case counts) is the most direct quantitative equity indicator; QQ3 captures how students experience equity of learning environment, access to procedures, and inclusion within the surgical team — and whether experiences differ by demographic characteristics.
Measure (quant): M-10 — Assessment-blueprint coverage: proportion of the five program competencies (and, if applicable, AAMC Core EPAs) covered by at least one valid, scored assessment instrument (status: Needs setup)
Explain (qual): QQ6 — What do clerkship leaders, faculty, and students perceive as the most significant gaps in the current assessment system, and what changes would they prioritize to strengthen competency-based evaluation?
M-10 (assessment blueprint coverage: proportion of the five competencies covered by at least one valid, scored instrument) provides a structural map of the assessment system's completeness; QQ6 gathers faculty and student perceptions of the most significant gaps and priority changes, grounding the blueprint analysis in stakeholder experience.
Top gaps
No performance standards defined for any assessment instrument — blocks Q1 entirely
IRB/data-governance clearance not confirmed — blocks all data collection
Rater training and calibration practices undescribed — limits interpretation of score variability
Operative case log accuracy and completeness unverified — undermines Q3 equity analysis
Produced by EvalSmart — draft for human review.
This EvalSmart starter package converts a program description into an evaluation-ready brief: priority evaluation functions, quantitative indicators, qualitative follow-up questions, evidence gaps, and prerequisite decisions requiring human review.
Standards and frameworks referenced above are included as contextual alignment prompts and must be verified by your institution before any accreditation or compliance use. EvalSmart cites them as organizing context, not as certification of compliance.
This is the short sample. A full EvalSmart run also produces a fuller executive brief and a comprehensive technical plan — every item tagged stated, inferred, or gap, and reviewed by a human.