STATUS: DRAFT — REQUIRES HUMAN REVIEW
Illustrative sample — a fictional research consultation service. This is an actual EvalSmart run (academic-library / ACRL pack), AI-drafted and stamped DRAFT for human review.
The library runs a by-appointment, one-on-one research consultation service — about a dozen subject librarians, 30–60-minute sessions, in person and online. Heading into its program review, it needs to show the service's impact, not just its session counts — but today it has only a low-response satisfaction survey and volume data. It brought four questions to an evaluator:
EvalSmart's job wasn't to answer these — it was to find out whether the evidence on hand can, and to name what's in the way.
Each of the four questions runs into something the current evidence can't yet support — and a specific measure or input that would close it.
The service measures satisfaction and volume, but no outcome data exists yet, so "consultations improve research skills" can't be evidenced — and the survey that does run has low, self-selected response. A standardized patron-reported outcome measure would answer q1, but a baseline has to be established first.
q1 · standardized patron-reported outcome measure — Needs setup · no outcome baseline yet · outcome instrument is an inferred candidate
That sessions run very differently across the ~dozen librarians is, in EvalSmart's own words, an untested assumption. The variance analysis that would settle it — outcome variance attributable to the librarian — needs consultation records every librarian fills in the same way, and completeness varies by librarian today.
q2 · librarian-variance analysis — Needs setup · needs uniform consultation records across librarians
Utilization by patron group, discipline, and modality is Blocked: there are no campus-population denominators yet, and the tracking data hasn't been audited for completeness. Any claim about who is underserved is premature until those two land.
q3 · utilization vs. campus population — Blocked · no denominators; tracking not yet audited
The program-review body, the evidence it requires, and the deadline aren't identified yet, so there's no target to design toward. And the qualitative strand that would explain the numbers is blocked on IRB clearance. Strengthening the assessment starts with pinning those down.
q4 · survey-reliability & the review-body bar — IRB clearance pending · target not yet defined
Beyond the four it was asked, EvalSmart also surfaced two analytic follow-ons the brief never named — response bias in the current survey (q5) and subgroup outcome differences (q6) — because both bear on whether the original four can be answered defensibly.
Before any data collection, six choices need a human — and they're the prerequisites the run flagged, made concrete:
Smaller calls follow once these are set (census vs. sampling; whether faculty/staff are in scope; whether a pre-consultation self-assessment enables a pre/post design). EvalSmart names each decision and what it unlocks; it does not make them for you — and it maps the plan to ACRL's SLHE principles only as organizing context, never as a compliance verdict (ACRL is a professional standards body, not an accreditor).
Stripped to its spine, each question runs into the gap blocking it and the human decision that clears the way:
That's the skeleton. The full technical plan — every indicator and qualitative question, each gap, each tagged stated / inferred / gap and carrying its ID — is the comprehensive report that accompanies the run. This page is the narrative; the comprehensive plan is the source of truth.
Generated by an EvalSmart academic-library run and reframed as a case study. The full executive brief and comprehensive technical plan (logic model, indicator matrix, qualitative protocol, full gap memo) accompany every run — every item tagged stated, inferred, or gap, and reviewed by a human.