Planning artifact — not an authorization. This plan does not authorize new data collection, participant-level linkage, cross-border data transfer, analysis of identifiable or minor-participant data, or any compliance certification. The prerequisites it names must be resolved before execution.
Illustrative sample — a fictional program, AI-drafted and stamped DRAFT for human review.
The questions this evaluation answers
The Youth Technology & Entrepreneurship Program is a globally delivered, ~12-week, team-based initiative for youth ages 10–18 that aims to build coding, entrepreneurship, and 21st-century skills through community-problem-focused app development and a tiered competition. This evaluation starter package frames a mixed-methods plan organized around six evaluation functions: demonstrating outcome evidence, assessing implementation and delivery, examining reach and equity, establishing measurement quality, enabling use and learning (CQI), and centering stakeholder voice. The plan cannot yet be fully executed: eight high-priority prerequisites — IRB/ethics review, participant identifier linkage, rubric standardization, non-response analysis, data linkage confirmation, fidelity data, language equivalence documentation, and cohort-size confirmation — must be resolved before the most important indicators can be computed. Once resolved, the dosage-response pattern and triangulated demonstrated-skill evidence represent the strongest feasible within-program contribution argument available without a comparison group.
q1 — Do participants actually gain skills over a season — demonstrated, not just self-reported confidence — and how can that be shown credibly? — The program cannot yet show demonstrated skill gain: outcomes rest almost entirely on retrospective self-reported confidence, which is susceptible to recall and response-shift bias, and no standardized, validated judging rubric with confirmed inter-rater reliability is yet in use — so skill gain cannot yet be distinguished from self-report artifacts or enthusiasm effects.
q2 — Is survey response low and non-random — are we primarily hearing from the most engaged participants and competition winners — and how does that bias our understanding of outcomes? — The program cannot yet show that its survey-based outcome estimates represent the full participant population: response-rate funnels by segment and respondent vs. non-respondent profile comparisons have not been computed, so the direction and magnitude of self-selection bias remain unquantified.
q3 — Do our measures mean the same thing across countries, languages, and age groups, such that cross-country comparisons are valid (measurement equivalence / cross-cultural validity)? — The program cannot yet show that cross-country survey comparisons are valid: measurement equivalence has not been tested, and the number of language versions and translation processes are insufficiently documented, so reported cross-country differences cannot yet be distinguished from translation or cultural-interpretation artifacts.
q4 — Is the program reaching and serving participants equitably across home computer/internet access, whether coding is taught at school, country/region, and age? — The program cannot yet show equitable reach or outcomes: per-site enrollment counts and equity-stratified completion rates are not yet reported, and equity gaps in outcomes cannot be computed until non-response bias is characterized and minimum cell sizes are confirmed.
q5 — Is mentoring and curriculum delivery consistent across teams and regions, and does that variation matter for participant outcomes? — The program cannot yet show whether delivery variation drives outcome variation: implementation fidelity data and mentor training records do not currently exist, and variance decomposition cannot be computed without confirmed data linkage.
q6 — How can the program demonstrate credible impact to funders in the absence of a comparison group? — The program cannot yet show credible impact: no comparison group exists, dosage-response analysis requires confirmed attendance data and linkage, and multi-season trend analysis requires comparable prior-season data — so the strongest feasible within-program contribution arguments are not yet computable.
q7 — How much does the program reach — how many participants complete a season, across which countries and demographic segments? — The program cannot yet report accurate reach: exact per-season cohort sizes, per-country and per-site counts, and equity-stratified completion rates are not yet confirmed, limiting funder reporting and power calculations for subgroup analyses.
q8 — Do participants sustain interest in STEM education and careers beyond the program season? — The program cannot yet show sustained STEM pathway outcomes: alumni consent-to-recontact and current contact information are unconfirmed, and alumni survey feasibility has not been established — so long-term outcomes cannot yet be distinguished from selection effects among alumni who remain in contact.
The decisions only you can make
EvalSmart names these; your team owns them. A few must be resolved before any data collection begins — the full checklist is in the executive brief.
File IRB/ethics review and confirm existing consent covers evaluation use before any data collection or analysis begins
Confirm governance-approved participant identifier for pre/post survey linkage with data owner and legal/ethics reviewer
Confirm judging rubric standardization and double-scoring feasibility before computing M-7, M-8, M-9
Compute non-response analysis before interpreting any survey-based outcome indicator
Confirm cross-source data linkage permissibility with legal/ethics reviewer before finalizing AP-6, AP-7, AP-8
Confirm secure, governance-approved qualitative data storage environment before scheduling interviews or focus groups
Priority measures — and what blocks them
Status — Ready: computable now; Needs setup: you supply a value or data first (a baseline, threshold, identifier, or linkage); Blocked: a decision or approval must happen first (IRB / privacy / consent, a scope or design decision, or instrument mapping).
Outcome Evidence
Evaluation quality (JCSEE): Accuracy
Measure (quant): M-7 — CANDIDATE demonstrated-skill evidence — app quality rubric score: judge-assigned rubric scores on the final mobile app (functionality, coding complexity, problem-relevance), disaggregated by region and equity strata (status: Needs setup)
Explain (qual): QQ1 — Why do participants who complete the season report varying levels of skill gain — what program experiences, mentor behaviors, and personal factors do they associate with feeling more or less capable in coding, entrepreneurship, and problem-solving?
M-7 (app quality rubric score) is the most direct demonstrated-skill indicator, moving beyond self-report; QQ1 explains why participants report varying skill gain and what program experiences drive it — together they triangulate the program's core effectiveness claim.
Implementation & Delivery
Evaluation quality (JCSEE): Accuracy
Measure (quant): M-13 — Implementation fidelity index: proportion of curriculum sessions delivered as intended per team/region, based on session logs or fidelity checklist (status: Needs setup)
Explain (qual): QQ5 — How do mentors understand their role, interpret the curriculum, and adapt their delivery — and what factors (training, time, confidence, regional context) shape the consistency or variability of their mentoring?
M-13 (implementation fidelity index) quantifies whether curriculum is delivered as intended across teams and regions; QQ5 explains how mentors interpret and adapt their delivery — together they address whether delivery variation drives outcome variation (q5).
Reach / Equity
Evaluation quality (JCSEE): Propriety
Measure (quant): M-11 — Equity gap in survey outcomes: difference in self-reported confidence change scores and judging rubric scores between equity subgroups (home computer/internet access: yes vs. no; coding taught at school: yes vs. no; age group; country income level) (status: Blocked)
Explain (qual): QQ4 — How do participants with limited home computer/internet access, or without school-based coding instruction, experience the program differently — what barriers do they encounter and what adaptations (if any) help them succeed?
M-11 (equity gap in survey outcomes) quantifies differences in outcomes across equity subgroups (access, school coding, age, country); QQ4 explains the barriers and adaptations experienced by participants with limited access — together they address whether the program reaches and serves all participants equitably (q4).
Measurement Quality
Evaluation quality (JCSEE): Accuracy
Measure (quant): M-10 — Cross-country measurement equivalence of survey scales: configural, metric, and scalar invariance tests for each multi-item survey scale across language/country groups (status: Needs setup)
Explain (qual): QQ3 — How do youth in different countries and language contexts understand and interpret key survey concepts (e.g., 'confidence in coding,' 'entrepreneurship,' 'data literacy') — do these terms carry equivalent meaning across cultural settings?
M-10 (cross-country measurement equivalence) statistically tests whether survey scales mean the same thing across language/country groups; QQ3 explores how youth in different cultural contexts interpret key survey concepts — together they are the prerequisite for any valid cross-country comparison (q3).
Use & Learning (CQI)
Evaluation quality (JCSEE): Utility
Measure (quant): M-17 — Multi-season trend in key outcomes: season-over-season change in completion rate, survey response rate, mean confidence change scores, and mean judging scores (status: Needs setup)
Explain (qual): QQ10 — What do program leaders, regional coordinators, and mentors believe constitutes credible evidence of impact — and what kinds of evidence would be most persuasive to their funders and communities?
M-17 (multi-season trend in key outcomes) tracks whether the program is improving over time; QQ10 surfaces what program leaders and funders consider credible evidence of impact — together they support the learning and accountability decisions the evaluation is designed to inform (q6).
Top gaps
IRB/ethics review and consent coverage not confirmed — blocks all data collection
No governance-approved participant identifier for pre/post linkage — blocks individual-level change analysis
Judging rubric not confirmed as standardized or double-scored — blocks demonstrated-skill evidence
Non-response analysis not yet computed — all current outcome estimates potentially severely biased
Cross-source data linkage permissibility unconfirmed — blocks dosage-response and variance decomposition analyses
Produced by EvalSmart — draft for human review.
This EvalSmart starter package converts a program description into an evaluation-ready brief: priority evaluation functions, quantitative indicators, qualitative follow-up questions, evidence gaps, and prerequisite decisions requiring human review.
Standards and frameworks referenced above are included as contextual alignment prompts and must be verified by your institution before any accreditation or compliance use. EvalSmart cites them as organizing context, not as certification of compliance.
This is the short sample. A full EvalSmart run also produces a fuller executive brief and a comprehensive technical plan — every item tagged stated, inferred, or gap, and reviewed by a human.