EvalSmart

Youth Technology & Entrepreneurship Program — Evaluation case study

STATUS: DRAFT — REQUIRES HUMAN REVIEW

Planning artifact — not an authorization. This plan does not authorize new data collection, participant-level linkage, cross-border data transfer, analysis of identifiable or minor-participant data, or any compliance certification. The prerequisites it names must be resolved before execution.

Illustrative sample — a fictional program, AI-drafted and stamped DRAFT for human review. This is a narrative hook; the full executive brief and comprehensive technical plan accompany every run.

The problem

The program is a ~12-week, mentor-led season in which teams of youth (ages 10–18), across many countries and languages, build a mobile app to solve a community problem and pitch it in a tiered competition. It is funded by grants and donors and reports to a board.

Leadership wants to show funders real impact — but the program can't yet do so credibly. Outcomes rest almost entirely on a retrospective self-report survey whose response has been low and probably not random; there is no demonstrated-skill measure, no comparison group, and the survey has never been checked for whether it means the same thing across languages and ages. And because participants are minors across many jurisdictions, the data-governance picture is unresolved.

What EvalSmart surfaced

From the program description alone, the run isolated the handful of issues that actually decide whether this evaluation can stand up — not another wall of indicators.

The program can't yet show skill — only confidence.

Every outcome rests on self-reported confidence, and confidence is not competence. Without a validated, consistently scored rubric for the apps and pitches, a "participants gained skills" claim can't be made. EvalSmart names the keystone that would fix it: a rubric-scored measure of the work students actually produce.

Outcome Evidence keystone · Needs setup — no validated skills rubric in place

The results may only reflect the most engaged.

Survey response is low and likely non-random — the participants who answer tend to be the most engaged and the competition winners. Until response rates by segment are documented and respondents are compared to non-respondents, reported gains may be biased upward and unrepresentative.

Non-response analysis · Needs setup — response rates by segment undocumented

Cross-country comparisons aren't yet valid.

The survey has never been tested for measurement equivalence across languages and age bands. Until it is, aggregated global findings and country-by-country comparisons may reflect translation or developmental artifacts rather than real differences in outcomes.

Measurement equivalence · Needs setup — invariance untested across languages/ages

With minors across jurisdictions, governance comes before data.

Participants are minors and data-protection rules vary by country; consent-for-evaluation, a stable participant identifier, and ethics review are unconfirmed. EvalSmart treats this as a prerequisite to execution, not a reason the plan can't be designed — the plan is produced and leads with these as binding prerequisites.

Execution prerequisite · resolve consent / linkage / ethics before any data collection

The decisions only you can make

EvalSmart names them · the program owns them

Before data collection, the prerequisites and inputs that need a human decision:

  1. Confirm consent-for-evaluation and the multi-jurisdiction data-governance picture, and confirm ethics/IRB review for collecting data from minors — before any data collection begins.
  2. Verify a stable participant identifier links survey, competition, and administrative records — the prerequisite to any individual-level pre/post or attrition analysis.
  3. Adopt and pilot a validated app-and-pitch rubric and establish inter-rater reliability — the prerequisite to any demonstrated-skill claim.
  4. Plan measurement-equivalence testing before any cross-country or cross-age comparison is reported.
  5. Provide cohort scale — enrollment, number of countries/sites, and participants per site — so subgroup analyses and power are feasible.
  6. Design a response-rate and non-response strategy so the outcome picture is representative.

These are judgment and governance calls. EvalSmart flags each one and explains what it unlocks — it does not make them for you. Standards (e.g. the JCSEE Program Evaluation Standards, AEA Guiding Principles) are cited as organizing context and must be verified by your organization, never as a compliance verdict.

From question to decision

Stripped to its spine, each concern runs into the gap blocking it and the human decision that clears the way:

That's the skeleton. The full technical plan — every indicator and qualitative question, each gap, each tagged stated / inferred / gap and carrying its ID — is the comprehensive report that accompanies the run. This page is the narrative; the comprehensive plan is the source of truth.


Generated by an EvalSmart run and reframed as a case study. The full executive brief and comprehensive technical plan (logic model, indicator matrix, qualitative protocol, full gap memo) accompany every run — every item tagged stated, inferred, or gap, and reviewed by a human.