Peer feedback is the engine of The Geneva Learning Foundation’s learning system. A new screening tool protects that engine, and its design answers three questions education specialists have been circling for decades: what peer assessment is for, how much verification a learning system can afford, and what a certificate can honestly claim.
“Ok ok ok ok” and the review that changed a clinic
In one of three sample courses used to pilot the system this article describes, a physician in Tanzania reviewed a colleague’s account of a woman whose symptoms had been dismissed for years. Her feedback did four things in three sentences. It connected the colleague’s story to her own practice. It named a specific gap in the account. It asked a question that pushed the analysis further. And it committed the reviewer herself to a change: she would start asking her own patients about these symptoms in the consultation room rather than at triage. Three weeks later, in the course forum, she reported that she had kept the habit every week since, and had begun educating patients alongside their spouses.
In the same course, another reviewer submitted feedback that read, in its entirety: “Ok ok ok ok ok ok ok ok oo kok ok ok ok ok.” It passed the platform’s automated word-count check.
Both reviews sat in the same dataset and counted identically in the completion statistics. The distance between them is the subject of this article. Peer learning produces extraordinary contributions and empty ones, and a system that cannot tell the difference at scale will eventually be defined by the empty ones, because the learners writing the extraordinary ones are the first to notice when nobody is protecting their effort. The Geneva Learning Foundation (TGLF) has built and piloted a screening system, the TGLF Reviewer, that reads every submission, every review, and every reflection in a course and produces evidence about which learners kept their commitments to their peers.
Calling this a system for catching cheaters misses what assessment is doing in the design. In TGLF courses, assessment exists to scaffold peer learning: the rubric, the ratings, and the comment prompts are structures that teach participants to give feedback worth receiving. The screening system keeps that scaffold load-bearing. Its purpose is not surveillance but protection of the conditions under which a health worker in one country writes three sentences that change practice in another.
Three commitments, and what breaks when one fails
TGLF courses follow a pattern the foundation calls structured peer learning, grounded in what Reda Sadki’s recent theoretical work names Structured Peer Praxis: the claim that under conditions of complexity, learning and implementation are one process, and that just enough structure makes that process reliable at scale (Sadki, 2026).
A course of this type asks each participant for three commitments. First, answer three learning questions, each demanding one specific, real situation from the participant’s own practice rather than a textbook case. The questions move deliberately from personal observation to systemic analysis to a concrete plan for the next four weeks. Second, review three colleagues’ answers against a rubric whose criteria match the three questions, with instructions that are explicit about what feedback is for: name a specific strength, ask a question that pushes thinking further, share a connection from your own work. Third, respond to the feedback received, stating what the participant took from it and what they will do differently.
Each review carries a rating on a five-level developmental scale: 0 (No answer), 1 (Getting started), 2 (Building), 3 (Strong), 4 (Exceptional). The levels describe how far a contribution has traveled, not how smart its author is. The gap between “Building” and “Strong,” the two most consequential neighbors on the scale, is the gap between an answer that has begun to meet the criteria and one that meets them fully: a “Building” response typically tells a real story but stops short somewhere, the reflection is missing, the action plan has no named beneficiary, the analysis stays general, while a “Strong” response is specific and complete across the criterion. Reviewers are told the scale describes the contribution rather than grades the person, and that “a community health worker in a rural village and a doctor in an urban hospital can both earn the highest rating.”
Accountability in this design runs horizontally. When a participant writes a thin review, the injured party is the colleague who spent an evening writing a real account of a real patient and received “Nice” in return, not an institution. Horizontal accountability is the design’s strength and its exposure at once: peer accountability produces the learning value, and peer accountability is exactly what a minority of participants will not honor. In the pilot course, 91.5% of learners received at least one review, and 99.8% of review comments cleared the platform’s automated substance threshold. That second number is the tell. A threshold that passes “Ok ok ok ok” fourteen times over is not measuring substance. Something else has to.
When two reviewers cannot agree, what exactly is being assessed?
Education researchers have asked for decades whether peer assessment can be trusted as measurement (Topping, 1998). TGLF’s own data gives the honest answer: no, and that is not what it is for.
The pilot course’s evaluation computed inter-reviewer agreement using weighted kappa, a statistic that asks how often two reviewers assign the same rating to the same work after subtracting the agreement you would expect from chance alone, with partial credit when they land one level apart rather than far apart. A kappa of 1.0 means perfect agreement, and 0 means the reviewers might as well have been rolling dice. Across the three rubric criteria, weighted kappa ranged from 0.06 to 0.17, which the measurement literature labels slight agreement, even though the raw numbers looked respectable: reviewers landed within one level of each other 77% to 81% of the time. In plain terms, two colleagues reading the same answer routinely disagreed about whether it was “Building” or “Strong,” the exact boundary where a developmental scale does its work. The report’s conclusion (The Geneva Learning Foundation, 2026) is worth quoting because institutions rarely say it aloud: the ratings should be read as “an encouragement signal, not a validated quality measure,” and “the more defensible outcome claim from the peer review instrument is not the numeric rating distribution but the qualitative content it elicited.”
An assessment specialist faces two options at this fork. The first is to chase psychometric respectability: more criteria, anchor training, double-marking, moderation panels. That road is well mapped and expensive, and its destination is a peer review process so constrained it stops producing the thing that made it valuable, the four-sentence review from Tanzania that changed the reviewer’s own practice. The second option follows Sadler’s (1989) insight that evaluative judgment is itself the thing being learned: stop asking peer review to be a measurement instrument and treat it as a learning mechanism whose by-product is a rich qualitative record. Applying criteria to a colleague’s work develops the reviewer’s judgment. Receiving three colleagues’ readings of your account shows you how your meaning lands. None of that requires the numbers to converge.
But the second option has a hidden dependency that the peer assessment literature rarely prices in. If the ratings are formative and the substance is the qualitative record, then the system’s integrity rests entirely on participation quality: real answers, real reviews, real responses. The reliability question does not disappear. It moves. It stops being “did two reviewers give the same score?” and becomes “did this learner actually do the activity in good faith?” And that question, unlike holistic quality judgment, is largely computable.
The hundred submissions one facilitator read by hand
Here is what answering the good-faith question cost before the Reviewer existed. In the pilot course, one TGLF facilitator personally read more than one hundred individual submissions to pass or fail each one as an integrity check. The course enrolled 1,110 people. TGLF’s flagship event series, Teach to Reach, drew 24,610 registrations in its most recent edition. Hand-reading does not survive contact with those numbers, and it spends the scarcest resource the foundation has, expert human attention, on the task where it adds least value: distinguishing “Ok ok ok ok” from a genuine contribution, a distinction a script can make.
The TGLF Reviewer is a deterministic screening pipeline, which means it is a set of explicit, inspectable rules rather than a machine-learning model: no AI reads learner work in its default configuration, and two runs on the same data produce identical results. It applies twelve rules across the three commitments. For authors: placeholder answers, minimal total effort, off-topic content, answers copied from one question into another, and formatting patterns typical of machine-generated text. For reviewers: the same comment sent to different colleagues, the same text pasted into every comment box, reviews too thin to act on, and praise so generic it could attach to any submission, which the system detects by measuring how much vocabulary the review shares with the work it claims to review. For reflections: missing, empty, or duplicated responses to the feedback received.
Three design choices matter more than the rules.
The thresholds come from the cohort, not from an administrator’s intuition. A review is “thin” relative to what colleagues in the same course actually wrote, and every threshold is reported with its provenance, so the team can see whether a norm came from configuration or from the cohort’s own distribution.
Every finding arrives with the learner’s own words as evidence. The system does not report that a learner “showed low engagement.” It reports that the learner’s shortest answer was eleven words, quotes it, states what the activity asked for, and assembles the whole thing into a paragraph ready to send. Two learners with the same finding receive identical wording, because the sentences come from templates rather than from a model’s paraphrase. In an integrity process this is a fairness property: the finding is inspectable, the evidence is the learner’s own text, and no algorithmic rhetoric stands between the two.
Nothing is suppressed silently. When a rule is switched off for a learner, because the course told participants a short answer was acceptable, or for the reason that follows, the exemption is written to a ledger with its justification. A human can always reconstruct why someone was not flagged.
The pilot showed why that ledger earns its keep. The screening initially flagged one reviewer for thin, generic, padded feedback: her three comment boxes contained the same short note. Read by a human, the note explained that she had been assigned a colleague’s work written in a language she could not read, and she asked the organizers to address the mismatch. She had violated nothing. The learning system had put her in a position where no honest review was possible, and the screening system was about to blame her for it. The fix was not to soften the rules but to teach the pipeline to see the situation: it now classifies the language of every submission and every review, records cross-language pairings as a design signal for the course team rather than a violation, and exempts the affected reviewer and author from every rule the pairing makes unfair. Twelve reviewers in the pilot cohort turned out to have been handed work they could not read. Under hand-reading at scale, most of those twelve would have been misjudged or missed. The machine did not replace human judgment in this story. It found the cases where human judgment was about to be wrong.
The work that requires judgment stays with people. Findings that touch on honesty, such as undeclared machine-generated text, carry a severity level that requires human review and is phrased as a question about how the work was developed, never as a verdict. A learning loop feeds the team’s decisions back into the system, which reports how often each rule turned out to be wrong and refuses to attach consequences to any finding a human has not confirmed. The facilitator who once read a hundred submissions now reads a one-page summary and the handful of cases that need reading. The recovered hours go where TGLF always wanted them: into supporting the learners doing the work rather than surveilling the ones who are not.
What a certificate can say with a straight face
TGLF delivers certificates, which forces a question most peer learning communities never have to answer in writing: what exactly is being attested?
The foundation’s certification specification for the pilot course names the professional behaviors participants practiced, reflective observation, pattern recognition, systemic analysis, action planning, and constructive solidarity in giving and integrating structured feedback, and then states plainly: “These are areas of deliberate practice, not certified competencies. The primer verifies engagement with a professional analysis process. It does not assess or certify clinical or programmatic competence.”
It would be easy to read sentences like those as defensive hedging. They are the opposite. In TGLF’s philosophy of education, being crystal clear about what the organization can assess, how, and why is an ethical requirement, part of the same epistemic discipline that makes the foundation report its own reliability statistics and label its own outcome claims as self-reported. The contrast is with formal education’s habitual register, which routinely makes totalizing claims about effectiveness and outcomes that its own evidence base does not support. TGLF certifies what its platform and team actually verify: that a named health worker completed a disciplined cycle of analysis, peer exchange, and planning about their own practice, documented in their own words. What that cycle changed is tracked by a different instrument entirely, the action plans with named beneficiaries and dates, and the follow-up accounts, in the pilot, 45.9% of analyzed projects committed to a specific, dated practice change, and each such claim is published with its limitation attached: self-reported, not independently verified.
The evidence on formal credentials makes this precision look less like modesty and more like advantage. Fifty years of personnel selection research finds that years of education predict job performance with a validity of 0.10, where validity is the correlation between the credential and later performance on a scale from 0 (no relationship) to 1 (perfect prediction), among the weakest signals ever measured, and that college grades add almost nothing to prediction beyond general cognitive ability (Schmidt & Oh, 2016). The correlation between grades and job performance, modest to begin with, decays to 0.05 within six years of graduation (Roth et al., 1996). Labor economists attribute roughly 30% of the earnings return to completing a degree to the credential itself rather than to anything learned, the sheepskin effect (Ferrer & Riddell, 2002). Employers report the same disconnect in their own dialect: job postings demand degrees that two thirds of the people already doing those jobs successfully do not hold (Fuller & Raman, 2017), and when large employers publicly dropped degree requirements, the actual change in who they hired amounted to fewer than one hire in 700 (Sigelman et al., 2024). In the AAC&U’s employer survey, the gap between the importance employers assign to oral communication and their confidence that graduates are well prepared in it ran to 30 percentage points (Finley, 2023). Closest to TGLF’s own field, the current Cochrane review of continuing education for health professionals, which pooled 215 studies of more than 28,000 practitioners, found that educational meetings only slightly improved compliance with desired clinical practice (median 4 percentage points) and only slightly improved patient outcomes (Forsetlund et al., 2021), with the specific question of whether interactive meetings outperform lecture-based ones downgraded to very low certainty. Which is to say: the dominant credentialed format in health worker education produces the same small, heterogeneous effects TGLF’s own model reports for a fraction of the cost.
None of this proves peer learning certifies competence either. What it dismantles is the assumption that formal assessment sets a gold standard against which TGLF’s precision needed excusing. A certificate that documents verified engagement, what the participant actually did, with whom, and with what stated plan, makes a claim its issuer can stand behind. And this is where the Reviewer changes the certification calculus: the specification’s phrase “genuine, substantive engagement” is only worth printing if genuineness and substance are checked. The screening system is what makes that check possible for a thousand learners at a time, which converts the certificate’s honesty from aspiration into audit trail.
Why automation made the feedback more personal, not less
Automating integrity screening made the feedback each learner receives more individual, not less.
Consider what the system actually sends. A learner who reviewed diligently but never responded to the colleague who reviewed her receives a message quoting her own record, thirteen of fifteen activities completed, the message notes, calling her diligent, naming exactly what is missing and why it matters to the colleague still waiting. A learner whose three answers total 92 words receives his own words back, alongside what colleagues in the same course typically wrote. The message about machine-formatted text asks how the work was developed and reminds the learner that the Honor Code asks for disclosure, not abstinence: declared AI use is context, never a violation. Every message states that nothing has been decided, and every message ends with a named person who will read the reply.
Each choice tracks what the integrity literature has found about adult professional learners. Clear expectations stated before the work, not after, remain the strongest single lever, which is why the Honor Code and the AI rules are course elements a learner must review before anything is submitted. Sanctions imposed without dialogue corrode the relationship the sanction is meant to protect, so the first contact is always a warning that quotes evidence and offers a right of reply. And treating AI use as a disclosure norm rather than a prohibition reflects both the pilot’s reality, 64.4% of respondents reported using generative AI in developing their answers, and Sadki’s own analysis of what he calls the transparency paradox: professionals inside punitive accountability structures are penalized for disclosing AI use and penalized for concealing it, so a learning organization’s obligation is to make transparency safe rather than to demand it while punishing it (Sadki, 2026).
The fairness properties run deeper than the wording. Peer review in these courses is anonymous, so no message a learner receives ever names another learner, even though the staff-facing evidence preserves full names for verification. The same finding always produces the same sentences. An unresolved identity, two learners sharing a name, is a hard stop: nobody is emailed on a guess.
Why a theory of learning needs an enforcement layer
Read through the lens of Structured Peer Praxis, the Reviewer is the enforcement layer for the theory’s central mechanism rather than an assessment add-on.
Sadki’s framework holds that learning at scale works through what he calls selected self-organization: just enough structure, rubrics, deadlines, matched criteria, reciprocal obligations, to let a network of practitioners produce knowledge together without an expert bottleneck (Sadki, 2026). The framework’s own diagnostic rubric interrogates any learning system on three counts: whether accountability runs horizontally between peers or vertically to a certifying authority, whether the process closes the praxis loop from reflection back into tracked action, and whether scale is treated as a diversity asset or a quality risk.
Every one of those commitments creates a specific vulnerability, and the Reviewer’s rules map onto them nearly one to one. Horizontal accountability functions only if reviews are real, so the system measures whether feedback engages the specific work it claims to review. Praxis closure functions only if the reflection step happens, so the system flags the learner who received a colleague’s labor and never closed the loop. Scale is an asset only if the network’s exchanges carry signal, so the system screens every exchange rather than sampling. Even the theory’s decolonial commitments surface concretely: the cross-language detection exists because a rule built on a monolingual assumption was about to penalize the multilingual reality of a global cohort, and the fix treats the mismatch as a course design signal rather than a learner failure.
The deeper point belongs to the assessment literature as much as to Sadki. A long line of work, from Sadler (1989) on formative assessment through Topping (1998) and Nicol and Macfarlane-Dick (2006) on peer feedback, argues that evaluative judgment is not a measurement problem but a learning outcome. TGLF’s data supports an operational version of that claim. The physician in Tanzania did not learn despite having to review three colleagues. The review was where her practice change was born, in her own sentence, addressed to someone else, which she then went and did. A system that protects the conditions under which that sentence gets written, by making empty reviews visible, by refusing to let a language mismatch masquerade as laziness, by reserving human judgment for the cases that need it, is using assessment exactly as the design intends: as scaffolding for the feedback that carries the learning.
What the machine cannot read
The Reviewer measures participation integrity, not depth, accuracy, or kindness. A superficial review of a specific answer passes. A factually wrong clinical claim passes. Every statement about local conditions and practice change remains self-reported, and the reports say so in their own limitations sections. Inter-reviewer reliability remains slight, and the pilot’s most actionable recommendation on that front, showing reviewers brief calibration examples before their first rating, is a human process improvement no pipeline can substitute. The screening system clears the routine so that human attention can concentrate where judgment, care, and expertise are irreplaceable: on the learners doing the work, and on the small number of hard cases where the evidence needs a person to read it.
Machines verify commitments and humans support learning: that division is what the whole exercise buys. The alternative was never a world without verification. It was a facilitator reading a hundred submissions by hand, a flagging process that would have sanctioned a reviewer who deserved an apology, and a certificate whose central claim, genuine engagement, nobody could actually check.
References
Ferrer, A. M., & Riddell, W. C. (2002). The role of credentials in the Canadian labour market. Canadian Journal of Economics, 35(4), 879-905. https://econ.queensu.ca/pub/jdi/deutsch/edu_conf/Ferrer.pdf
Finley, A. (2023). The career-ready graduate: What employers say about the difference college makes. American Association of Colleges and Universities. https://www.aacu.org/research/the-career-ready-graduate-what-employers-say-about-the-difference-college-makes
Forsetlund, L., O’Brien, M. A., Forsén, L., Mwai, L., Reinar, L. M., Okwen, M. P., Horsley, T., & Rose, C. J. (2021). Continuing education meetings and workshops: Effects on professional practice and healthcare outcomes. Cochrane Database of Systematic Reviews, (9), CD003030. https://doi.org/10.1002/14651858.CD003030.pub3
Fuller, J. B., & Raman, M. (2017). Dismissed by degrees: How degree inflation is undermining U.S. competitiveness and hurting America’s middle class. Harvard Business School. https://www.hbs.edu/ris/Publication%20Files/dismissed-by-degrees_707b3f0e-a772-40b7-8f77-aed4a16016cc.pdf
Nicol, D. J., & Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. Studies in Higher Education, 31(2), 199-218. https://doi.org/10.1080/03075070600572090
Roth, P. L., BeVier, C. A., Switzer, F. S., III, & Schippmann, J. S. (1996). Meta-analyzing the relationship between grades and job performance. Journal of Applied Psychology, 81(5), 548-556. https://doi.org/10.1037/0021-9010.81.5.548
Sadki, R. (2026). Structured peer praxis: A theory of networked knowledge production and local action at scale [Unpublished manuscript]. The Geneva Learning Foundation.
Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119-144. https://doi.org/10.1007/BF00117714
Schmidt, F. L., & Oh, I.-S. (2016). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 100 years of research findings [Working paper]. https://home.ubalt.edu/tmitch/645/session%204/Schmidt%20%26%20Oh%20validity%20and%20util%20100%20yrs%20of%20research%20Wk%20PPR%202016.pdf
Sigelman, M., Fuller, J. B., & Martin, A. (2024). Skills-based hiring: The long road from pronouncements to practice. The Burning Glass Institute and Harvard Business School. https://www.burningglassinstitute.org/research/skills-based-hiring-2024
The Geneva Learning Foundation. (2026). Course evaluation and field intelligence reports, pilot peer learning cohort [Internal reports].
Topping, K. (1998). Peer assessment between students in colleges and universities. Review of Educational Research, 68(3), 249-276. https://doi.org/10.3102/00346543068003249
