Backed by research
The science behind Teacher Supreme
Third-period algebra. You have just asked how anybody got x = 4 out of 3x + 7 = 19. A student who never shows work says they took the 7 off both sides first, because the 7 was not attached to the x the way the 3 was, and only then divided. That is the real reasoning, said out loud, by a student who normally writes down an answer and nothing else. You have a moment to decide what comes next, before it closes and the room moves on.
Most of us say “Good job!” — or “Nice work,” “Keep it up,” “I like that,” “Way to go” — and move to the next desk. Those five phrases cover most of the praise spoken in most classrooms, and there is nothing wrong with any of them. They are fast, they are warm, and they keep the room moving.
The trouble is what the student is left holding. Stop that same student in the hallway at lunch and ask what they did well in algebra today. The answer is “I don’t know — my teacher said good job.” The compliment survives; the thing that earned it does not. And because the student cannot name what they did, they cannot choose to do it again on tonight’s homework, or on Friday’s quiz when the equation has fractions in it, or in April on the state test. Nothing about tomorrow changed.
The other option takes no longer. Say the student’s name, then say what you saw: “Marcus, you explained every step instead of jumping to the answer. That is showing your reasoning — show your steps again whenever a problem gets hard.” Ask Marcus at lunch now and the answer is “I showed my steps — I said why I subtracted before I divided.” They can name the move, so they can run it again on purpose on Friday, when the equation has fractions in it and the answer is no longer obvious — which is the only thing that was ever going to raise the grade.
Same breath. Same teacher. Same student. The difference is in the wording, not in effort, caring, or years on the job — and it has been measured. In the study at the bottom of this page, one line of praise decided whether ten-year-olds picked the harder problem or the easier one, and whether they got better or worse after a setback.
Nobody hands a new teacher the second sentence. There are 24 things worth noticing in a classroom and three ranked moves for each — 99 sentences — and no one can hold 99 sentences in mind while also teaching a lesson, watching the door, and tracking who has not spoken today. Teacher Supreme keeps them and shows you the right one while the moment is still open, with the study it came from attached. Here is that evidence, with sources you can check.
Where a claim is our own reasoning rather than something a study measured, the card says so plainly. We would rather be checkable than impressive.
How to read these numbers
Researchers measure how much something helps students using an “effect size” — a single number on a shared scale, so any two teaching ideas can be compared fairly.
Two different kinds of number appear below, and they do not mean the same thing. When researchers randomly pick some classes to try something new and compare them to classes that did not, 0.20 or higher counts as large for a real school — that bar comes from Kraft (2020), who lined up 1,942 results from 747 randomized school trials measured on standardized tests and found the middle one sat at 0.10.
The other kind — Hattie’s figures, like the 0.52 for teacher-student relationships — mostly comes from studies that measured whether two things travel together, not whether one caused the other. Those numbers run larger for that reason, and Kraft’s own first rule is that they cannot be lined up against his bar. So we do not do it: a Hattie figure on this page tells you a practice keeps good company with learning, not that using it buys you that much gain.
The practices TeacherSupreme helps you notice and coach — explaining reasoning, summarizing, where-to-next feedback — are the practices the 24 indicators name, and every figure below is about the practice. Teacher Supreme itself has not been studied. No effect size on this page is a measurement of this product.
Why this exists
The argument, and the studies it rests on. Where a finding is contested or narrower than it is usually quoted, that is said here rather than left for you to discover.
Teachers already know what works. That was never the problem.
When Mary Budd Rowe taught teachers to pause about three seconds after asking a question instead of the one second most of us manage, student answers got three to seven times longer and students who had stopped speaking started again. Learning that one pause took six to twelve hours of practice. Seven or eight teachers in ten reached it. And most drifted back to their old pace by the third or fourth week unless someone was there to talk it through with them. One pause. That is what a new habit costs a teacher who is already running.
Coaching works — and it thins out the moment it scales.
The strongest answer we have is putting a coach beside a teacher. A review of sixty studies that could support cause found coaching moved classroom instruction by about 0.49 of a standard deviation and student achievement by about 0.18; after the authors’ own correction for publication bias, roughly 0.34 and 0.14. The number that matters here is the next one. Programs reaching a hundred teachers or more did markedly worse than small ones — about 0.34 against 0.63 on instruction, and 0.10 against 0.28 on achievement. That comparison is across different programs rather than one program grown large, so read it as a pattern and not a law. But the pattern is the whole problem: you cannot put a person next to every teacher every day.
Kraft, M. A., Blazar, D. & Hogan, D. (2018), Review of Educational Research 88(4), 547–588
What a teacher says in that second is not a small difference.
A 2020 re-analysis pulled 435 primary studies out of 32 existing reviews — 994 measurements, more than sixty thousand participants — and found feedback averaging about 0.55, or 0.48 once outliers were removed. The average is not the finding. Feedback that carries real information about the task, and about how a student can steer their own learning, reaches 0.99. Feedback that only rewards or punishes — a mark, a “good job” — sits at 0.24. Simply supplying the right answer reaches 0.46. Two limits we will state rather than let you find: the 0.99 rests on 42 of those 994 measurements and compares different studies rather than running them head to head, and the authors’ own test flags publication bias in the journal articles. Treat these as upper bounds. The gap between them is still the difference between two sentences a teacher could say in the same second.
The technology to be there in that second did not exist until now.
This is the part that is ours, not a citation. A teacher does not need another course, another binder, another August workshop. They need the right words in their hand while the student is still standing there — and until recently nothing could do that. It could not hear a name in the middle of a room, hold what happened to that student last Tuesday, and hand back a sentence in the time it takes to turn around. That is new. Teacher Supreme is what happens when someone who spent thirty-five years in those classrooms gets to build with it.
Not a research finding — the reason this was built. James Morris, San Diego Unified, 1991–2026
What we do not claim
No trial has tested whether Teacher Supreme raises achievement. We have not run one, and we will not imply otherwise.
- The research above is about feedback, about wait time, and about coaching — not about this product. It is the reason we built what we built. It is not a measurement of it.
- The link between warmer teacher–student relationships and higher achievement is real but small, and it is a correlation: roughly r = .35 with engagement and r = .17 with achievement across 189 studies and about a quarter of a million students. Conflict runs almost as strongly the other way. The authors say plainly that their data do not permit conclusions about cause, and they warn that part of the engagement figure may come from the same person rating both things.
- One well-known finding we rely on inside the product — that praising the process beats praising the person — has a replication that failed on most of its measures, though it reproduced the original’s main outcome and the original authors dispute the rest. We name that on the science page rather than leave you to find it.
- Anyone who tells you their software raises test scores should be asked for the trial. So should we, when there is one.
A human mind holds about four things at once. A class has thirty.
Miller’s famous “seven, plus or minus two” was never a hard limit — Miller offered it as a rough estimate. Reviewing decades of experiments, Cowan (2001) put the real ceiling lower: about four chunks, roughly three to five, measured when a person cannot rehearse the items or group them into larger ones. Cognitive-load theory supplies the consequence for learning: working memory is the bottleneck when someone is taking in new material, and capacity spent on information extraneous to the task — noise, clutter, things to keep track of — is capacity no longer available for the learning itself (Paas & van Merriënboer, 2020). No study measures a teacher trying to carry a full classroom’s allergies, IEP pages, missing work and who has not spoken in a week while also teaching the lesson; that extension is ours, not a finding. But the direction is not in question. Memory is the scarce resource in a classroom, and every bit of it spent on retrieval is not spent on the student standing in front of you.
In the app: This is the whole design. Teacher Supreme takes the remembering and leaves the teaching alone — the teacher does the noticing and the judging, and the app carries the allergy, the accommodation, the promise made on Tuesday, and who has not been reached this week.
Source: Cowan, N. (2001), “The magical number 4 in short-term memory,” Behavioral and Brain Sciences 24(1), 87–114; Paas & van Merriënboer (2020), Current Directions in Psychological Science 29(4), 394–398 →Students who feel known by their teacher tend to do better — one of the steadier patterns in the research.
Hattie’s current synthesis puts the association between a strong teacher–student relationship and achievement at about 0.52 — one of the larger figures in those rankings. That number comes from research measuring whether the two travel together, not from trials that assigned some students a better relationship, so here is exactly what it supports: students who feel known by their teacher tend to do better, and independent meta-analyses point the same way. It is not a gain a teacher is promised, and 0.52 cannot be lined up against the 0.20 bar in the box above — that bar is for randomized trials, and this is not one.
In the app: Teacher Supreme makes those relationships visible — the parent brief surfaces the specific, documented interactions that tell a student they are seen and remembered.
Source: Hattie, Visible Learning (current synthesis) — correlational research, not a randomized trial →Positive interactions should substantially outnumber corrections — the exact ratio is not the point.
Two classroom studies raised teachers’ rate of praise over reprimands and watched what happened: students were on task more and disrupted less (Cook et al., 2017 — a small study, six teachers; Caldarella et al., 2023, a larger one). Neither found a target number to hit. The popular 4:1 and 5:1 “magic ratios” were reviewed by Sabey, Charlton & Charlton (2019), who found no empirical support for any specific ratio while still backing the direction: stay well positive. So TeacherSupreme shows you your own balance of positive to corrective interactions — your pattern, never a target to hit.
In the app: Teacher Supreme logs every interaction by type and shows each teacher their own positive-to-correction balance on the daily recap. When one student’s positives slip behind their corrections, it names that student and offers a specific positive to use — turning an invisible habit into something a teacher can see and act on.
Source: Cook et al. (2017) & Caldarella et al. (2023), Journal of Positive Behavior Interventions; Sabey, Charlton & Charlton (2019), Journal of Emotional and Behavioral Disorders →Warmer teacher-student relationships go with higher achievement — modestly, and nobody can say which way it runs.
Göktaş and Kaya (2023) pooled 17 earlier meta-analyses — a meta-analysis of meta-analyses — and found a modest positive correlation between positive teacher–student relationships and academic achievement (r = .21), with a smaller negative one where the relationship is conflictual (r = −.14). It is correlational, so it cannot tell you which way the arrow points: a student who is already doing well may also draw warmer treatment. What it does say is that across hundreds of studies the two travel together.
In the app: Teacher Supreme holds the details a teacher would otherwise have to carry in their head — the peanut allergy, the extended-time accommodation, the brother in the hospital, the promise made on Tuesday — so they are in front of you when that student is. The research above is about relationships between teachers and students. It is not a study of this app.
Source: Göktaş, E., & Kaya, M. (2023), “The Effects of Teacher Relationships on Student Academic Achievement: A Second Order Meta-Analysis,” Participatory Educational Research 10(1), 275–289 →Feedback is among the highest-impact things a teacher can do — especially when it says what to do next.
A 2020 meta-analysis of 435 studies and more than 61,000 participants (Wisniewski, Zierer & Hattie) — far larger than the 131-study review that had been the benchmark — found feedback overall averages an effect size of 0.48, above the 0.20 that counts as large in randomized school trials (Kraft, 2020), though his bar comes from trials measured on standardized tests and this analysis pools a wider mix. The gap inside that average is the real story: feedback that tells a student what was right or wrong AND adds information about the process and about managing their own learning reaches 0.99, while feedback that only rewards or reprimands, carrying almost nothing about the task, sits at 0.24. Formative assessment built the same way shows similarly strong effects (Black & Wiliam, 1998).
In the app: Snap-Grade, included in Plus, reads a worksheet on the spot, marks it, names the precise item to revisit, and whispers the next move — the high-information feedback the research says matters most, turning grading from after-hours paperwork into in-the-moment coaching.
Source: Wisniewski, Zierer & Hattie (2020), “The Power of Feedback Revisited,” Frontiers in Psychology 10:3087; Black & Wiliam (1998), Phi Delta Kappan 80(2) →Schools are not under-spending on teacher development. They are spending it in August.
TNTP’s The Mirage studied three large districts and one charter network and put professional-development spending at roughly $18,000 per teacher per year — about 19 school days, a tenth of the year, and an estimated $8 billion annually across the fifty largest districts. It found no link between those programs and improved classroom performance. Those are 2015 figures from four systems, not a national census, and they are the best-known numbers in the field. The failure is not effort or money; it is timing. A workshop in August is being asked to change a decision a teacher makes in two seconds in November, with no reminder in between.
In the app: Teacher Supreme puts the coaching inside the two seconds. The teacher says what they just saw; the words come back while the student is still standing there. It is professional development delivered at 10:42 in room 214 rather than in a lecture hall before the students arrive.
Source: TNTP (2015), The Mirage: Confronting the Hard Truth About Our Quest for Teacher Development →Under ESSA we qualify at Level 4 — and we will not claim the three above it.
ESSA defines four evidence levels. Level 1 needs an experimental study, Level 2 a quasi-experimental one, Level 3 a correlational study with controls. Teacher Supreme has none of those, and says so. Level 4 — “demonstrates a rationale based on high-quality research findings… that such activity, strategy or intervention is likely to improve student outcomes” — is the tier written for exactly this case, and it carries a real obligation: explore the research, state the mechanism in a logic model, and plan how you will evaluate it. All three are documented in the evidence brief below. Because it is professional learning rather than curriculum, Title II-A is the usual funding home; confirm allowability with your own federal programs office.
In the app: The evaluation is built from numbers the app already produces and the teacher already owns: their positive-to-corrective balance, and the share of their roster reached in a week. A school can baseline both in two weeks, compare against a matched group, and read three outcomes at 30, 60 and 90 days.
Source: ESSA §8101(21)(A) evidence definitions; U.S. Department of Education, Using Evidence to Strengthen Education Investments →One line of praise, said once, changed what ten-year-olds chose to do next.
Mueller and Dweck (1998) ran six studies with 412 fifth graders in all, most of them working through sets of Raven’s matrices puzzles, and told them all the same thing about their score. Then one group heard six more words — “You must be smart at these problems” — and the other heard “You must have worked hard at these problems.” That was the entire difference. Offered a choice next, 67% of the students praised for being smart picked a task that would let them look good rather than the one they would learn more from. After a round of deliberate failure, those same students quit sooner, enjoyed the work less, said the problem was that they were not smart enough — and then scored below their own opening round. The students praised for their effort stayed with it longer, put the failure down to effort, and finished higher than they started. In one of the studies, 38% of the students praised for intelligence misreported their score to children they did not know, against 13% of those praised for the work they did and 14% of a control group. The samples were not a narrow slice: three of the six studies were majority African American and Hispanic, and a fourth was evenly split. Person praise hands a student something to protect. Process praise hands a student something to repeat.
In the app: This is the rule the entire coaching library is written to, and it is enforced in the code, not just intended. Every Coaching Moment names what the student actually did, names the strategy, then says when to use it again, addressed to that student by name — “Marcus, you explained every step instead of jumping to the answer. That is showing your reasoning — show your steps again whenever a problem gets hard.” Twenty-four indicators, ninety-nine ranked moves, each one carrying the name of the researcher it came from. In class no one reads any of this: the teacher says what they see, and the words arrive on the student’s card while the moment is still happening.
Source: Mueller, C. M., & Dweck, C. S. (1998), Journal of Personality and Social Psychology 75(1), 33–52 →For principals and district leaders
One page for the budget meeting \u2014 the research, the logic model, and a 90-day evaluation you run yourself.
The evidence brief carries everything on this page plus what a school actually needs to act: why the timing of coaching matters more than its quantity, the ESSA Level 4 rationale written out in full, a logic model you can paste into a plan, and an evaluation built from numbers the app already produces. It also states plainly what we are not claiming \u2014 there is no study of Teacher Supreme, and no figure in it is a gain a teacher is promised.
The ask is small: five to ten teachers, ninety days, three outcomes you choose. Every teacher starts on a 28-day free trial and no card is needed to begin.
Download the evidence brief (PDF) →What that looks like at 10:14 on a Tuesday
The difference is not effort, or caring, or years on the job. It is the sentence.
Three real interactions, in three different rooms. Every line on the right comes out of Teacher Supreme word for word. Each one does the same three things the research asks for: it names what the student actually did, it names the strategy so the student can find it again, and it says when to use it next.
English 9
A student cuts a three-page chapter down to the one thing that changed between the two characters, and leaves out the weather, the bus ride, and the dog.
What most of us say
“Nice work!”
What Teacher Supreme shows you
“You kept the big idea and cut the small details. That is summarizing — cut the small details again the next time a reading is too long to hold in your head.”
Biology
A student’s onion-root-tip slide keeps drifting out of focus and every cell looks like a smudge. Instead of copying the lab partner’s drawing, they re-mount the slide, re-focus, and find a cell caught in anaphase.
What most of us say
“You’re so smart.”
What Teacher Supreme shows you
“You stayed with the task after it got hard instead of setting it down. That is staying with it — stay with the task again the next time the work starts to drag.”
U.S. History
A student argues the tariff hurt farmers, gets pushed on it, and goes back to the primary source to find the line about grain prices rather than restating the claim louder.
What most of us say
“Great answer!”
What Teacher Supreme shows you
“You went back to the page and found proof. That is backing your position with evidence. When someone asks ‘how do you know?’ — go back to the reading and find the evidence that supports your claim.”
Nobody can hold ninety-nine of these in their head while teaching. That is the point — you do not have to. You say what you see, and the words come to you.
The research behind all twenty-four indicators
Thirteen have an effect size. The other eleven have a literature. Both are research.
An effect size is one narrow product of one kind of study — a meta-analysis of how strongly something relates to an outcome. Most of what a teacher watches for has never been reduced to one, and never needed to be: it has been defined, measured and studied instead. The thirteen indicators that carry a number are above, with their figures and their limits. Here is the work behind the other eleven — the off-task and disengagement behaviours — and each one names its source on its own card inside the app, too.
Task avoidance
Teacher-rated task-avoidant behaviour predicts later literacy skills across time.
Liao, Georgiou, Zhang & Nurmi (2013), Learning and Individual Differences 25
Low / no effort
The engagement literature calls this disaffection: “passivity, lack of initiation, lack of effort, and giving up.” Its validated teacher item is “when faced with a difficult assignment, this student doesn’t even try.”
Skinner, Kindermann & Furrer (2009), Educational and Psychological Measurement 69(3)
Rushed work
Finishing fast and finishing well trade against each other — “an inescapable property of choice behaviour.”
Heitz (2014), Frontiers in Neuroscience 8:150
IDK response
Children who attribute failure to something fixed in themselves stop trying on problems they could otherwise solve.
Dweck & Reppucci (1973), Journal of Personality and Social Psychology 25(1)
Intellectual guessing
Answering too fast to have considered the question is a named, measured behaviour — rapid-guessing, as opposed to solution behaviour.
Wise & Kong (2005), Applied Measurement in Education 18(2)
Delayed response
Slow to start is its own measure — latency to initiate an academic task — and it moves when the request is made differently. A single-subject study, so read it as the source of the construct.
Wehby & Hollahan (2000), Journal of Applied Behavior Analysis 33(2)
Disrupts others
Disruption for a class reaction is one of the two dimensions a validated teacher instrument measures at school.
Rayfield, Eyberg & Foote (1998), Educational and Psychological Measurement 58(1)
Side conversation
“Off-task verbal” is a coded category in the direct-observation system school psychologists use: verbal behaviour that interferes with classroom functioning.
Behavior Observation of Students in Schools; validity study, Alperin et al. (2023), Psychology in the Schools 60
Device misuse
Non-academic internet use was common among students who brought laptops to class and was inversely related to their class performance — measured by logging real use, not by asking.
Ravizza, Uitvlugt & Fenn (2017), Psychological Science 28(2)
Copying work
Copying is the most-studied form of academic dishonesty, with decades of measurement behind it.
Whitley (1998), Research in Higher Education 39(3); school-level items in Šorgo et al. (2015)
Wandering / daydreaming
Zoned out is a researched state, not a judgement: mind-wandering, defined for participants as being lost in thoughts unrelated to the task. Out of seat is the other half, coded as off-task motor.
Unsworth & McMillan (2017), Cognitive Research 2:32
Defiance / back talk
Noncompliance has a formal definition: when a child, actively or passively but purposefully, does not do what an adult has asked. That review is about children generally, so the school-specific measure is the oppositional dimension of the teacher instrument above.
Kalb & Loeber (2003), Pediatrics 111(3)
Twelve entries for eleven indicators, because Evidence-Based Defense sits with them: argumentation from evidence has its own literature (Osborne, Erduran & Simon, 2004, Journal of Research in Science Teaching 41(10), 994–1020) and no single published effect size. Citing a source is not a claim that its authors endorse this product, and an indicator being research-grounded is not a claim that Teacher Supreme has been validated as a diagnostic instrument. It has not.
The microphone, kept honest
A classroom is full of voices. Here is how only yours counts.
Listening is teacher-activated only, and two physical checks sit behind it. The first asks whether the voice was close to the phone. The second asks whether its pitch sat inside a band you enrolled by reading three sentences once — what is kept is four numbers on your own device, no recording, deleted in one tap. A child’s speaking pitch sits measurably above an adult’s, which is why plain arithmetic is enough. Both checks claim only that a voice was consistent with yours — never identity — and both fail open: with nothing to go on, the card comes through. The checks can only remove a wrong trigger, never one of yours.
When it mishears a name, it resolves against your roster by sound — the students actually in the room — and when more than one student could fit, it asks instead of choosing. Where live listening does not exist at all (no iPhone browser allows it), the same logging runs from a held button with the words made on the phone itself: audio never leaves the device. Names spoken aloud in a classroom are exactly the data that should never travel, so the architecture is built so they cannot.
This is professional development that arrives in the moment it is useful, instead of in a workshop in August. Every move names its source, so you always know why you are saying it.
Start your 28-day free trial