How it works

Most apps grade. Ours diagnoses.

Under the hood is a real student-modeling engine — the same kind of math the SAT and NWEA MAP run on — built specifically for the way your child actually answers questions.

A different way to read an answer

Fast or slow, right or wrong — four very different kids

A single score hides four different stories. We read all four separately.

Fast, correct

Genuine mastery

The skill is automatic — they're not thinking hard about it anymore.

Raise the difficulty. Re-check it in a few weeks, then leave it be.
Slow, correct

Knows it, hasn't automated it

Invisible to grade-book apps — and a real risk on a timed exam.

Fluency drills at the same difficulty, with gentle time pressure.
Fast, wrong

A confident misconception

Not a guess — they're fluently applying the wrong rule.

Re-teach that exact misconception — more of the same question only reinforces it.
Slow, wrong

A genuine knowledge gap

They searched for a method and came up empty.

Step the difficulty down, check the prerequisite skill, offer a hint.

Only one of these four is really addressed by "assign more questions on the weak topic" — which is what most apps do to all four.

Our unfair advantage

Wrong answers with a diagnosis built in

Because we write our own questions, every wrong option is tagged, at creation, with the exact misconception that produces it.

⅓ + ⅙ = ?

B 2⁄9
🔎 added the tops and bottoms separately

⅓ + ⅙ = ?

D 1⁄18
🔎 multiplied numerators AND denominators, like multiplying fractions

Same right answer, two students, two completely different next steps — because we know why each of them was wrong, not just that they were.

Not just more drilling

Missions, not worksheets

Learning-science research calls repetitive drilling on a weak topic "wheel-spinning" — kids who grind the same skill without real progress usually don't improve with more of the same. So every mission mixes four kinds of practice instead.

Growth. The weakest 1–2 topics, one notch below where they last struggled — never a wall of "you failed this."
Fluency. Topics they get right but slowly — the same difficulty, a gentle nudge on pace.
Stretch. Their strongest topic, one tier harder — the actual unlock target, so leveling up feels earned.
Keep-sharp. Anything mastered a while ago gets a quiet re-check — so it doesn't quietly fade.

Every mission blends all four — weighted toward growth, never all one kind.

Behind every question

It takes a whole panel, not one model

A roster of specialists, each with one job, and a few of them deliberately trying to catch each other out. Teaching a kid math well was never going to be simple, so we stopped building as if it were.

🧮

The Solvers

Two different AI families work every problem blind — separately, with no idea what the "right" answer is supposed to be.

✍️

The Editors

A separate panel checks the wording, the reading level, and whether the hint actually helps — not just whether the math is right.

🎭

The Impersonators

Told to believe one exact wrong idea and answer as that student would — proof a question actually catches the mistake it's built to catch.

🗣️

The Plain-Talkers

Rewrite a hard question in simpler words and re-test it. If kids would answer differently, that's a wording problem, not a math one — fixed before it ships.

📚

The Historian

Remembers every mistake this bank has ever made and feeds the lesson back into how the next question gets written.

📊

The Calibrator

Checks the model's own difficulty judgment against real results from real kids on real published exams — never just its own opinion.

One of the Impersonators, on the job

"You believe fractions add straight across — tops with tops, bottoms with bottoms. Answer ⅓ + ⅙ as that student would."

The AI answers
2⁄9
— the exact wrong option the question was built to catch it with.

None of this is improvised. It's grounded in decades of published testing and learning-science research — the same item-response math the SAT runs on, the hint-partial-credit studies from Worcester Polytechnic's ASSISTments project, and the "wheel-spinning" finding that more of the same question doesn't help a stuck kid, which is why ours never just does that.

Under the hood

Confidence grows with every test

Instead of one score, the model holds a range — how sure it is about what your child knows. Twenty questions in, that range narrows fast. It never mistakes one lucky guess (or one bad day) for the whole story.

After the first test
Wide range — still learning who they are.
After 20 questions
Tight, confident — the plan gets sharp.
See it diagnose a real question →