Project review pack v0.1 · draft

Every learner keeps making the same mistakes.
Nobody can tell them which ones.

This is the working design for an English-learning system built around a single job: find out exactly what a learner is weak at, practise that, and prove the weakness is gone. It is not built yet. This document exists so that teachers can tell me where the thinking is wrong.

Try it — what the system actually does 1 / 3

Answer these three as your weakest student would, not as you would.

75grammar units
1,100questions needed
82%day-one accuracy
0learners tested

The idea

What this project is trying to do

And, more importantly, what it is deliberately not trying to do.

The problem, in one paragraph

An Iranian adult who has studied English for six years is not short of material. They have books, apps, YouTube, a teacher, and now ChatGPT. What they do not have is an answer to a simple question: which specific things am I still getting wrong, and what should I do about them today? So they study broadly, revise things they already know, and three years later they are still dropping articles and still using the present perfect with yesterday.

The core idea

Stop teaching English in general. Continuously build a picture of what one particular learner cannot do, aim the practice there, and keep re-testing until the mistake genuinely stops.

How it would work

1

A short test places the learner

Around 20 adaptive questions. This gives a reliable CEFR band and three to five areas worth checking. It does not, and cannot, produce a full diagnosis.

2

Every answer becomes evidence

Each question is tagged with the exact grammar point it tests, and each wrong option is tagged with the misconception behind it. A wrong answer names the error rather than just recording a failure.

3

Repeated evidence becomes a diagnosis

One mistake is a slip. Several mistakes on the same point, across different question types, is a weakness. The system waits until it is confident before telling the learner anything.

4

Practice aims at the weakness, then comes back

Short lessons and targeted exercises, then the same grammar point returns days and weeks later in harder formats — until the learner can produce it, not just recognise it.

Who it is for, first

Iranian adults roughly 20–35, around B1, who have studied for years and still repeat basic mistakes. Not children, not beginners, not IELTS candidates, not every learner in the world. One group first.

What is deliberately left out of the first version

No IELTS or TOEFL simulation. No speaking assessment. No teacher marketplace. No mobile apps. No vocabulary system at first, because spaced repetition already solves that well and building it would prove nothing. The first version tests one thing only: whether a system can find a learner's real weaknesses, and whether anyone cares.

Honest status

Nothing here has been built or tested with a real learner. Every number in these documents is a reasonable starting guess, chosen so it can be corrected later against real data. Treat all of it as claims to be attacked, not conclusions.

What I'm asking you

Seven questions for a teacher

You teach English. I do not, and this design makes claims about how learning works that I am not qualified to make alone.

Please do not be encouraging

If you tell me this is a nice idea, you have given me nothing. Reply with the question number and the word "no" wherever it applies. Disagreement is the entire purpose of sending you this.

1The answer I want most§2.5

Which grammar points do your Iranian students get wrong persistently — where you taught the rule properly, they understood it in class, and three months later they are still making the same mistake? A list of five is worth more to me than everything else in this document combined.

2Granularity§2.2 · §2.3

The design breaks grammar into 60–90 small units for the A2–B2 range. Too many? Too few? Take "present perfect" and try to split it using the rules in §2.2 — does that produce sensible units, or nonsense?

3Definition of mastery§5.2

I have defined "mastered" as: correct at error-correction level, on two separate days at least a week apart. For a B1 learner and a point like article use — too strict, too loose, or roughly right?

4Naming the mistake§3.5

When a student chooses a wrong answer, can you usually name the misconception behind it? Could you write three named misconceptions for a point like present perfect? The whole product depends on this being possible for a normal teacher, not just an expert.

5A number, not an opinion§5.5

Roughly how many minutes would it take you to write fifteen practice questions for one grammar point, with every wrong option labelled and a short explanation written? I need to multiply your number by 1,100 to find out whether this business can exist.

6The question ladder§3.2

The design assumes correcting an error shows more knowledge than filling a gap, which shows more than multiple choice. Do you agree with that order? Some teachers argue error-correction is easier, because the mistake has already been isolated for the student.

7Where is this naive?anywhere

Where does this design misunderstand how learners actually behave — motivation, honesty, effort, the difference between a classroom and a tired person on a phone at 11pm?

If you only answer one

Answer question 1. Everything else I can eventually work out on my own. That one I cannot.

Hypothesis

Startup Thesis v0.1

A set of beliefs written down so they can be attacked. Not a plan.

1Core hypothesis

We believe that many Iranian adults spend years learning English through classes, books, apps, videos and AI tools, but their learning is fragmented. They repeatedly make the same grammar and vocabulary mistakes because existing tools focus mainly on teaching new material rather than identifying individual weaknesses and systematically revisiting them until they are actually learned.

Our hypothesis is that learners will get better results, and some will pay, for a structured English-learning product that continuously identifies their weak points, creates a personalized learning path around those weaknesses, and uses repeated testing and spaced review to confirm real improvement.

This thesis is a hypothesis to test, not a statement of fact.

2Initial target customer

Iranian adults roughly 20–35 years old, around B1 English level, who have studied English for several years but still repeatedly make grammar and vocabulary mistakes and feel that their learning is unstructured or inefficient.

They may learn through language institutes, self-study books, ChatGPT, YouTube, apps such as Duolingo, private teachers, or a combination of these.

We are intentionally not trying to serve every learner, age group, level or objective in the first version.

3The problem

The main problem is not lack of content. There is already an enormous amount available. The problem is that learners often do not know exactly:

  • which concepts they genuinely understand,
  • which concepts they repeatedly fail,
  • which mistakes were simple slips versus real knowledge gaps,
  • what they should study next,
  • when they should review something they previously learned,
  • and whether a weakness has actually been resolved.

As a result, learners may spend significant time studying while still repeating the same mistakes months later. Their learning history is also scattered between teachers, books, apps, notes, flashcards and AI conversations. None of these keeps a unified model of what the learner knows, what they are weak at, and what should be reviewed next.

4Proposed solution

A system centred on a continuously updated learner weakness model. Initially it would:

  1. assess the learner's current English ability,
  2. break grammar and vocabulary into small, precisely tagged concepts,
  3. collect evidence from the learner's answers,
  4. identify concepts where the learner is probably weak,
  5. recommend short lessons and exercises for those weaknesses,
  6. bring previously missed concepts back at appropriate intervals,
  7. distinguish isolated mistakes from repeated evidence of weakness,
  8. track confidence and mastery over time,
  9. and show the learner measurable progress.

Instead of answering "What lesson comes next?" the product should increasingly answer "What should this particular learner practise next, and why?"

5Why existing alternatives may be insufficient

Traditional classes provide structure but usually cannot continuously personalize practice for every individual learner. Books provide high-quality structured material, but they do not remember the learner's mistakes or adapt future exercises to them. Apps such as Duolingo provide repetition and engagement, but their learning path may not provide the detailed weakness diagnosis we want to build. General-purpose AI tools such as ChatGPT can explain almost anything, but the learner must decide what to ask, maintain their own discipline, and organize their learning themselves.

The differentiation is therefore not more English content. It is diagnosis + structured learning + persistent weakness tracking + adaptive review.

6Initial product wedge

The first product should focus on grammar weakness detection and review for intermediate learners. A learner takes a diagnostic assessment, receives a small number of identified weaknesses, completes targeted lessons and exercises, then meets those concepts again through spaced review.

Vocabulary, IELTS/TOEFL preparation, speaking, academic English, business English and teacher dashboards may become future product areas. None are required to prove the initial thesis.

7Business model hypothesis

The initial business should aim to become a sustainable paid product rather than depend on venture funding. Possible models: paid access to structured grammar courses, premium diagnostic assessments, subscriptions for continuous personalized review, or a combination.

Pricing must fit the Iranian market and must be tested with real users rather than assumed. Content should be original or properly licensed. The business should not depend on reproducing copyrighted grammar books without permission.

8Why this could become defensible

The long-term value may come from the accumulated understanding of the learner rather than from individual lessons:

Concept → evidence → confidence → weakness
        → interventions → review history → mastery

If this model accurately determines what a learner should practise next, the product becomes increasingly personalized and difficult to replace with a static course or an isolated AI conversation. The quality of the learning model therefore matters more than the number of features.

9Critical assumptions that must be tested
#AssumptionWho can test it
1Problem. Learners genuinely feel repeated mistakes and fragmented learning are important problems.Learners
2Demand. The problem is painful enough that they will use a product for it repeatedly.Learners
3Willingness to pay. A meaningful share will pay for diagnosis and continuous review.Learners
4Learning model. We can identify weaknesses accurately enough that users trust the recommendations.Teachers
5Measurement. We can tell an occasional slip apart from a genuine weakness.Teachers
6Outcome. Adaptive review beats simply completing lessons in order.Teachers
7Content. We can produce enough tagged content at a reasonable cost.Teachers
8AI. AI can help without making the product dependent on an unreliable provider.Build & test
9Infrastructure. Practical, legal payment and AI infrastructure exists for the Iranian market.Research
10Retention. Learners keep returning for reviews rather than finishing a few lessons and leaving.Learners
Four of these are yours

Assumptions 4 to 7 are pedagogy, not business. They cannot be tested by talking to learners or by writing more code. They need teachers.

10What would make us change direction
  • learners do not consider repeated mistakes a meaningful problem,
  • they understand the problem but are unwilling to pay to solve it,
  • personalized weakness recommendations do not improve engagement or learning,
  • users regularly disagree with the system's weakness diagnosis,
  • high-quality tagged content is prohibitively expensive to produce,
  • or infrastructure constraints make the economics impractical.

The purpose of the first phase is not to prove the idea is good. It is to discover as cheaply and quickly as possible whether these assumptions are true.

Design

Learning Model v0.1

How the system decides what someone is weak at. Every number is a starting value chosen to be reasonable, not one proven correct.

Scope: grammar only, Iranian adults around B1, web only. Out of scope: vocabulary, speaking, writing assessment, IELTS/TOEFL, instructors, children.

1What this document decides

Before any code is written, the system needs written answers to five questions:

  1. What is the smallest thing a learner can be weak at?
  2. What counts as evidence that they are weak at it?
  3. When do we say they have mastered it?
  4. When does mastery expire?
  5. What do we show them, and when are we confident enough to show it?

Everything else — the placement test, the lesson order, the review schedule, the dashboard — is a consequence of these five answers.

2The concept taxonomy

2.1 Three levels, fixed

Domain              e.g. Verb Tenses
  └── Concept            e.g. Present Perfect
        └── Micro-concept    e.g. Present Perfect vs Past Simple
                                  with time markers

The micro-concept is the unit of measurement. Scores, weakness labels and review scheduling all attach to micro-concepts. Concepts and domains are computed by rolling up their children, and exist for display only.

2.2 Granularity rule

Fixed granularity, variable reporting. Every learner is measured at micro-concept level. Beginners are shown results rolled up; advanced learners see micro-concepts. A taxonomy that changes shape per learner cannot be authored, tested, or compared across users.

The test for a valid micro-concept — all three must be true:

  1. You could write a useful explanation of it in three to five sentences.
  2. You could write at least 15 distinct practice items that test it and not its siblings.
  3. A real learner could plausibly be right on it and wrong on its sibling.

If (3) fails, merge with the sibling. If (2) fails, it is too narrow.

2.3 Target size

For A2–B2, expect 60–90 micro-concepts. At 300 the granularity is too fine and there will never be enough data per unit to say anything with confidence. At 20 we are back to "Grammar: 67%."

This should not be invented from scratch. Cambridge's English Grammar Profile is a free, CEFR-mapped inventory of grammar features drawn from learner corpus data. Use it as the skeleton, then cut it to what matters for Persian speakers.

2.4 Prerequisites

Each micro-concept lists prerequisites. Used for one thing in v0.1: never recommend a concept whose prerequisites are Weak or Unknown. Fixing "Present Perfect vs Past Simple" is wasted effort if the learner cannot form past participles.

2.5 The Persian-interference list

One extra field per micro-concept: is this a known difficulty for Persian speakers? Articles, present perfect and preposition choice are strong candidates. Cheap to add, a real differentiator against Duolingo, and it gives the placement test better starting assumptions.

Question 1 for teachers

That list is currently three guesses made by someone who is not a teacher. Your five persistent-error grammar points would replace it with something real.

3The evidence model

3.1 The observation record

Every answer creates one immutable row. Nothing is deleted or overwritten.

FieldNotes
learner_id
item_id
micro_concept_idThe item's primary concept
secondary_concept_ids[]See 3.4
format_levelL1–L4, see 3.2
is_correct
distractor_idWhich wrong option was chosen — see 3.5
response_time_ms
position_in_sessionFor fatigue detection
strength_beforeThe model's estimate before this answer
predicted_p_correctWhat the model expected — see §8

The last two fields cost nothing to store and are what will later let us prove the model works. They must exist from day one; they cannot be reconstructed afterwards.

3.2 The question ladder

Not all correct answers mean the same thing.

LvFormatGuessIf rightIf wrong
L14-option multiple choice~25%0.51.0
L2Fill the gap, base word given~5%0.81.0
L3Find and correct the error~2%1.01.0
L4Write your own sentence~0%1.50.8

Wrong answers are weighted at or near full value everywhere, because a wrong answer is hard to fake. Correct answers at L1 are worth half, because a quarter of them are luck. L4 wrong answers are weighted down, because free writing fails for many reasons that are not the target concept. Mastery cannot be reached on L1 and L2 evidence alone.

Question 6 for teachers

Is this ordering right? Some argue error-correction is easier, because the mistake has already been isolated.

3.3 Slip versus error

You cannot design questions so that careless mistakes are rare. Slips are a permanent feature of human performance, and treating every slip as a weakness means telling people they are bad at things they already know — after which they stop believing anything the system says.

So slips are modelled, not eliminated. A wrong answer is flagged a probable slip if two or more hold:

  • response time under 40% of that learner's median for that format
  • prior strength on the concept was already ≥ 0.85
  • the answer came after item 25 in one session (fatigue)
  • the chosen wrong option carries no named misconception — an arbitrary answer

A probable slip counts at 0.3 weight instead of 1.0. It is never discarded. If a learner "slips" on the same point three times, it is not a slip.

3.4 Questions that test more than one thing

  • Wrong answers give evidence to the primary concept only.
  • Correct answers give full weight to the primary, and 0.4 to each secondary.

Getting something right is evidence you knew all of it. Getting it wrong is evidence about only one thing.

3.5 Tagging the wrong options — the most important content decision

Every wrong option must be tagged with the misconception it represents.

Q102 — primary: present-perfect-vs-past-simple

  A  "I have seen him yesterday"  → perfect-with-finished-time
  B  "I saw him yesterday"        → CORRECT
  C  "I had seen him yesterday"   → overuse-of-past-perfect
  D  "I am seeing him yesterday"  → untagged (arbitrary)

With this, a wrong answer stops being "he failed Q102" and becomes "he applies present perfect to finished time expressions." That sentence is what the whole product promise rests on.

Cannot be retrofitted

This costs roughly 20% more authoring time per question and cannot be added to 5,000 existing questions later. It has to be a hard requirement from question number one — which is why question 5 matters so much.

4The mastery score

4.1 Two numbers, not one

  • strength — the estimated chance the learner would get a fresh error-correction question on this point right.
  • confidence — how much we trust that estimate.

One number is not enough. A strength of 0.42 from two answers and 0.42 from forty answers mean completely different things, and only one of them should ever be shown to a learner.

4.2 Updating strength

observed  = 1 if correct else 0
w         = format_weight × slip_multiplier
α         = 0.30 × w

strength ← strength + α × (observed − strength)

New concepts start at 0.5, adjusted by the placement test and lowered if the point is on the Persian-interference list.

4.3 Updating confidence

confidence = min(1, Σ weights over last 60 days / 6)
             × format_diversity_bonus
             × recency_factor
  • format diversity: 1.0 if evidence exists at two or more question types, 0.7 if only one.
  • recency: decays from 1.0 to 0.5 over 90 days since the last answer.

The intent: roughly six weighted answers across at least two question types before the system claims to know anything.

5The five states
StateCondition
Unknownconfidence < 0.35
Weakconfidence ≥ 0.35 and strength < 0.50
Developingconfidence ≥ 0.35 and 0.50 ≤ strength < 0.85
Masteredsee below
Needs reviewwas mastered, strength decayed below 0.75

A learner cannot be labelled Weak until there is enough evidence. Until then the point is Unknown, and Unknown points are shown as "not tested yet," never as a weakness.

5.2 Mastery — all four must hold

  1. strength ≥ 0.85
  2. confidence ≥ 0.70
  3. at least one correct answer at L3 or L4
  4. correct answers on two calendar days, at least 7 days apart

Condition 4 is the one that matters. Getting a structure right once tells you it was in working memory. Getting it right twice across a gap tells you it was retained. Without it, "mastered" quietly means "recently studied."

Question 3 for teachers

Seven days, two occasions, at error-correction level. Right bar for a B1 learner? What would you set instead?

5.3 Marking something weak

A point becomes Weak when confidence crosses 0.35 with strength below 0.5 — roughly two full-weight wrong answers with no offsetting correct ones, or more if slips are involved. It falls out of the evidence rules rather than being a separate hard-coded counter.

5.4 Decay and re-testing

Decay:        strength − 0.02 per idle week, floor 0.60
Re-probe at:  strength ≤ 0.80  (roughly 25 days)
Re-probe is always L3 or L4, never L1

First wrong  → Needs review (not Weak). Probe again within 3 days.
Second wrong → Weak. Back into the active learning queue.
Correct      → strength = 0.95, next interval × 1.8

A single wrong answer after a long gap is exactly when a slip is most likely, which is why one failure does not undo mastery.

5.5 Never repeat the question

A point may be re-probed many times, but with a different question each time, and no question may be shown to the same learner twice within 90 days. Otherwise we are measuring memory of one sentence, not knowledge of a structure — a failure invisible in the metrics and fatal to the product.

This is where the money goes — question 5

Minimum 15 questions per micro-concept across all four types. At 75 micro-concepts that is a floor of roughly 1,100 authored questions before launch, every one with tagged wrong options. Almost certainly the largest single cost in the business, larger than engineering.

6The placement test

A 20-question test cannot diagnose 75 micro-concepts. It is arithmetically impossible, and pretending otherwise produces a report full of confident nonsense a learner would immediately recognise as wrong. So the placement test produces:

  • a coarse CEFR band (A2 / B1 / B2), which is reliable
  • a low-confidence starting assumption for each area
  • three to five candidate weak areas, framed as "areas to check first"

It does not produce the weakness map. That emerges over the first two weeks, and the product should say so out loud: "This is a starting point. Your real map builds as you practise."

Shape

  • 20–25 questions, 12–15 minutes, adaptive
  • start at B1; move up after two consecutive correct, down after two consecutive wrong
  • cover breadth across all areas, not depth
  • mix L1 and L2 only, to keep it short
  • over-sample the Persian-interference list — that is where the signal is
  • placement answers stored at 0.6 weight; test conditions differ from practice conditions
7What the learner sees
  1. Never show an Unknown point as a weakness. Show it as untested.
  2. Never show raw numbers. articles-specific-reference: 0.31 is for the database, not a person.
  3. At most three weaknesses at a time. A list of 22 things you are bad at is a reason to close the tab.
  4. Progress words, not deficit words: Solid / Getting there / Needs work / Next up. Avoid "Very weak."
  5. State the misconception, not the score. "You use present perfect with finished time expressions like yesterday" is the output that justifies the product existing.
  6. Show the evidence. When claiming a weakness, show the two or three sentences where it happened. This is what makes the diagnosis believable rather than magical.
8How we find out whether the model is any good

A. Calibration. Every answer stores what the model predicted. Bucket the predictions and compare predicted rate against actual rate — in a well-calibrated model, questions predicted at 0.7 are answered correctly about 70% of the time. Then compare against a dumb baseline: predicting the learner's own overall accuracy for everything. If the concept model does not beat that baseline, the weakness engine is decoration.

B. Slip rate. Measure the share of wrong answers landing on points already at strength ≥ 0.85. Above roughly 15%, either the slip rules are too loose or the mastery threshold is too low.

C. Concurrent validity. Have 20 learners take our placement test and a validated free test (Oxford Placement Test, EF SET) within a few days, and correlate the results. Below about r = 0.7, our test is not measuring level.

9Parameters to tune

All the arbitrary numbers in one place. Configuration, not code.

Parameterv0.1
Learning rate α0.30
Confidence denominator6 answers
Confidence window60 days
Single-format penalty×0.7
Slip weight multiplier0.3
Slip fast-response threshold40%
Weak threshold< 0.50
Mastery strength≥ 0.85
Mastery confidence≥ 0.70
Mastery spacing2 days, ≥7 apart
Decay rate−0.02/wk
Decay floor0.60
Re-probe trigger≤ 0.80
Interval growth×1.8
Question cooldown90 days
Min questions per point15
10Open questions
  1. Session length. How many questions per day is realistic for a working adult in Iran? This drives how quickly the product feels useful.
  2. Time to a first useful map. Answered by simulation — see The evidence.
  3. Grading free writing. L4 needs a grader. Given AI access constraints for Iran, this may need to be rule-based, human, or deferred.
  4. Explanation quality. The model tells you what is wrong. Writing the three-sentence explanation of why, for 75 points, for a Persian speaker, is a content problem this document does not solve.
  5. Motivation. Nothing here addresses why someone opens the app on day 12. That is a separate document and it may matter more than all of this.

Evidence

What the simulation found

The only part of this document that is not an opinion. Also the weakest kind of evidence there is.

Two hypotheses were written down and there was no evidence for any of them. Three ways to get evidence, running in parallel:

TrackTestsStatus
Learner interviews
12–15 people who are not friends
Assumptions 1–3Not started
Content cost pilot
3 points, 15 tagged questions each, timed
Assumption 7Not started
Model simulation
no learners required
Open question 2Done

The simulation runs the rules from The model against invented learners: 75 grammar points, a plausible spread of ability, a 5% slip rate, 20 practice questions per day, 500 simulated learners per data point. It asks when the system first produces three confident weaknesses, and how many of those three are real.

1 — Speed is not the problem

Time to a first weakness map is days, not weeks. Around day 2 when practice is guided by the placement test, day 7 when questions are chosen at random. The thresholds do not need loosening.

2 — The design is missing its most important rule

Changing only the rule for which question to show next moved time-to-diagnosis from day 1 to day 16, and accuracy by 18 points. That is a larger effect than any parameter in the table, and the model document does not mention it at all. There is a hole where its most consequential decision should be.

3 — Day-one accuracy is about 82%

Of the three weaknesses shown to a brand-new user, roughly one is wrong. They will notice. This is a direct threat to assumption 4 of the thesis.

Demanding more evidence barely helps: raising the confidence threshold from 0.35 to 0.85 moved accuracy from 81% to 83%. The problem is the quality of the evidence, not the amount. A correct 4-option answer currently pushes the estimate toward "knows it," when a quarter of those answers are guesses.

4 — Counting correct answers properly helps a little

80% 85% 90% 95% 100% Day 1 Day 3 Day 7 Day 14 Day 30 model as written with correction

Share of shown weaknesses that are real, at 20 practice questions per day.

Two changes for v0.2:

The product consequence matters more than the parameters

At 82% on day one, the product cannot open with a confident weakness map. The first screen has to say "three areas to check," and the real map has to earn its confidence over the first week. That is a change to the product itself, not a configuration value.

What this evidence is not

The simulation uses invented learners. It tests whether the rules are internally coherent — not whether they describe how Iranians actually learn English. It can rule a design out. It cannot prove one right. That is what teacher review and learner interviews are for.