This is the working design for an English-learning system built around a single job: find out exactly what a learner is weak at, practise that, and prove the weakness is gone. It is not built yet. This document exists so that teachers can tell me where the thinking is wrong.
Answer these three as your weakest student would, not as you would.
The idea
And, more importantly, what it is deliberately not trying to do.
An Iranian adult who has studied English for six years is not short of material. They have books, apps, YouTube, a teacher, and now ChatGPT. What they do not have is an answer to a simple question: which specific things am I still getting wrong, and what should I do about them today? So they study broadly, revise things they already know, and three years later they are still dropping articles and still using the present perfect with yesterday.
Stop teaching English in general. Continuously build a picture of what one particular learner cannot do, aim the practice there, and keep re-testing until the mistake genuinely stops.
A short test places the learner
Around 20 adaptive questions. This gives a reliable CEFR band and three to five areas worth checking. It does not, and cannot, produce a full diagnosis.
Every answer becomes evidence
Each question is tagged with the exact grammar point it tests, and each wrong option is tagged with the misconception behind it. A wrong answer names the error rather than just recording a failure.
Repeated evidence becomes a diagnosis
One mistake is a slip. Several mistakes on the same point, across different question types, is a weakness. The system waits until it is confident before telling the learner anything.
Practice aims at the weakness, then comes back
Short lessons and targeted exercises, then the same grammar point returns days and weeks later in harder formats — until the learner can produce it, not just recognise it.
Iranian adults roughly 20–35, around B1, who have studied for years and still repeat basic mistakes. Not children, not beginners, not IELTS candidates, not every learner in the world. One group first.
No IELTS or TOEFL simulation. No speaking assessment. No teacher marketplace. No mobile apps. No vocabulary system at first, because spaced repetition already solves that well and building it would prove nothing. The first version tests one thing only: whether a system can find a learner's real weaknesses, and whether anyone cares.
Nothing here has been built or tested with a real learner. Every number in these documents is a reasonable starting guess, chosen so it can be corrected later against real data. Treat all of it as claims to be attacked, not conclusions.
What I'm asking you
You teach English. I do not, and this design makes claims about how learning works that I am not qualified to make alone.
If you tell me this is a nice idea, you have given me nothing. Reply with the question number and the word "no" wherever it applies. Disagreement is the entire purpose of sending you this.
Which grammar points do your Iranian students get wrong persistently — where you taught the rule properly, they understood it in class, and three months later they are still making the same mistake? A list of five is worth more to me than everything else in this document combined.
The design breaks grammar into 60–90 small units for the A2–B2 range. Too many? Too few? Take "present perfect" and try to split it using the rules in §2.2 — does that produce sensible units, or nonsense?
I have defined "mastered" as: correct at error-correction level, on two separate days at least a week apart. For a B1 learner and a point like article use — too strict, too loose, or roughly right?
When a student chooses a wrong answer, can you usually name the misconception behind it? Could you write three named misconceptions for a point like present perfect? The whole product depends on this being possible for a normal teacher, not just an expert.
Roughly how many minutes would it take you to write fifteen practice questions for one grammar point, with every wrong option labelled and a short explanation written? I need to multiply your number by 1,100 to find out whether this business can exist.
The design assumes correcting an error shows more knowledge than filling a gap, which shows more than multiple choice. Do you agree with that order? Some teachers argue error-correction is easier, because the mistake has already been isolated for the student.
Where does this design misunderstand how learners actually behave — motivation, honesty, effort, the difference between a classroom and a tired person on a phone at 11pm?
Answer question 1. Everything else I can eventually work out on my own. That one I cannot.
Hypothesis
A set of beliefs written down so they can be attacked. Not a plan.
We believe that many Iranian adults spend years learning English through classes, books, apps, videos and AI tools, but their learning is fragmented. They repeatedly make the same grammar and vocabulary mistakes because existing tools focus mainly on teaching new material rather than identifying individual weaknesses and systematically revisiting them until they are actually learned.
Our hypothesis is that learners will get better results, and some will pay, for a structured English-learning product that continuously identifies their weak points, creates a personalized learning path around those weaknesses, and uses repeated testing and spaced review to confirm real improvement.
This thesis is a hypothesis to test, not a statement of fact.
Iranian adults roughly 20–35 years old, around B1 English level, who have studied English for several years but still repeatedly make grammar and vocabulary mistakes and feel that their learning is unstructured or inefficient.
They may learn through language institutes, self-study books, ChatGPT, YouTube, apps such as Duolingo, private teachers, or a combination of these.
We are intentionally not trying to serve every learner, age group, level or objective in the first version.
The main problem is not lack of content. There is already an enormous amount available. The problem is that learners often do not know exactly:
As a result, learners may spend significant time studying while still repeating the same mistakes months later. Their learning history is also scattered between teachers, books, apps, notes, flashcards and AI conversations. None of these keeps a unified model of what the learner knows, what they are weak at, and what should be reviewed next.
A system centred on a continuously updated learner weakness model. Initially it would:
Instead of answering "What lesson comes next?" the product should increasingly answer "What should this particular learner practise next, and why?"
Traditional classes provide structure but usually cannot continuously personalize practice for every individual learner. Books provide high-quality structured material, but they do not remember the learner's mistakes or adapt future exercises to them. Apps such as Duolingo provide repetition and engagement, but their learning path may not provide the detailed weakness diagnosis we want to build. General-purpose AI tools such as ChatGPT can explain almost anything, but the learner must decide what to ask, maintain their own discipline, and organize their learning themselves.
The differentiation is therefore not more English content. It is diagnosis + structured learning + persistent weakness tracking + adaptive review.
The first product should focus on grammar weakness detection and review for intermediate learners. A learner takes a diagnostic assessment, receives a small number of identified weaknesses, completes targeted lessons and exercises, then meets those concepts again through spaced review.
Vocabulary, IELTS/TOEFL preparation, speaking, academic English, business English and teacher dashboards may become future product areas. None are required to prove the initial thesis.
The initial business should aim to become a sustainable paid product rather than depend on venture funding. Possible models: paid access to structured grammar courses, premium diagnostic assessments, subscriptions for continuous personalized review, or a combination.
Pricing must fit the Iranian market and must be tested with real users rather than assumed. Content should be original or properly licensed. The business should not depend on reproducing copyrighted grammar books without permission.
The long-term value may come from the accumulated understanding of the learner rather than from individual lessons:
Concept → evidence → confidence → weakness
→ interventions → review history → mastery
If this model accurately determines what a learner should practise next, the product becomes increasingly personalized and difficult to replace with a static course or an isolated AI conversation. The quality of the learning model therefore matters more than the number of features.
| # | Assumption | Who can test it |
|---|---|---|
| 1 | Problem. Learners genuinely feel repeated mistakes and fragmented learning are important problems. | Learners |
| 2 | Demand. The problem is painful enough that they will use a product for it repeatedly. | Learners |
| 3 | Willingness to pay. A meaningful share will pay for diagnosis and continuous review. | Learners |
| 4 | Learning model. We can identify weaknesses accurately enough that users trust the recommendations. | Teachers |
| 5 | Measurement. We can tell an occasional slip apart from a genuine weakness. | Teachers |
| 6 | Outcome. Adaptive review beats simply completing lessons in order. | Teachers |
| 7 | Content. We can produce enough tagged content at a reasonable cost. | Teachers |
| 8 | AI. AI can help without making the product dependent on an unreliable provider. | Build & test |
| 9 | Infrastructure. Practical, legal payment and AI infrastructure exists for the Iranian market. | Research |
| 10 | Retention. Learners keep returning for reviews rather than finishing a few lessons and leaving. | Learners |
Assumptions 4 to 7 are pedagogy, not business. They cannot be tested by talking to learners or by writing more code. They need teachers.
The purpose of the first phase is not to prove the idea is good. It is to discover as cheaply and quickly as possible whether these assumptions are true.
Design
How the system decides what someone is weak at. Every number is a starting value chosen to be reasonable, not one proven correct.
Scope: grammar only, Iranian adults around B1, web only. Out of scope: vocabulary, speaking, writing assessment, IELTS/TOEFL, instructors, children.
Before any code is written, the system needs written answers to five questions:
Everything else — the placement test, the lesson order, the review schedule, the dashboard — is a consequence of these five answers.
Domain e.g. Verb Tenses
└── Concept e.g. Present Perfect
└── Micro-concept e.g. Present Perfect vs Past Simple
with time markers
The micro-concept is the unit of measurement. Scores, weakness labels and review scheduling all attach to micro-concepts. Concepts and domains are computed by rolling up their children, and exist for display only.
Fixed granularity, variable reporting. Every learner is measured at micro-concept level. Beginners are shown results rolled up; advanced learners see micro-concepts. A taxonomy that changes shape per learner cannot be authored, tested, or compared across users.
The test for a valid micro-concept — all three must be true:
If (3) fails, merge with the sibling. If (2) fails, it is too narrow.
For A2–B2, expect 60–90 micro-concepts. At 300 the granularity is too fine and there will never be enough data per unit to say anything with confidence. At 20 we are back to "Grammar: 67%."
This should not be invented from scratch. Cambridge's English Grammar Profile is a free, CEFR-mapped inventory of grammar features drawn from learner corpus data. Use it as the skeleton, then cut it to what matters for Persian speakers.
Each micro-concept lists prerequisites. Used for one thing in v0.1: never recommend a concept whose prerequisites are Weak or Unknown. Fixing "Present Perfect vs Past Simple" is wasted effort if the learner cannot form past participles.
One extra field per micro-concept: is this a known difficulty for Persian speakers? Articles, present perfect and preposition choice are strong candidates. Cheap to add, a real differentiator against Duolingo, and it gives the placement test better starting assumptions.
That list is currently three guesses made by someone who is not a teacher. Your five persistent-error grammar points would replace it with something real.
Every answer creates one immutable row. Nothing is deleted or overwritten.
| Field | Notes |
|---|---|
| learner_id | |
| item_id | |
| micro_concept_id | The item's primary concept |
| secondary_concept_ids[] | See 3.4 |
| format_level | L1–L4, see 3.2 |
| is_correct | |
| distractor_id | Which wrong option was chosen — see 3.5 |
| response_time_ms | |
| position_in_session | For fatigue detection |
| strength_before | The model's estimate before this answer |
| predicted_p_correct | What the model expected — see §8 |
The last two fields cost nothing to store and are what will later let us prove the model works. They must exist from day one; they cannot be reconstructed afterwards.
Not all correct answers mean the same thing.
| Lv | Format | Guess | If right | If wrong |
|---|---|---|---|---|
| L1 | 4-option multiple choice | ~25% | 0.5 | 1.0 |
| L2 | Fill the gap, base word given | ~5% | 0.8 | 1.0 |
| L3 | Find and correct the error | ~2% | 1.0 | 1.0 |
| L4 | Write your own sentence | ~0% | 1.5 | 0.8 |
Wrong answers are weighted at or near full value everywhere, because a wrong answer is hard to fake. Correct answers at L1 are worth half, because a quarter of them are luck. L4 wrong answers are weighted down, because free writing fails for many reasons that are not the target concept. Mastery cannot be reached on L1 and L2 evidence alone.
Is this ordering right? Some argue error-correction is easier, because the mistake has already been isolated.
You cannot design questions so that careless mistakes are rare. Slips are a permanent feature of human performance, and treating every slip as a weakness means telling people they are bad at things they already know — after which they stop believing anything the system says.
So slips are modelled, not eliminated. A wrong answer is flagged a probable slip if two or more hold:
A probable slip counts at 0.3 weight instead of 1.0. It is never discarded. If a learner "slips" on the same point three times, it is not a slip.
Getting something right is evidence you knew all of it. Getting it wrong is evidence about only one thing.
Every wrong option must be tagged with the misconception it represents.
Q102 — primary: present-perfect-vs-past-simple
A "I have seen him yesterday" → perfect-with-finished-time
B "I saw him yesterday" → CORRECT
C "I had seen him yesterday" → overuse-of-past-perfect
D "I am seeing him yesterday" → untagged (arbitrary)
With this, a wrong answer stops being "he failed Q102" and becomes "he applies present perfect to finished time expressions." That sentence is what the whole product promise rests on.
This costs roughly 20% more authoring time per question and cannot be added to 5,000 existing questions later. It has to be a hard requirement from question number one — which is why question 5 matters so much.
One number is not enough. A strength of 0.42 from two answers and 0.42 from forty answers mean completely different things, and only one of them should ever be shown to a learner.
observed = 1 if correct else 0
w = format_weight × slip_multiplier
α = 0.30 × w
strength ← strength + α × (observed − strength)
New concepts start at 0.5, adjusted by the placement test and lowered if the point is on the Persian-interference list.
confidence = min(1, Σ weights over last 60 days / 6)
× format_diversity_bonus
× recency_factor
The intent: roughly six weighted answers across at least two question types before the system claims to know anything.
| State | Condition |
|---|---|
| Unknown | confidence < 0.35 |
| Weak | confidence ≥ 0.35 and strength < 0.50 |
| Developing | confidence ≥ 0.35 and 0.50 ≤ strength < 0.85 |
| Mastered | see below |
| Needs review | was mastered, strength decayed below 0.75 |
A learner cannot be labelled Weak until there is enough evidence. Until then the point is Unknown, and Unknown points are shown as "not tested yet," never as a weakness.
Condition 4 is the one that matters. Getting a structure right once tells you it was in working memory. Getting it right twice across a gap tells you it was retained. Without it, "mastered" quietly means "recently studied."
Seven days, two occasions, at error-correction level. Right bar for a B1 learner? What would you set instead?
A point becomes Weak when confidence crosses 0.35 with strength below 0.5 — roughly two full-weight wrong answers with no offsetting correct ones, or more if slips are involved. It falls out of the evidence rules rather than being a separate hard-coded counter.
Decay: strength − 0.02 per idle week, floor 0.60
Re-probe at: strength ≤ 0.80 (roughly 25 days)
Re-probe is always L3 or L4, never L1
First wrong → Needs review (not Weak). Probe again within 3 days.
Second wrong → Weak. Back into the active learning queue.
Correct → strength = 0.95, next interval × 1.8
A single wrong answer after a long gap is exactly when a slip is most likely, which is why one failure does not undo mastery.
A point may be re-probed many times, but with a different question each time, and no question may be shown to the same learner twice within 90 days. Otherwise we are measuring memory of one sentence, not knowledge of a structure — a failure invisible in the metrics and fatal to the product.
Minimum 15 questions per micro-concept across all four types. At 75 micro-concepts that is a floor of roughly 1,100 authored questions before launch, every one with tagged wrong options. Almost certainly the largest single cost in the business, larger than engineering.
A 20-question test cannot diagnose 75 micro-concepts. It is arithmetically impossible, and pretending otherwise produces a report full of confident nonsense a learner would immediately recognise as wrong. So the placement test produces:
It does not produce the weakness map. That emerges over the first two weeks, and the product should say so out loud: "This is a starting point. Your real map builds as you practise."
A. Calibration. Every answer stores what the model predicted. Bucket the predictions and compare predicted rate against actual rate — in a well-calibrated model, questions predicted at 0.7 are answered correctly about 70% of the time. Then compare against a dumb baseline: predicting the learner's own overall accuracy for everything. If the concept model does not beat that baseline, the weakness engine is decoration.
B. Slip rate. Measure the share of wrong answers landing on points already at strength ≥ 0.85. Above roughly 15%, either the slip rules are too loose or the mastery threshold is too low.
C. Concurrent validity. Have 20 learners take our placement test and a validated free test (Oxford Placement Test, EF SET) within a few days, and correlate the results. Below about r = 0.7, our test is not measuring level.
All the arbitrary numbers in one place. Configuration, not code.
| Parameter | v0.1 |
|---|---|
| Learning rate α | 0.30 |
| Confidence denominator | 6 answers |
| Confidence window | 60 days |
| Single-format penalty | ×0.7 |
| Slip weight multiplier | 0.3 |
| Slip fast-response threshold | 40% |
| Weak threshold | < 0.50 |
| Mastery strength | ≥ 0.85 |
| Mastery confidence | ≥ 0.70 |
| Mastery spacing | 2 days, ≥7 apart |
| Decay rate | −0.02/wk |
| Decay floor | 0.60 |
| Re-probe trigger | ≤ 0.80 |
| Interval growth | ×1.8 |
| Question cooldown | 90 days |
| Min questions per point | 15 |
Evidence
The only part of this document that is not an opinion. Also the weakest kind of evidence there is.
Two hypotheses were written down and there was no evidence for any of them. Three ways to get evidence, running in parallel:
| Track | Tests | Status |
|---|---|---|
| Learner interviews 12–15 people who are not friends | Assumptions 1–3 | Not started |
| Content cost pilot 3 points, 15 tagged questions each, timed | Assumption 7 | Not started |
| Model simulation no learners required | Open question 2 | Done |
The simulation runs the rules from The model against invented learners: 75 grammar points, a plausible spread of ability, a 5% slip rate, 20 practice questions per day, 500 simulated learners per data point. It asks when the system first produces three confident weaknesses, and how many of those three are real.
Time to a first weakness map is days, not weeks. Around day 2 when practice is guided by the placement test, day 7 when questions are chosen at random. The thresholds do not need loosening.
Changing only the rule for which question to show next moved time-to-diagnosis from day 1 to day 16, and accuracy by 18 points. That is a larger effect than any parameter in the table, and the model document does not mention it at all. There is a hole where its most consequential decision should be.
Of the three weaknesses shown to a brand-new user, roughly one is wrong. They will notice. This is a direct threat to assumption 4 of the thesis.
Demanding more evidence barely helps: raising the confidence threshold from 0.35 to 0.85 moved accuracy from 81% to 83%. The problem is the quality of the evidence, not the amount. A correct 4-option answer currently pushes the estimate toward "knows it," when a quarter of those answers are guesses.
Share of shown weaknesses that are real, at 20 practice questions per day.
Two changes for v0.2:
At 82% on day one, the product cannot open with a confident weakness map. The first screen has to say "three areas to check," and the real map has to earn its confidence over the first week. That is a change to the product itself, not a configuration value.
The simulation uses invented learners. It tests whether the rules are internally coherent — not whether they describe how Iranians actually learn English. It can rule a design out. It cannot prove one right. That is what teacher review and learner interviews are for.