Where a language app actually stops
Not "apps are bad" — apps are excellent at the thing they do. The model puts their ceiling at B1, and the reason is structural: nothing in an app corrects a sentence you have never said before.

Every estimate on this site can return a refusal. The most common one is this: app-only study does not reliably reach B2.
That is a strong claim and it deserves an argument rather than an assertion.
What the model says
Our estimates multiply an FSI baseline by a factor for how you study. The app-only factor is not one number — it grows with the level you are aiming at:
| Aiming at | App-only multiplier |
|---|---|
| A1 | 1.3× |
| C1 | 3.5× |
And above B1 it stops being a multiplier at all, because the model marks the level unreachable.
Concretely, Spanish to B2 at an hour a day, five days a week:
| Method | Hours | Reaches B2? |
|---|---|---|
| Living in-country | 345 | yes |
| Intensive class (FSI-style) | 438 | yes |
| 1-on-1 tutor | 460 | yes |
| Structured self-study course | 695 | yes |
| App only | 1,292 | no |
The 1,292 is shown struck through on the method chart, and the point of showing it is that it is not an answer. However many hours go in, the bar does not arrive.
Why the multiplier grows
This is the part worth understanding, because it explains why apps feel so effective at first and so stuck later.
At A1, an app is doing almost the right thing. The task is recognition and recall of a fixed set of forms. An app is a very good flashcard with a good scheduler, and the 1.3× penalty is mostly the absence of pronunciation feedback.
At B2, the task has changed. B2 is producing sentences you have not seen before, under time pressure, that are correct — and finding out when they are not. The skill is generating novel output and having it corrected.
An app cannot correct a sentence you invent, because it does not know what you were trying to say. It can only check whether you reproduced a sentence it already had. That gap does not close with more content or better software; it is what the medium is.
Hence the shape: cheap at the bottom, expensive in the middle, impossible at the top.
What this is not
Not "apps do not work". They are the best tool available for vocabulary and early forms, they are the reason many people start at all, and the free ones are genuinely free. Our recommendations include zero-commission options for exactly this reason.
Not a precise threshold. B1 is where the model places the ceiling; the real edge is fuzzy, depends on the app, and some people push past it by supplementing. A person doing app work plus conversation practice is not an app-only learner, and the model treats them differently.
Not a claim about any specific product. The ceiling is about the format, not the brand.
The practical version
If your target is A2 — travel, courtesy, reading signs — an app can take you there and the price is right.
If your target is B2 — work, study, arguing a point — the app is a component and not the plan. What is missing is corrected output, and the cheapest way to buy it is not another subscription. It is a person, for fewer hours than you think.
Work out what your own target costs by method, or read how the method factors were set, including which of them are estimates rather than published figures. Several are, and they are marked.
Sources
Every figure above, and where it came from. A post that traces other people's unsourced numbers does not get to have any of its own.
- Foreign Language Training — US Department of State, Foreign Service Institute
The class-hour bands every estimate on this site starts from, and the definition of the proficiency they lead to (ILR Speaking-3 / Reading-3).
- Guided learning hours — Cambridge English
The shape of the curve between CEFR levels. Cambridge publishes cumulative hours per level (A2 180–200, B1 350–400, B2 500–600, C1 700–800, C2 1,000–1,200); our level factors are normalised against the C1 band and every one lands inside its published range.