🗣️ Learning 27 min read

Best AI Tools for Language Learning in 2026

What the research says actually works, which skills AI barely touches, and how to practise so you improve instead of just chatting.

By whichaibest.com Team

Quick Answer

The best AI tools for language learning in 2026 are ChatGPT voice mode for free conversation practice, Speak or ELSA for pronunciation scoring, and a structured course like Duolingo for the syllabus. Research shows AI helps most with vocabulary and speaking confidence, and least with the long grind of grammar.

The best AI tool for language learning in 2026 is ChatGPT with voice mode, because it gives you free, on-demand conversation practice in almost any language, within a daily voice cap. Pair it with a structured course for the syllabus, and add Speak or ELSA if pronunciation is your weak point. That's the short version, and for most people it's the right answer.

But the more useful question is whether any of this actually works, and the research on that has got a lot better in the last two years. It turns out AI helps a great deal with some parts of learning a language and barely at all with others, and knowing which is which will save you months.

What Are the Best AI Tools for Language Learning in 2026?

Sorted by what each one is genuinely good at:

If you want a wider view of how these models differ before picking one, our model comparison pages break each one down by what it's actually good at.

Does AI Actually Help You Learn a Language?

Yes, and by 2025 the evidence stopped being anecdotal.

Mengdi Li, Yinyu Wang and Xiaorong Yang pooled 41 experimental and quasi-experimental studies covering 3,515 participants for a meta-analysis published in the Journal of Computer Assisted Learning in August 2025. Across 48 independent effect sizes, generative AI chatbots produced a moderate to large positive effect on second language acquisition: 0.576, with a 95 percent confidence interval of 0.385 to 0.768 and p below 0.001.

For context, an effect size approaching 0.6 in education research is the kind of result that gets a method taken seriously. A separate 2025 meta-analysis by Lyu, Lai and Guo in the International Journal of Applied Linguistics reached the same broad conclusion, finding a medium effect for chatbots on second language learning.

One detail from the Li analysis is worth pausing on. They found no meaningful difference by educational stage or learning environment. Whether you're a university student or teaching yourself at the kitchen table, the effect held. That's unusual and encouraging.

Here's the part that should temper your expectations, though, and it's buried in the moderator analysis. The large effects clustered around specific conditions: Indonesian as the learner's first language, a target language other than English, ChatGPT as the tool, vocabulary as the goal, and interventions lasting 1 to 7 days. If you're an English speaker working on Spanish over six months, you're outside most of those conditions. The 0.576 figure is an average across a literature that leans heavily toward learners who don't look like you. Still a real effect. Just don't read it as a promise.

And there's one number you should actively distrust. Go looking for evidence on AI and learning and you'll hit a 2025 meta-analysis reporting an effect of 0.867 on learning performance, which is huge. It's been retracted. Springer Nature published a retraction note for Jin Wang and Wenxiang Fan's paper after Magnus Ingebrigtsen and Marko Lukic raised problems with it. The analysis had pooled a study that was itself retracted, mislabelled others, and rested on trials too small to carry that much weight. Another critic, the Finnish researcher Ilkka Tuomi, had pointed out that 33 of its 51 studies had fewer than 35 students in the ChatGPT group. According to the note, the authors didn't respond to the journal's correspondence about the retraction.

That paper was about education generally rather than language specifically, but it gets quoted in language learning posts constantly, and it's still circulating. Retractions always travel slower than the headline did. So if you see someone claiming AI has a near transformative effect on learning, check whether that's where the number came from.

It's also a good argument for trusting the smaller figure. A 0.576 from a language-specific meta-analysis that held up beats a 0.867 from one that didn't.

And if you want a general-education number that hasn't fallen over, there is one, in the same journal that published the retraction. Xinning Wu, Pei Zhu, Jinliang Zhang and colleagues pooled 35 experimental studies covering 4,193 participants in Humanities and Social Sciences Communications, and landed on a moderately positive effect of 0.670. They also ran publication bias checks, including Begg's test for funnel asymmetry, and found no significant bias.

Notice how much smaller that is than the number still doing the rounds. 0.670 across education generally, 0.576 for language specifically, against a retracted 0.867. The honest read is that AI helps, and helps consistently, but it isn't the step change the louder headlines sold you.

Does What You Learn With AI Actually Stick?

This is the honest weak spot in the whole field. The evidence that AI improves what you produce is a lot stronger than the evidence that it improves what you keep.

The cleanest test of that distinction comes from Yizhou Fan, Luzhen Tang, Dragan Gasevic and colleagues, writing in the British Journal of Educational Technology in 2025 under a title that gives away the conclusion: "Beware of metacognitive laziness." They put 117 university students into four conditions for a writing task. One group got ChatGPT, one got a human expert, one got writing analytics tools, and one got no extra help at all.

The ChatGPT group improved their essay scores more than the others. That part matches every optimistic headline you've read.

Then the researchers measured what the students actually knew afterwards. Knowledge gain and transfer showed no significant difference across the four conditions. Post-task intrinsic motivation didn't differ either. What did differ significantly was the frequency and sequence of self-regulated learning processes, meaning the ChatGPT group went about the work differently. Better output, same knowledge.

Be careful how far you carry that finding, though, because I'd rather flag the limit than oversell it. This was academic essay writing with university students, not second language acquisition. It isn't a direct test of learning Spanish. But it measures the exact mechanism that should worry a language learner, which is whether help with the output turns into something you retain, and I couldn't find an L2-specific study that tests transfer as cleanly.

Now put it next to a detail from the meta-analysis further up this page. Those large effects clustered in interventions lasting 1 to 7 days. Short studies measure the near end of learning almost by definition. So the strongest evidence for AI in language learning sits precisely where retention is least tested. Two separate findings, pointing the same direction, and neither one gets mentioned in the marketing.

None of this means stop. It means the way you use the tool decides whether anything sticks. A few habits that put the effort back where it belongs:

Where Does Spaced Repetition Fit Into This?

Underneath everything else, ideally. The section above landed on retrieval as the thing that actually builds retention, and spaced repetition is just that same idea run on a schedule instead of left to memory.

It's also the best-evidenced technique in the entire field, which is a strange thing to have to say about something most AI tool round-ups skip completely.

S. K. Kim and Stuart Webb gathered the evidence for a 2022 meta-analysis in Language Learning, pooling 98 effect sizes from 48 experiments covering 3,411 learners. They compared three separate things: spaced practice against massed practice, longer gaps against shorter ones, and equal gaps against expanding ones. The headline holds. Spreading your encounters with a word across time beats bunching them together, and longer gaps beat shorter ones on delayed tests, which is the measurement that actually counts.

A companion meta-analysis adds a useful twist. Takumi Uchihara, Stuart Webb and Akifumi Yanagisawa pooled 45 effect sizes from 26 studies covering 1,918 participants, published in Language Learning in 2019 and indexed at ERIC. They found a medium correlation of r = 0.34 between repetition and incidental vocabulary learning. The moderator analysis is the useful bit, though, and it doesn't go the way you'd guess. In these incidental learning studies, the link between repetition and learning was stronger when the encounters were massed (r = 0.38) than when they were spaced (r = 0.23). So the strong case for spacing comes from Kim and Webb's delayed-test results, not from this one.

So here's the awkward part for AI tools. A chatbot has no memory of the word you fumbled three weeks ago, and no scheduler deciding when to put it back in front of you. Every conversation starts from nothing. That is precisely the massed-practice failure mode: a pile of encounters in one sitting, then silence.

Which is why the sensible recommendation isn't to use an AI tutor instead of flashcards. It's to use both, for different jobs. The chatbot is where you produce language and get it corrected. A spaced repetition system is where individual items come back at widening intervals until they stick.

The practical version costs nothing. Keep the error log this article already recommended, then actually feed it somewhere. After a session, ask the model to turn your own mistakes into question and answer pairs, and drop those into a free spaced repetition app. Anki is free on desktop and Android, and most other flashcard tools run some version of the same scheduling logic.

One warning. Don't let the AI spit out three hundred cards for you in one go. Cards covering things you never got wrong are cards you don't need, and a bloated deck is the quickest route to abandoning the whole thing. Your own errors are the highest-value material you have, because they come pre-filtered by what you personally struggle with.

Which Skills Does AI Help Most, and Which Least?

This is where the picture gets sharper, and where most tool roundups stop being useful.

Li and colleagues found vocabulary was a significant moderator, meaning the gains were strongest when vocabulary was the learning target. That lines up with a controlled trial by Zhihui Zhang and Xiaomeng Huang, published in Heliyon in 2024, which randomly split 52 high school English learners into chatbot and control groups.

Participants averaged 44.71 percent on the pre-test before being split. On the immediate test the chatbot group scored 73.93 percent on productive vocabulary against 66.24 percent for the control, a gap of 7.7 points. Useful, but not dramatic. The interesting result came two weeks later. On the delayed test the chatbot group held at 67.95 percent while the control had dropped to 55.55 percent, widening the gap to 12.39 points. Both differences were significant at p below 0.001.

So the retention advantage grew over time rather than fading. That's the strongest single argument for using an AI tutor: not that you learn words faster, but that you keep them. Worth remembering it's one study of 52 teenagers, so treat it as a promising signal rather than settled fact.

The caveat nobody mentions: Li and colleagues found intervention durations of just 1 to 7 days had a large moderating effect. Most of the studies in this literature are short. So we have good evidence that AI helps over days and weeks, and much thinner evidence about what happens over a year. Treat long-term claims with more caution than the short-term ones.

Can AI Make You Reading Material at Your Level?

It will happily try. Whether the result is actually easier to read is a different question, and the one direct test of it came back awkward.

The pitch is genuinely appealing. Finding material pitched just above where you are is the perennial problem in language learning. Real articles are too hard, textbook dialogues are too dull, and graded readers run out fast. An AI that rewrites any text to your level should solve all three at once.

Dennis Murphy Odo tested that directly, publishing in Language Learning & Technology in 2025. He compared how well second language readers understood three versions of the same passage: the original, a version simplified by a human, and one simplified by AI.

Neither simplified version came out easier to comprehend than the original once he accounted for how familiar readers already were with the topic.

Sit with that for a second, because it's doing two things at once. It isn't a finding that AI simplification is bad. The AI held its own against the human. What it questions is whether simplifying the text was the lever anyone thought it was, since topic familiarity was the one background factor linked to comprehension, weakly, while the version of the text made no measurable difference.

That's a useful correction if you've been asking an AI to dumb down articles for you. Shortening the sentences and swapping in commoner words doesn't reliably make a text land, because the thing making it hard may be that you don't know the subject.

So a better way to use it. Instead of "rewrite this at B1 level", try asking for the background first. What is this article assuming I already know? Give me the five terms and the context before I read it. Then read the original. You're attacking the part the research points at instead of the one it found made no difference, and you keep the real language instead of a flattened version of it.

And if you do want easier text, ask it to write something new on a topic you already understand well, rather than to compress something you don't. Familiar subject, unfamiliar language is the combination worth chasing.

Which Languages Does AI Handle Badly?

Every roundup you'll read treats "language learning" as one thing. It isn't. The tool that gives you fluent, well-corrected Spanish will quietly hand you worse Thai, and nothing in the interface tells you which one you're getting.

The reason is training data. Models learn from what's on the internet, and the internet is wildly lopsided. Spanish, French, German, Mandarin and Japanese are all over it. Khmer, Lao, Burmese and Javanese are not, and the models are correspondingly thinner on them.

This gap is documented well enough that there's now a benchmark built specifically to measure it. AI Singapore's AI Products team, working with the HELM group at Stanford's Center for Research on Foundation Models, built SEA-HELM, presented at ACL Findings in 2025. It grades models across Filipino, Indonesian, Tamil, Thai and Vietnamese over 35,136 test cases, on five fronts including local linguistics, cultural knowledge and safety, rather than relying only on translated English tests.

Their conclusion is blunt. Significant gaps remain between high-resource languages like English and the mid to low-resource languages of Southeast Asia, several of which are official languages with millions of speakers. A benchmark had to be built from scratch precisely because English scores told you almost nothing about how a model performs in Thai.

There's a genuinely hopeful finding in there too. Dedicated fine-tuning of much smaller models, in the 7 to 9 billion parameter range, narrowed the gap substantially against far larger systems like GPT-4o and DeepSeek-R1. Size isn't what's missing. Attention is. That's why purpose-built model families like SEA-LION, which covers 11 Southeast Asian languages, can beat a bigger general model on its home turf.

Now hold that beside the Li meta-analysis from earlier, because the two together say something more interesting than either alone. Remember that the largest measured effects came from Indonesian first-language learners studying a target language other than English. So this isn't a simple story of AI being bad at Asian languages. It's about direction. These tools are strong at explaining a well-resourced language to you, in whatever language you speak. They're much weaker at correcting and assessing your output in a low-resource one.

What to do with that. If you're learning Spanish, French, German, Italian, Japanese or Mandarin, the research in this article applies to you fairly directly. If you're learning Thai, Vietnamese, Tagalog, Khmer, Lao or Burmese, use AI for explanation and reading practice, and be far more sceptical of its grammar corrections and its confidence about what sounds natural. Cross-check anything that matters with a native speaker or a proper dictionary. A model that's shaky in your target language won't hedge. It'll just be wrong, fluently.

What About Learning a Whole New Writing System?

This is the biggest blind spot in every AI language learning roundup, including the friendly ones. If your target language uses a script you can't already read, chat-based AI helps with roughly half the job and is close to useless at the other half. And it won't tell you which half you're in.

Worth being clear that this is a different problem from the low-resource one above. Mandarin and Japanese are both extremely well represented in training data. The section above said the research in this article applies fairly directly if you're learning them, and for speaking and grammar that's true. Characters are the exception. The gap here isn't a shortage of data. It's that learning a script is partly a motor and visual skill, and a text box cannot reach that.

There's a specific finding that makes this concrete rather than hand-wavy. Connie Qun Guan, Ying Liu, Derek Ho Leung Chan, Feifei Ye and Charles Perfetti ran a set of learning experiments published in the Journal of Educational Psychology in 2011, comparing handwriting practice for Chinese characters against reading-only study and then, in a second experiment with 37 adult learners, against an alphabetic pinyin typing tutor. The typing condition was designed to control for the fact that both groups were moving their hands.

The two routes built different things. Handwriting strengthened orthographic representation and the link between a character and its meaning. Pinyin typing strengthened phonology instead.

Sit with what that implies for how most people now practise. If your entire study routine is typing romanised input into a chat window and reading what comes back, you're training the phonological channel and largely skipping the orthographic one. You will get better at recognising how a word sounds and at holding a conversation. You will not build the same grip on the characters themselves, and you'll notice it the first time you try to write something without a keyboard predicting it for you.

The second problem is invisible and affects your free tier directly.

Models don't read characters, they read tokens, and scripts are not tokenised equally. Aleksandar Petrov, Emanuele La Malfa, Philip Torr and Adel Bibi measured this across a large set of languages for a paper at NeurIPS in 2023. The same passage translated into different languages produced token counts differing by up to 15 times. Even character-level and byte-level models, which you'd expect to even things out, still showed over 4 times the difference for some language pairs.

Three consequences you'll actually feel:

Petrov's team frame this as a fairness problem, and they're right, but for a learner it's simply a planning fact. Budget more sessions, or shorter ones, if your target script is token-expensive.

So what is AI genuinely good for here?

And what it's bad at, which is where people get burned:

The stack that works is unglamorous. Handwriting practice on paper or a dedicated app for production, because the Guan and Perfetti result shows handwriting builds character knowledge that typing doesn't. A spaced repetition deck for retention. AI for explanation, mnemonics, graded reading and example sentences, which it does better and faster than anything else available to a self-learner. And a native speaker or a real dictionary as the check on anything you're about to commit to memory, for the same reason set out in the section on languages AI handles badly. A model that's wrong about a stroke count won't hesitate. It'll just be wrong, neatly.

Does AI Get Politeness and Register Right?

Often no, and this is a different failure from the one above. A model can have excellent grammar in your target language and still quietly teach you to sound rude.

Politeness, formality and indirectness are what linguists call pragmatics. It's the gap between what a sentence means and what it does. Grammatically perfect Japanese in the wrong register lands badly with a manager. Korean honorifics are not decoration, they encode the relationship between the people speaking. Thai has particles that change the whole temperature of a request. None of that shows up in a grammar check, so none of it gets corrected.

Two pieces of research put numbers on how thin the ground is here.

Bolei Ma, Yuting Li and colleagues surveyed the field in Pragmatics in the Era of Large Language Models, published in 2025 and covering 58 papers, mostly from the ACL Anthology. They start from the premise that pragmatic competence remains insufficiently evaluated. The detail that matters for learners is the coverage gap. Only 19 percent of the resources they reviewed, roughly 11 of 57, include any non-English language at all. So this is barely measured in English, and hardly measured anywhere else.

Martina Manna, Maka Eradze and Federica Cominetti came at it from the training side in Frontiers in Education in 2025. They note that only just over 5 percent of Llama 3's training data was non-English. What comes out the other side is sometimes called translationese: text that is grammatically correct and pragmatically off. Politeness strategies, register shifts and norms about how direct you are supposed to be are all highly culture-specific, and a model trained overwhelmingly on English carries English defaults into every language it speaks.

Their warning to learners is the part worth sitting with. Practise long enough against a model with English politeness norms baked in and you can absorb them. You end up fluent and subtly wrong in a way that's hard to unlearn, because nothing ever flagged it. They also cite work showing smaller models slipping back into English when a task gets hard, which tells you how thin that non-English layer can be.

What to do about it:

This matters more the further your target language sits from English. If you're learning a Southeast Asian language, our guide to AI tools for SEA users covers which models handle the region's languages best.

Are AI's Corrections Actually Right?

The advice further down this page tells you to ask the AI to correct every mistake. Fair enough. But that assumes the corrections are correct, and the research says they're right more often than they're wrong while also flagging things that were never broken.

Jaganov, Blake, Villegas and Carr tested this properly in a 2025 paper on dynamic assessment of grammatical accuracy in learner writing. Running 200 test sentences, GPT-4o hit an F1 of 0.743, but the split underneath that number is the interesting bit: recall was 0.861 and precision only 0.654. In plain terms, it catches most of your real mistakes and also flags a fair number of things that were fine.

That imbalance has a name in the literature: overcorrection. And it shows up clearly in other languages too. Sha Lin evaluated three models on learner Chinese for PLoS One in 2024 and found LLMs tend to overcorrect, rewriting sentences with extra conjunctions, particles and modal verbs to make them flow, where a human marker would have made the smallest possible edit. The scores were sobering too. On the CGED-2021 dataset, GPT-4 managed an F0.5 of just 17.55.

Here's the detail that ties back to the language gap above. On that same Chinese benchmark, ERNIE-4 scored 30.28 against GPT-4's 17.55, and it beat GPT-4 on every dataset tested. A model built and trained in the language wins on its home turf, same as SEA-LION does for Southeast Asian languages. If your target language has a strong local model, it's worth trying for corrections specifically.

So what do you actually do with this? Three things:

Worth noting these are studies of written correction, not live voice conversation, where nobody has published comparable numbers. Given voice adds transcription errors on top, assume it's no better.

Does AI Help If You're Too Nervous to Speak?

This might be the most underrated reason AI practice works, and it has nothing to do with the quality of the AI.

Zafer Susoy ran a neat study on this, published in Frontiers in Psychology in January 2026. Forty-eight first-year English Language Teaching undergraduates at a Turkish university each sat two comparable speaking exams two weeks apart, one assessed by an AI chatbot and one by human instructors, with the order counterbalanced so practice effects couldn't skew it.

Anxiety before the AI exam was lower, but only modestly: a mean of 98.48 against 102.94 for the human exam, t(47) = 2.67, p = 0.01, with a small effect size of d = 0.39. On its own that's a shrug.

The interesting result is what anxiety did to their scores. In the human-assessed exams, anxiety and achievement correlated at r = -0.500, p below 0.01. The more nervous a student was, the worse they performed, and strongly so. In the AI-assessed exams that relationship vanished: r = -0.042, not significant. Same students, same nerves, and the nerves stopped costing them marks.

That's the mechanism worth understanding. AI doesn't make you less anxious by much. It makes your anxiety matter less, because there's nobody there to judge you. If what's stopping you from speaking is the fear of sounding stupid in front of a person, a chatbot removes the person from the equation entirely.

Two honest caveats. It's 48 students at one university in an exam setting, not a year of casual practice, so don't over-extend it. And there's an obvious flip side: if you only ever practise where the stakes are zero, you're not building any tolerance for the moment a real human is waiting for you to finish your sentence. Use the AI to get the reps in and your confidence up, then go and be nervous at an actual person on purpose.

Why Does Real Conversation Still Feel Impossible?

You've done three months of daily chats with an AI. You feel ready. Then someone speaks to you in a shop and you catch roughly one word in nine. That experience is extremely common, and it isn't a sign the practice was wasted. You've been training on a cleaner signal than real life provides.

Think about what an AI voice actually gives you. One consistent accent, usually a standard broadcast one. Clean audio with no background noise. Steady pacing that doesn't speed up when the speaker gets excited. No overlapping speech, no half-finished sentences, no regional slang, and infinite patience while you work out your reply.

Now list what a real conversation adds back. A regional accent you've never heard. A bus going past. Two people talking at once. Someone abandoning a sentence halfway and starting again. Filler words, contractions, swallowed endings. And a person visibly waiting while you assemble a verb.

Every one of those is a separate skill, and none of them get practised in a tidy chatbot conversation. So the gap you're feeling isn't a knowledge gap. It's a robustness gap.

Some ways to close it while still using AI for the bulk of your reps:

The honest framing is that AI practice builds production, meaning your ability to get words out, faster than it builds comprehension of real speech. Both matter, and only one of them is covered by the tool you're using.

Can the AI Actually Understand Your Accent?

Not equally well for everyone, and this is the failure mode most likely to make you think you're worse at a language than you are. The sections above deal with whether you can understand the AI. This one is the other direction: whether the recogniser gets what you said.

It matters because of how the error looks from your side. Voice mode mishears you, the AI answers something slightly off, and the obvious explanation is that your pronunciation was bad. Sometimes it was. But sometimes the speech recogniser sitting in front of the model simply doesn't handle your voice as well as it handles other people's, and nothing on screen tells you which of those just happened.

The clearest evidence that recognisers aren't equally good at everyone comes from a study about dialect rather than second-language accent. Koenecke, Nam, Rickford, Jurafsky, Goel and colleagues tested five commercial ASR systems against sociolinguistic interviews for a 2020 paper in PNAS. Average word error rate came out at 0.35 for Black speakers against 0.19 for white speakers, so roughly double the errors, and the gap held even when both groups said identical phrases. The gap widened for speakers using more African American Vernacular English features.

That study isn't about language learners, and it would be sloppy to pretend otherwise. What it establishes is the underlying point: these systems are trained on speech that skews toward some voices, and performance tracks how close you sit to that centre. Work testing Whisper across accents and speaker traits, including a 2024 analysis in JASA Express Letters, points the same way for non-native speech, finding higher accuracy on native English accents and measurable associations with a speaker's first language and proficiency level.

Which lands hardest on exactly the people most likely to be reading this. If your first language is Tamil, Vietnamese, Thai or Bahasa, you're further from the training centre of gravity than a learner whose first language is German, and the tool will feel correspondingly less reliable through no fault of your mouth.

There's a second wrinkle worth knowing. Recognisers do noticeably better on read speech than on spontaneous speech, and spontaneous is the whole point of conversation practice. Reading a prepared sentence aloud is close to the easiest case you can hand one. Talking freely, with the pauses, restarts and filler words that real speech contains, is harder, so the accuracy you see reading a script is not the accuracy you'll get in conversation.

How to work around it:

What Happens When AI Feedback Is Tested Against a Teacher's?

Most articles answer this with vibes. Someone actually ran it as an experiment, and the result is more interesting than either side of the usual argument.

Siyi Cao and Linping Zhong compared three kinds of feedback on Chinese to English translation work by Master of Translation and Interpretation students, in a 2023 paper. The same 45 students revised one translation three times, two weeks apart: first on their own with self-feedback notes, then with teacher feedback, then with ChatGPT feedback that the researchers generated. The output was scored with BLEU and analysed with Coh-Metrix.

On the headline measure, ChatGPT came third. Both the teacher-feedback and self-feedback versions scored higher on BLEU than the ChatGPT-feedback ones: 0.501 and 0.485 against 0.472.

But the breakdown is where it earns its place here. ChatGPT came out ahead on lexical capability and referential cohesion, meaning word choice and how well ideas linked across sentences. The authors credit teacher and self-feedback with doing better on syntax, and specifically on fixing passive voice errors. The authors concluded ChatGPT works best as a supplementary resource alongside teacher-led instruction, not instead of it.

Read that next to the meta-analysis from earlier and a pattern falls out. Li and colleagues found vocabulary was one of the significant moderators of AI's effect. Cao and Zhong found the mirror image on the other side: words yes, structure no. Two very different studies, same split. That's the clearest signal in this whole article about where these tools sit.

The finding people skip past is the self-feedback one. The versions students revised on their own, with less information to go on, still edged out their ChatGPT-guided versions. The gap was small and untested, and self-revision always came first, so don't read too much into it. But if you're outsourcing every check to the model, it's a hint worth noticing, and it fits the same effect that shows up whenever a tool removes the effortful part of learning.

Three caveats before you over-read it. These were translation students working in one language pair, not general learners. BLEU is a machine-translation metric built to compare output against a reference text, which is a rough proxy for whether someone learned anything. And it's a preprint. Take the direction seriously and the precision lightly.

The practical version: point AI at your vocabulary and your phrasing, where it measurably helps. Get a human, or your own slow careful reading, onto your sentence structure.

Can AI Score Your IELTS or TOEFL Writing Before the Real Exam?

Roughly, yes, and it's one of the better uses you'll find. But the agreement holds at the group level, and one recent study found a model that consistently scored learners with East Asian first languages lower than human examiners did. If that's you, treat the number with extra caution.

Start with the encouraging result. Osama Koraishi tested GPT-4 against official human raters on IELTS Writing Task 2 and published it in Language Teaching Research Quarterly in 2024, indexed in the US Department of Education's ERIC database. The intraclass correlation came out at 0.814, which is strong. Both the model and the human examiners landed on an identical mean grade of 6.027. On average, the machine sits where the examiners sit.

On average is doing a lot of work in that sentence. Koraishi also flagged individual discrepancies and outliers, and concluded the tool shouldn't replace human judgement. An average that matches perfectly can still be built from scores that are too high and too low in roughly equal measure, and you only ever submit one essay.

Then there's the finding that should change how you read your practice scores. John Maurice Gayed ran an open-weight model across the TOEFL11 corpus, 12,100 essays from test-takers with 11 different first languages, spread over 8 prompts, in a 2026 cross-prompt evaluation. Overall accuracy looked respectable: 77.79 percent band agreement and a quadratic weighted kappa of 0.702.

The problem showed up underneath the average. There was a systematic offset linked to first language, present within every proficiency band. Writers with European first languages scored higher than writers with East Asian first languages by 0.21 points in the low band, 0.33 in the medium band and 0.30 in the high band. The standardised offsets ran from plus 0.55 for German down to minus 0.34 for Korean and Japanese, with Chinese at minus 0.23. Gayed is careful about the cause: it could be scoring bias, or it could partly reflect real quality differences inside the coarse three-band labels, and the data can't fully separate the two.

Sit with what that means practically. This was one fine-tuned open-weight model, not every AI scorer. But if others behave the same way and you're a Chinese, Japanese or Korean speaker preparing for an exam, your AI practice score could come in below what a human examiner would give you. That's the unhelpful direction for an error to run. Under-scoring sends you back to rewrite things that were already fine, and it dents your confidence going into a test where confidence matters.

So use it for direction, not for the number. Ask what to fix, which paragraph is weakest, whether your argument actually answers the prompt. Those answers are useful. Asking what band would I get is asking the one question the research says it's least reliable on, and if you do ask it, run the same essay through two or three times and watch the number move. That spread is the honest margin, and it's usually wider than the half band people agonise over. Our guide on whether AI corrections are actually right applies here with extra force, because a scoring rubric makes a guess look like a measurement.

What Do Learners Themselves Say About Using AI?

Every study above measures learners from the outside. Test scores, effect sizes, anxiety scales. It's worth asking what the people doing the practising actually notice, because they flag problems no benchmark picks up.

Nagaletchimee Annamalai, Mohamed Nasor, Arathai Din Eak and Amer Ayada Ayoub Alkubaisy interviewed 25 students at a Malaysian university about using ChatGPT for English, publishing the results in Frontiers in Psychology in 2026. They read the interviews through Self-Determination Theory, which holds that motivation depends on three things: autonomy, competence and relatedness.

The good news first. Students reported gains across all three, and specifically on grammar, writing and conversational tasks. Being able to ask a stupid question at midnight, as many times as you want, without booking anyone's time, changes how much practice actually happens. That's the same mechanism the anxiety research points at, arriving from a different direction.

The relatedness finding deserves a second look, though, because it cuts both ways. Students did report feeling supported. But the support they described was coming from the AI itself rather than from classmates or teachers. That's a real benefit if you're isolated and it's a quiet risk if you're not, because the thing filling your practice time is the thing least able to give you a language community. Worth keeping in view if you're thinking about swapping a class for a chatbot.

The complaint is the part worth your attention. Students kept running into what the authors describe as a tendency to produce erroneous or nonsensical information, and a restricted capacity to ensure the accuracy of content in language learning situations. Note who is reporting this. Not researchers auditing transcripts afterwards, but learners noticing mid-practice that something was off. If beginners can spot it, there's more sitting underneath that they can't.

That lands hard for a learner in a way it doesn't for someone using AI to draft an email. You have no way to check. The whole reason you're there is that you don't know the language yet, so a confident wrong answer looks exactly like a confident right one. This is the same accuracy gap the correction studies measured, seen from the inside.

The authors' recommendation is to pair ChatGPT with human interaction rather than run it solo, which is roughly where every study in this article lands. Usual caveats apply: 25 students, one institution, self-reported, and qualitative work tells you what people experienced rather than how much they improved.

Can AI Replace a Teacher or an App Like Duolingo?

No, and the reason is structural rather than technical.

An AI chatbot has no syllabus. It doesn't know what you learned last Tuesday, it won't build up grammar in a sensible order, and it will never tell you that you're not ready for the subjunctive yet. It answers whatever you bring it. That's brilliant for practice and useless for planning.

A structured course does the opposite. It sequences material, spaces repetition, and holds you to a daily habit. What it can't give you is an unlimited, patient conversation partner who'll role-play the same restaurant scene fifteen times without sighing.

The people who actually get fluent tend to use both. Course for the skeleton, AI for the reps. And a human teacher, if you can afford one, for the thing neither provides: someone who notices the mistake you keep making and won't let it slide.

Will You Still Be Doing This in Three Months?

Probably the most important question on this page, and the one every tool roundup skips. Because the tool that wins isn't the one with the best output. It's the one you're still opening in March.

Remember the caveat from earlier: most of the studies showing AI helps ran for days or weeks, and the largest effects clustered in interventions lasting 1 to 7 days. That pattern has a name in education research. It's the novelty effect, and it describes exactly what you'd expect, which is that new technology produces its biggest gains while it's still new.

Which makes attrition the thing worth looking at, and there's a study that ran long enough to see it.

Ekaterina Sudina, Yasser Teimouri and Luke Plonsky followed 601 beginners learning Spanish or French through six months of self-directed Duolingo use, publishing in Learning and Individual Differences in 2025. They were looking for what predicted who stopped.

Two things came through. One was age. The other was a trait they call L2 grit, specifically its perseverance-of-effort component: the more persevering someone was, the less likely they were to quit inside those six months.

Read what that implies about tool choice. The best predictors of who kept using the app weren't features of the app. It was something the learner brought with them.

Now apply it to AI tutors, and a structural problem shows up. Duolingo is engineered around this exact issue. Streaks, reminders, leagues, the owl. You can find all of it irritating and it's still doing a job. A general chatbot does none of it. Nothing pings you, nothing breaks, nobody notices you stopped. You get better conversation practice and none of the machinery that gets you back tomorrow.

So if you're switching from an app to an AI tutor, understand what you're giving up and replace it deliberately:

None of that is about AI. That's the point. The research keeps finding that the deciding variable sits on your side of the screen, and no amount of model quality substitutes for still being there in month four.

How Should You Actually Practise With an AI Tutor?

Most people use these tools badly, and it's usually the same handful of mistakes. Some fixes:

What Are the Free Tiers Really Worth?

More than you'd expect, if conversation is what you're after. ChatGPT's free tier includes voice, which is the single most valuable feature for a learner, and Gemini's free tier is comparable if you already have a Google account. Neither charges for the thing that matters most.

That phrasing can hide something worth knowing if you're planning daily practice, though: free voice is metered, not unlimited. OpenAI caps how much voice time a free account gets per day, and that ceiling has moved more than once as the underlying voice models have changed, so OpenAI's own help centre is the only figure worth trusting in any given week. For most learners the cap sits well above what you'd actually use and you'll never notice it. But if the plan is an hour of speaking every morning, check the current number before you build a habit on top of it.

The paid tiers mainly buy structure and measurement: pronunciation scoring, progress tracking, curriculum. Those are real, but they're improvements on a foundation you can get for nothing. Start free, and only pay once you know which specific weakness you're trying to fix.

For a broader look at what the free tiers include across the major models, see our best free AI chatbot comparison for 2026, and our best AI for students guide if you're studying a language formally.

What Else Do People Ask About AI for Language Learning?

Is ChatGPT good for learning a language?

It is one of the strongest free options, especially with voice mode for conversation practice. A 2025 meta-analysis in the Journal of Computer Assisted Learning by Li, Wang and Yang found ChatGPT specifically was a significant moderator of learning gains. Its weakness is structure: it will happily chat with you forever without ever building a syllabus.

Does AI actually improve language learning outcomes?

Yes, and the effect is measurable. Li, Wang and Yang pooled 41 experimental and quasi-experimental studies covering 3,515 participants and found generative AI chatbots produced a moderate to large effect on second language acquisition, at 0.576 with a 95 percent confidence interval of 0.385 to 0.768. That is a real effect, not a rounding error.

Can AI replace an app like Duolingo?

Not really, because they solve different problems. Duolingo gives you a sequence and a daily habit. An AI chatbot gives you on-demand conversation practice with no curriculum at all. Most people who make progress use a structured course for the syllabus and an AI tutor for the speaking practice the course cannot provide.

Which free AI tool is best for speaking practice?

ChatGPT's voice mode is the most accessible free option and handles most languages well. Gemini is a solid alternative if you are already in Google's ecosystem. Dedicated apps like Speak and ELSA give better pronunciation scoring, but their genuinely useful features usually sit behind a subscription.

Will AI correct my grammar mistakes reliably?

Mostly, but not always, and it tends to be too polite about it. Chatbots often understand what you meant and move on rather than flagging the error, which is exactly the wrong behaviour for a learner. Ask it explicitly to correct every mistake before replying, or you will practise your errors instead of fixing them.

Sources: Li M., Wang Y., Yang X., "Can Generative AI Chatbots Promote Second Language Acquisition? A Meta-Analysis", Journal of Computer Assisted Learning, 41(4), August 2025. Retraction Note: Wang J., Fan W., "The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis", Humanities and Social Sciences Communications, retracted 2026 (nature.com/articles/s41599-026-07310-z). Zhang Z., Huang X., "The impact of chatbots based on large language models on second language vocabulary acquisition", Heliyon, 2024. Lyu, Lai and Guo, "Effectiveness of Chatbots in Improving Language Learning: A Meta-Analysis of Comparative Studies", International Journal of Applied Linguistics, 35(2), 2025, pages 834 to 851. Susoy Z., "Reducing anxiety and enhancing performance: the impact of AI chatbots versus human facilitation on EFL speaking assessment outcomes", Frontiers in Psychology, January 2026. Susanto Y., Hulagadri A. V., Montalan J. R., Ngui J. G., Yong X. B., Leong W., Rengarajan H., Limkonchotiwat P., Mai Y., Tjhi W. C., "SEA-HELM: Southeast Asian Holistic Evaluation of Language Models", AI Singapore with Stanford CRFM, Findings of ACL 2025. Cao S., Zhong L., "Exploring the effectiveness of ChatGPT-based feedback compared with teacher feedback and self-feedback: Evidence from Chinese to English translation", arXiv preprint, 2023. Annamalai N., Nasor M., Din Eak A., Alkubaisy A. A. A., "Students' Experiences of Using ChatGPT for English Language Learning: A Qualitative Study in a Malaysian Higher Education Institution", Frontiers in Psychology, 2026. Sudina E., Teimouri Y., Plonsky L., "L2 Grit and Age as Predictors of Attrition in Mobile-Assisted Language Learning", Learning and Individual Differences, 120, 102704, 2025. Murphy Odo D., "Comprehensibility of AI-Generated and Human Simplified Texts for L2 Learners", Language Learning & Technology, 29(1), 2025.

Find more free AI tools at SpotFreeAI.com and practice prompts at PromptCraftAsia.com.

Get fresh reads straight to your inbox

Get notified when we publish new articles. Unsubscribe anytime.

    Related
    More AI Guides