LingoQ: AI-Powered ESL Workplace Practice
- LingoQ is an AI-mediated system that converts workplace English interactions into personalized quizzes using real work queries.
- It employs a desktop-to-cloud-to-mobile architecture that logs queries, applies LLM filters, and delivers context-based quizzes on smartphones.
- The system enhances language learning by connecting work-related tasks with tailored review sessions, leading to measurable gains and improved self-efficacy.
LingoQ is an AI-mediated system for workplace English as a second language practice that converts users’ English-related interactions with LLMs into personalized quizzes. It is designed for non-native English-speaking information workers who already rely on LLM assistants for translation, lookup, proofreading, and related tasks, but whose day-to-day assistance use does not necessarily translate into sustained language learning. LingoQ links these activities by capturing work queries, generating review materials from them, and delivering short quizzes on smartphones. In a three-week deployment study with 28 ESL workers, participants valued the relevance of quizzes that reflected their own context, engaged with the system throughout the study, and showed improved self-efficacy together with learning gains for beginners and, potentially, for intermediate learners (Yang et al., 22 Sep 2025).
1. Problem setting and design objective
LingoQ addresses a specific mismatch between workplace language use and conventional ESL study. Non-native English speakers performing English-related tasks at work struggle to sustain ESL learning, despite their motivation, and study materials are often disconnected from their work context. At the same time, workers rely on LLM assistants to address immediate needs, yet these interactions may not directly contribute to their English skills. LingoQ responds to this gap by allowing workers to practice English using quizzes generated from their LLM queries during work (Yang et al., 22 Sep 2025).
The system’s defining premise is that work queries are not merely task-support interactions but potential learning material. Rather than introducing a separate curriculum detached from the workday, LingoQ uses the user’s own requests as the source of practice content. This places relevance at the center of the system: the quiz item is not an externally authored exercise but a transformed trace of situated workplace activity. A plausible implication is that LingoQ reframes LLM assistance from a purely compensatory tool into a mechanism for retrieval and review.
2. System architecture
LingoQ comprises three tightly integrated components: a desktop interface for English-related work queries, a backend pipeline that generates and curates questions, and a mobile interface for review and practice (Yang et al., 22 Sep 2025).
| Component | Role | Functions |
|---|---|---|
| LingoQuery (Desktop App) | LLM-powered chatbot interface running on users’ computers | Handles English-related work queries; offers smart input shortcuts and intent-specific prompts for “lookup,” “translate,” and “proofread”; captures selected text and a screenshot; lets users mark responses for review |
| Backend Pipeline (Cloud Service) | Processing and personalization layer | Receives conversations and context; processes, filters, and transforms user-English interaction data into validated quiz questions; selects and curates quizzes personalized to each user |
| LingoQuiz (Mobile App) | Smartphone review interface | Presents short, bite-sized quizzes, typically 10 items per session; supports review at any time from the user’s own work queries |
LingoQuery is described as similar to ChatGPT, but specialized for English-related work queries. Its context-capture functionality is central to the system design: selected text and screenshots can be attached to interactions, and marked assistant responses influence quiz generation priority. LingoQuiz, by contrast, is optimized for micro-moments, delivering short review sessions on smartphones. The architecture therefore separates task execution and later review while preserving continuity between them.
This desktop-to-cloud-to-mobile arrangement operationalizes a specific pedagogical loop. Work generates data; the backend turns that data into validated exercises; the smartphone app makes later retrieval possible outside the original work context. The significance of the design lies in its attempt to preserve contextual specificity while changing time and device.
3. Personalized quiz generation pipeline
The quiz generation pipeline begins with logging all English-related queries and assistant responses from LingoQuery. These are tagged by intent—“lookup,” “translate,” “proofread,” or “text”—and may be enriched with screenshots and inferred work context. Users can also mark specific assistant responses to flag especially relevant query-answer pairs for future quizzes (Yang et al., 22 Sep 2025).
Question generation proceeds in three stages. First, an LLM-based filter determines whether a query is English-related. Additional context is extracted from screenshots or application metadata via GPT-4o image understanding. For each eligible query-response pair, the system generates two multiple-choice, fill-in-the-blank questions using prompts designed after standardized tests such as TOEFL, TOEIC, and GRE. Each generated question is structured in JSON with a stem, key, distractors, explanation, and rationale. The context from screenshots or text is used so that the questions remain directly rooted in the user’s actual work scenario.
Second, LingoQ applies automated quality evaluation and refinement. Generated questions are assessed by an LLM-powered module on two criteria: answerability, defined as whether the question is unambiguous and includes a correct, unique answer, and proficiency, defined as whether it appropriately challenges the learner without being trivial. Questions that fail are revised in a feedback-refinement loop up to two times, and items failing after three iterations are discarded.
Third, validated questions are stored in a per-user question pool and assembled into quiz sessions of 10 items: 7 new questions and 3 from previously attempted items. Selection is weighted for low repetition, wrong answers, prior marking, or recent inactivity. Incorrectly answered questions reappear within sessions, and items re-emerge in future sessions. The system therefore combines contextual generation with re-exposure logic rather than treating each work interaction as a one-off exercise.
The technical approach extends beyond quiz generation alone. Intent classification is handled by an LLM classifier that assigns a query type from the four categories, and response generation in LingoQuery is tailored to intent: “lookup” yields a dictionary-style entry, “translate” yields side-by-side native and translated text with rationale, and “proofread” yields highlighted revision output with rationale. This makes the source interactions more structured before they become learning material.
4. Deployment study and usage patterns
The reported evaluation was a three-week deployment study with 28 information workers in South Korea, aged 25–48, from fields including IT, healthcare, business, and academics. Participants ranged from CEFR A1 to C1, including 3 beginners at CEFR A1 and 7 advanced users at C1. The study procedure comprised a pre-study survey and proficiency test, three weeks of regular LingoQ use, and a post-study survey and proficiency test (Yang et al., 22 Sep 2025).
Usage statistics indicate sustained interaction with both the work-facing and review-facing components. In LingoQuery, participants submitted 3,325 messages, approximately 119 per participant, and used the app on 13.2 days per participant on average out of 21 days. Query types were distributed as translate (38.2%), lookup (12%), proofread (8.6%), with the remainder plain text or follow-ups. Participants marked assistant responses 13.4 times on average per user. The context-capture shortcut was used for 6.9% of queries.
In LingoQuiz, 2,708 English-language queries produced 5,682 generated questions, of which 3,290 survived quality checks, approximately 118 per user. Participants solved 7,155 questions in total, including repeats, or approximately 256 per participant. The average number of quiz days was 13.4 per user, closely matching LingoQuery usage. Each quiz took approximately 9.3 minutes; most sessions were completed fully; and users averaged 1.04 quizzes per day, above the minimum required rate of 0.5 per day. Repeat exposure covered 28% of questions, and accuracy increased with repeats from 82.6% to 89.5% and then 92.4%.
Temporal usage patterns are also reported. LingoQuery peaked during workdays, especially in the late afternoon, whereas LingoQuiz usage was highest in the evenings. This pattern is consistent with the system’s intended division between work-time assistance and later review. The study therefore documents not only raw engagement volume but a diurnal separation between query production and quiz consumption.
5. Learning outcomes, self-efficacy, and perceived relevance
The study reports statistically significant overall improvement on the English proficiency test, with an average increase of +1 point and . The strongest gains were observed for CEFR A participants, who improved by +4 points on a 28-item test. No significant change was found for CEFR B or C overall, but among CEFR B participants, higher system usage was associated with more gain. The paper therefore characterizes the results as showing learning gains for beginners and, potentially, for intermediate learners (Yang et al., 22 Sep 2025).
Self-efficacy also improved significantly. On QESE, overall pre-post change was significant with and , and both reading and writing subscales improved. Qualitative feedback linked these changes to increased confidence in handling work English, especially among beginners and intermediate learners. In parallel, participants rated LingoQ quizzes as much more relevant to actual work than standard ESL apps or courses, with Wilcoxon and . LingoQ was also perceived as more helpful for practical on-the-job English skill, with and , and more sustainable for long-term study, with and ; 86% indicated willingness to continue use.
A 30-sample manual review of generated questions further evaluated the automated quality filter. Compared to experts, automated question evaluation achieved Answerability and Proficiency 0. Reviewers noted that domain-specificity occasionally challenged experts, and some generated items were considered equivalent or superior to established TOEIC- or TOEFL-style items. The system’s context-anchored stems were identified as a key differentiator from generic ESL question banks.
The paper also discusses the “Noticing Hypothesis” in relation to observed behavior. Knowing that queries would be turned into quizzes led users to pay closer attention and reflect on English gaps, shifting the act of querying from passive task-solving to an integrated learning opportunity. This suggests that LingoQ’s effects may derive not only from later retrieval practice but also from altered attention at the time of work.
6. Limitations, extensions, and broader significance
Several constraints and future directions are explicit in the reported findings. A small subset of participants experienced after-hours review of work-specific material as “bringing work home,” which could deter engagement outside work. Privacy of work content and the desirability of after-hours detachment are therefore important considerations for workplace deployment. The architecture is also described as extensible: question formats could be adapted, speaking and listening could be integrated, and the same pipeline could be applied to other skills such as technical writing, coding, or legal research (Yang et al., 22 Sep 2025).
The broader significance claimed for LingoQ lies in “bridging tools & learning.” It integrates two prevalent behaviors—LLM assistance at work and smartphone language study—into a single loop in which work activity produces later practice material. The system is presented as a way to transform potential learning losses from LLM convenience into personalized learning opportunities while surfacing knowledge gaps, boosting retention, and increasing worker confidence. The paper further argues that the context-anchored automated pipeline supports “English for Specific Purposes” for any profession and that the general interaction pattern is transferable to other knowledge-intensive professions and other languages, assuming reasonable LLM support.
Taken together, these claims locate LingoQ within a broader shift in AI-mediated language learning: from generic content delivery toward contextually grounded, work-derived practice. Its central contribution is not a new quiz format alone, but a mechanism for converting situated LLM use into repeated, personalized review. A plausible implication is that LingoQ represents a model of language learning embedded in workflow rather than appended to it.