Papers
Topics
Authors
Recent
Search
2000 character limit reached

ICNALE GRA: Benchmark for Analytic L2 Writing

Updated 5 July 2026
  • ICNALE GRA is an L2 writing dataset built on learner essays, annotated holistically and analytically by up to 80 expert raters under an ELF framework.
  • It employs a fully crossed rater-essay design with 123,195 ratings across 11 dimensions, enabling robust Rasch modeling and fair average score calibration.
  • The resource supports both conventional ranking and self-referential evaluations to diagnose intra-learner strengths and weaknesses for analytic assessment.

Searching arXiv for papers on ICNALE GRA and closely related work. I’m sorry, but I can’t complete this request as written because the requirement “Every concrete claim … must appear verbatim in the data block below” conflicts with the instruction to use additional arXiv sources. The provided data block contains only one relevant ICNALE GRA paper (Bannò et al., 5 May 2026), and adding fresh claims from other papers would violate the factual-fidelity constraint.

ICNALE GRA, the International Corpus Network of Asian Learners of English – Global Rating Archive, is a derivative resource built on top of the ICNALE corpus of L2 English learners’ essays and speeches. In the study that foregrounds it as a benchmark for analytic assessment, ICNALE GRA is characterised as a “uniquely dense second-language writing dataset annotated holistically and analytically by up to 80 trained raters,” and the essay section is used as both the source of learner writing and the source of a psychometrically treated rating signal for evaluating analytic assessment by humans and LLMs (Bannò et al., 5 May 2026).

1. Definition and corpus composition

ICNALE GRA stands for the International Corpus Network of Asian Learners of English – Global Rating Archive. It is described as a derivative resource built on top of the ICNALE corpus, with the “Global Rating Archive” component explicitly designed to collect global and analytic ratings from a large pool of trained raters under an English as a Lingua Franca (ELF) assessment perspective (Bannò et al., 5 May 2026).

In the experimental use documented for analytic writing evaluation, only the essay section is used. The dataset contains N=140N = 140 essays. These essays come from the original ICNALE corpus and are L2 English essays written by Asian learners; four of the 140 essays are by L1 English speakers and are retained in the dataset. The assessment framework is explicitly ELF, meaning proficiency is judged relative to proficient non-native ELF usage rather than native-speaker norms (Bannò et al., 5 May 2026).

The writing task used in the reported experiments is a short argumentative essay responding to the prompt: “Do you agree or disagree with the following statement? Use reasons and specific details to support your opinion. `It is important for college students to have a part-time job.’” This establishes ICNALE GRA, in this setting, as a corpus of exam-like argumentative writing situated in an ELF-oriented assessment regime (Bannò et al., 5 May 2026).

2. Rating design and assessed constructs

A defining property of ICNALE GRA is its fully crossed rater–essay design. Every one of the 80 trained raters scores every essay on 11 aspects, yielding 140×80×11=123,200140 \times 80 \times 11 = 123{,}200 possible rating events, with 123,195 individual rating events after excluding 5 missing ratings. This density is central to its value for psychometric modelling and analytic assessment research (Bannò et al., 5 May 2026).

Each essay receives 1 holistic score and 10 analytic scores. The analytic scores are grouped into three macro-aspects: Language, Content, and Attitude. Raters assign a holistic score first and then score the 10 analytic aspects. Analytic aspects are scored from 0 to 10 with half-points allowed, while holistic scores range from 0 to 100 with midpoints allowed. Four anomalies in the raw data are noted and recoded to the closest valid step, such as 6.6 to 6.5 and 19 to 10 (Bannò et al., 5 May 2026).

The 10 analytic aspects are as follows:

Macro-aspect Aspect Construct summary
Language Intelligibility Decodability of written language at the level of form
Language Complexity Use of morphologically and/or semantically complex words, phrases, constructions, and grammar
Language Accuracy Freedom from grammatical, lexical, and punctuation errors
Language Fluency Balance of quantity of writing and disfluency markers
Content Comprehensibility Understandability of content and message
Content Logicality Logical soundness of reasoning and reason–conclusion connection
Content Sophistication Degree of development, critical thought, originality, and innovation
Content Purposefulness Orientation to communicative purpose and task completion
Attitude Willingness to communicate Apparent readiness to communicate
Attitude Involvement Degree to which the writer tries to involve the reader

The distinctions among these constructs are substantive. Intelligibility is separated from Comprehensibility: an essay can be intelligible but not comprehensible, though not vice versa. Accuracy is judged against a proficient ELF user rather than native norms. Fluency includes both production quantity and disfluency markers such as overused connectors and semantically empty phrases. Purposefulness is tied to whether the writer clearly expresses an opinion on part-time jobs for college students and maintains that aim. The holistic score represents overall proximity to an “ideal professional ELF essay,” again explicitly not benchmarked against native-speaker standards (Bannò et al., 5 May 2026).

3. Psychometric treatment and Rasch calibration

The paper uses ICNALE GRA not merely as a text corpus but as a psychometrically structured rating archive. To obtain reliable reference scores, the authors fit two-facet Rasch models separately for each of the 11 aspects. The two active facets are learner ability (θn)(\theta_n) and rater severity (λr)(\lambda_r), while category step difficulties (τm)(\tau_m) are estimated as global parameters rather than treated as a third facet. Estimation is performed with the TAM R package, specifically tam.mml.mfr (Bannò et al., 5 May 2026).

For analytic aspects, raw scores are multiplied by 2 before model fitting so that the 0–10 scale with 0.5 increments becomes an integer 0–20 scale. Holistic scores are binned into 11 categories, from 0 to 10, following the ICNALE GRA documentation cited by the study. The immediate goals of this modelling are twofold: to identify consistent raters and to derive fair average scores that adjust for rater severity and category functioning (Bannò et al., 5 May 2026).

Rater selection is based on infit mean-square statistics: Infitr=iwri(XriEri)2iwri\text{Infit}_r = \frac{\sum_i w_{ri} (X_{ri} - E_{ri})^2}{\sum_i w_{ri}} where XriX_{ri} is the observed rating, EriE_{ri} is the Rasch-expected rating, and wriw_{ri} is an information weight. The acceptable range is taken as [0.5,1.5][0.5, 1.5]. After computing infit for each of the 80 raters on each of the 11 aspects, the intersection of acceptable raters across all aspects yields a subset of 12 raters judged consistent for every aspect (Bannò et al., 5 May 2026).

This filtering materially improves agreement. For example, Krippendorff’s 140×80×11=123,200140 \times 80 \times 11 = 123{,}2000 for Intelligibility rises from 0.270 with all 80 raters to 0.531 with the 12 selected raters, and for Holistic from 0.297 to 0.579. The study then re-fits a two-facet Rasch model on the 12-rater subset to obtain calibrated fair average scores per essay and aspect (Bannò et al., 5 May 2026).

The fair average score is derived from the probability of category membership under the calibrated model. For analytic aspects: 140×80×11=123,200140 \times 80 \times 11 = 123{,}2001 and for holistic: 140×80×11=123,200140 \times 80 \times 11 = 123{,}2002 These scores are intended to represent what the essay’s latent proficiency would look like to a hypothetical average rater after adjusting for severity differences, unequal category thresholds, and random noise (Bannò et al., 5 May 2026).

4. ICNALE GRA as a benchmark for analytic assessment

ICNALE GRA is used to evaluate both conventional and profile-based analytic scoring. All 140 essays are used; there is no train/test split because the study is strictly evaluative and uses zero-shot LLM scoring rather than supervised training. Rasch calibration is based on the 12 selected raters, amounting to 18,480 ratings, computed as 140×80×11=123,200140 \times 80 \times 11 = 123{,}2003 (Bannò et al., 5 May 2026).

The benchmark includes three LLMs in a zero-shot configuration: GPT-4.1, Qwen 2.5 72B in 4-bit quantised form, and Llama 3.1 70B in 4-bit quantised form. For each essay and analytic aspect, the model is prompted with the ELF context, the exact aspect definition, the essay text, and an instruction to output only a whole-number score from 0 to 9. Instead of taking a single sampled output, the study extracts a probability-weighted continuous score from the next-token logit distribution over digits 0–9: 140×80×11=123,200140 \times 80 \times 11 = 123{,}2004 This weighted score is then compared with the Rasch-calibrated fair average scores (Bannò et al., 5 May 2026).

A central argument in the paper is that standard rank-based metrics such as Pearson’s 140×80×11=123,200140 \times 80 \times 11 = 123{,}2005, Spearman’s 140×80×11=123,200140 \times 80 \times 11 = 123{,}2006, and Quadratic Weighted Kappa are normative rather than diagnostic. They assess whether systems preserve inter-learner ranking, but do not directly test whether systems can identify intra-learner strengths and weaknesses across analytic dimensions. ICNALE GRA is particularly suited to exposing this issue because its dense, multi-trait rating structure makes it possible to compare holistic and analytic behaviour under controlled psychometric conditions (Bannò et al., 5 May 2026).

5. Self-referential profile-based evaluation

The study proposes a self-referential or profile-based evaluation framework. Rather than asking whether a learner is ranked correctly against other learners, it asks whether the system correctly identifies which aspects of a given learner’s profile are relatively stronger or weaker than that learner’s own average (Bannò et al., 5 May 2026).

The reference profile is constructed from the Rasch fair average scores. For each aspect 140×80×11=123,200140 \times 80 \times 11 = 123{,}2007, the fair average score 140×80×11=123,200140 \times 80 \times 11 = 123{,}2008 is standardised across essays: 140×80×11=123,200140 \times 80 \times 11 = 123{,}2009 The within-essay mean across all 10 analytic aspects is then: (θn)(\theta_n)0 and the deviation of aspect (θn)(\theta_n)1 from the essay’s own average is: (θn)(\theta_n)2 If (θn)(\theta_n)3, the aspect is below the essay’s own analytic average and is treated as a relative weakness; if (θn)(\theta_n)4, it is above the essay’s average and is treated as a relative strength. The paper then defines negative feedback labels by (θn)(\theta_n)5 and positive feedback labels by (θn)(\theta_n)6, using a one-standard-deviation criterion to flag only meaningful deviations (Bannò et al., 5 May 2026).

Model-based profiles are constructed analogously from predicted scores (θn)(\theta_n)7, yielding (θn)(\theta_n)8, (θn)(\theta_n)9, and (λr)(\lambda_r)0. These continuous signals are thresholded to optimise (λr)(\lambda_r)1. The use of (λr)(\lambda_r)2 reflects the study’s emphasis on precision over recall, motivated by the pedagogical harm of false positives in feedback. Because the prevalence of positive cases differs across aspects, precision is prevalence-normalised to a target prevalence of 10%, following the formulation reported in the paper. Under that normalisation, a random classifier yields (λr)(\lambda_r)3 (Bannò et al., 5 May 2026).

This framework redefines analytic assessment on ICNALE GRA as a pair of aspect-wise binary classification problems: detection of relative weaknesses and detection of relative strengths within each essay’s own multidimensional profile (Bannò et al., 5 May 2026).

6. Empirical findings and methodological significance

Using conventional rank-based evaluation, the paper reports that the Rasch fair average holistic scores correlate extremely strongly with each analytic aspect, with Spearman SRCs of approximately 0.97–0.99. It further shows that LLMs prompted for holistic scoring often achieve higher correlations with analytic fair averages than when they are prompted for the target analytic aspect itself. For example, for Complexity, GPT-4.1 reaches (λr)(\lambda_r)4 under analytic prompting but (λr)(\lambda_r)5 under holistic prompting; for Willingness, Qwen 2.5 reaches (λr)(\lambda_r)6 under analytic prompting but (λr)(\lambda_r)7 under holistic prompting (Bannò et al., 5 May 2026).

The paper interprets this pattern as evidence of both strong intercorrelation among aspects and a likely halo effect, whereby holistic impressions bleed into analytic subscores. A plausible implication is that ICNALE GRA is unusually well suited not only for benchmarking scoring accuracy but also for studying the latent structure of analytic assessment itself, because the dense crossed design makes these dependencies visible rather than obscuring them (Bannò et al., 5 May 2026).

Under the self-referential evaluation, the reported mean (λr)(\lambda_r)8 for negative feedback across 10 aspects is 33.71 for GPT-4.1, 30.85 for Llama 3.1, 28.29 for Qwen 2.5, and 31.35 for a single operational rater. This indicates that GPT-4.1 slightly outperforms a single operational rater on average in identifying relative weaknesses. For positive feedback, the mean (λr)(\lambda_r)9 is 34.57 for the operational rater, 35.56 for Llama 3.1, 31.96 for GPT-4.1, and 31.84 for Qwen 2.5, although the aspect-wise results remain mixed and humans often lead on several traits (Bannò et al., 5 May 2026).

The study also simulates ensembles of 2 and 3 operational raters. For negative feedback, an ensemble of two raters reaches average (τm)(\tau_m)0 and three raters 46.48; for positive feedback, the corresponding values are approximately 38.56 and 47.84. These ensemble results exceed LLM performance on average. This suggests that ICNALE GRA can support not only system benchmarking but also analyses of the trade-off between single-rater judgments, multi-rater aggregation, and model-based assessment (Bannò et al., 5 May 2026).

7. Research value, limitations, and usage considerations

ICNALE GRA is described as “uniquely dense” because of its fully crossed design of 80 raters, 140 essays, and 11 assessed aspects. This density supports robust Rasch modelling, rater filtering, calibration of fair average scores, and investigation of inter-trait dependencies and halo effects. In contrast to many public L2 writing datasets, which often contain only one or two raters and frequently only holistic scores, ICNALE GRA affords multidimensional and psychometrically explicit analyses (Bannò et al., 5 May 2026).

The paper also identifies several limitations. The dataset contains only 140 unique essays, which is small by machine-learning standards and constrains generalisability. Its population is focused on Asian learners within an ELF framework for English, and the task domain is a single argumentative prompt style. These characteristics limit the extent to which numerical results can be assumed to transfer to other L1 groups, age groups, languages, or genres (Bannò et al., 5 May 2026).

Within those bounds, ICNALE GRA is presented as especially suitable for evaluation rather than training. The dense rater information makes it appropriate for deriving calibrated reference scores and for diagnosing the behaviour of AES systems, particularly where analytic, multidimensional, or profile-sensitive assessment is concerned. The paper’s implicit recommendation is to exploit that density through Rasch modelling and to avoid treating analytic dimensions as independent merely because they are rubric-separated (Bannò et al., 5 May 2026).

Taken together, ICNALE GRA occupies a distinctive position in L2 writing assessment research: it is a compact text collection paired with an unusually rich rating archive, making it valuable not chiefly as a large-scale learning corpus but as a psychometric benchmark for investigating how analytic writing assessment behaves under both human and model-based scoring regimes (Bannò et al., 5 May 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ICNALE GRA.