SciTrek: Long-Context QA for Scientific Articles
- SciTrek is a benchmark for long-context question answering, focusing on extracting and reasoning over metadata distributed across multiple scientific articles.
- It employs an SQL-backed reasoning framework to generate explicit reasoning paths, enabling tasks like counting, sorting, and filtering citations and authors.
- Built from clusters of scientific articles across diverse subjects, it mimics real scientific workflows and challenges retrieval-heavy models with structured queries.
Searching arXiv for the specified paper to ground the article in the cited source. SciTrek, styled in the source paper as \textsf{SciTrek}, is a benchmark for long-context question answering over scientific articles that evaluates whether LLMs can locate, combine, filter, count, sort, and reason over metadata distributed across multiple full-text papers rather than merely retrieve a sentence from a long prompt. Its central methodological feature is that every question is derived from an SQL query executed over a database constructed from article metadata, yielding automatic answer generation, explicit reasoning structure, and fine-grained diagnosability. The benchmark is introduced in "Who Gets Cited Most? Benchmarking Long-Context LLMs on Scientific Articles" (Li et al., 25 Sep 2025).
1. Scope and motivation
SciTrek was introduced to address four limitations that the authors identify in prior long-context benchmarks: frequent reliance on non-scientific text, overemphasis on retrieval tasks such as "needle-in-a-haystack" search, lack of explicit reasoning structure, and limited scalability because of manual annotation requirements (Li et al., 25 Sep 2025). In response, SciTrek uses real scientific articles as natural contexts, requires multiple information-processing skills rather than retrieval alone, exposes an explicit reasoning path through SQL programs, and supports automatic generation over article collections of arbitrary length.
The benchmark is motivated by the structure of actual scientific workflows. The relevant tasks involve reading full papers rather than snippets, tracking authors, references, citations, titles, and cross-paper relations, aggregating evidence across many documents, and identifying sparse facts distributed over long inputs. The questions are described as "relatively superficial" because they focus on metadata rather than deep scientific content, but this is deliberate: the benchmark is intended to test prerequisite capabilities such as counting authors, identifying cited papers, sorting articles by metadata, handling negation, and following citation or authorship relations across documents.
A common misconception is that long-context competence can be established by retrieval-heavy evaluations alone. SciTrek is designed around the opposing view that structured, compositional operations over natural long contexts are a distinct challenge. The paper therefore frames failures in long-range information localization, aggregation over many documents, counting and comparison, sorting and ranking, compound logical conditions, negation handling, citation/reference tracking, and multi-hop relational reasoning as central failure modes rather than edge cases.
2. Corpus construction and context design
SciTrek draws its articles from Semantic Scholar. The construction process begins with seed papers from eight subjects: Computer Science, Economics, Electronic Engineering, Math, Physics, Biology, Finance, and Statistics. For each subject, the authors choose two seed articles, and each seed has more than 100 citations since 2020. This produces 16 seed-based clusters (Li et al., 25 Sep 2025).
For each seed article, related articles are retrieved from Semantic Scholar, two-hop related articles are included using the citation graph, and the collection process randomly samples 10 first-hop related articles and 5 second-hop related articles for each first-hop article. Only articles with available PDFs are retained. This yields 16 article clusters and 662 scientific articles with PDFs. The PDFs are converted to markdown using Marker.
From these 662 articles, the benchmark creates 2,612 article collections across different lengths: 2,027 via random sampling and 585 via graph traversal. The graph-based expansion uses depth-first and breadth-first search in order to preserve citation structure. Each collection contains at least four articles, no two collections share more than half of their articles, and contexts are constructed or truncated to target token budgets of 64K, 128K, 512K, and 1M. The authors note that contexts beyond 1M tokens could also be constructed.
Each benchmark instance consists of a context formed by concatenating multiple full-text scientific articles, a natural-language question, and a short factual answer. The evaluation prompt instructs the model to output only the final answer, to use comma-separated items if multiple values are required, and to return NULL if it cannot find the answer. The answers are precise, verifiable outputs rather than summaries: integers, lists of integers, author names, article titles, sorted sequences, or aggregate values.
3. SQL-backed reasoning framework
For each article collection, SciTrek constructs a relational database with three tables. The articles table contains article_id, article_title, title_word_count, author_count, and reference_count. The article-author table contains relation_id, article_id, author_name, and author_position, where author_position starts from 0 for the first author. The citing-cited table contains relation_id, article_id_citing, and article_id_cited (Li et al., 25 Sep 2025). These tables encode titles, title word counts, authors and author order, reference counts, and citation links among articles.
The SQL layer is the benchmark’s symbolic reasoning backbone. The supported commands include MAX, MIN, SUM, AVG, COUNT, DISTINCT, ORDER BY, ASC, DESC, GROUP BY, and WHERE. The supported operators include comparison operators such as =, >, <, >=, <=, <>, and LIKE; arithmetic operators +, -, *, /, and %; and logical operators AND, NOT, OR, BETWEEN, and IN. The manually designed query templates cover Author Count, Author List, Reference Count, Title List, Title Word Count, Author Relation, and Citation Relation.
For each collection, 10 templates are randomly selected, placeholders are instantiated with collection-specific values, the resulting queries are executed against the database, and the execution results become the gold answers. Representative templates include SELECT MAX(author_count) FROM articles, SELECT title_word_count FROM articles ORDER BY author_count ASC, SELECT author_name FROM article_author WHERE author_position = {author-position}, SELECT SUM(title_word_count) FROM articles WHERE reference_count = {reference-count}, SELECT author_count FROM articles WHERE title_word_count % 2 = 1 ORDER BY title_word_count DESC, and a relational filtering query that counts articles that are cited by others but do not cite any other articles. The explicit SQL programs make it possible to decompose each question into record identification, filtering, joining or relating tables, aggregation or sorting, and final extraction.
To present the benchmark as ordinary question answering rather than SQL execution, the authors use Qwen2.5-Coder-32B-Instruct to translate SQL queries into natural-language questions. The conversion is validated by a round-trip procedure: converting SQL to natural language, converting the question back into SQL, and checking whether the regenerated SQL returns the same execution result on the database. This loop is repeated up to 10 times per SQL-collection pair, and pairs for which no valid question is obtained are discarded. The reported success rate for this pipeline is 82.9%.
4. Benchmark composition, skills, and validation
SciTrek has a test set of 2,121 question-answer pairs and a training set of 19,543 question-answer pairs, with the training set mainly used for post-training experiments (Li et al., 25 Sep 2025). The benchmark uses four context-length tiers: 64K, 128K, 512K, and 1M. In the test set, the 64K tier contains 112 instances with an average of 4.2 articles per context, average question length 14.6 words, and average answer length 6.2 words. The 128K tier contains 728 instances with averages 5.7, 16.9, and 7.9. The 512K tier contains 667 instances with averages 22.7, 18.1, and 30.9. The 1M tier contains 614 instances with averages 46.3, 18.6, and 63.6. As contexts become longer, the number of articles per context rises sharply and answer length also rises.
The benchmark’s skill categories are Aggregating, Sorting, Filtering, Filtering + Aggregating, Filtering + Sorting, and Relational Filtering. The number of SQL templates per skill is reported as 20, 27, 107, 107, 106, and 20 respectively. This distribution indicates that the benchmark is dominated by filtering-based compositional questions.
Coverage of SQL commands and operators is broad across the test set. The paper reports SELECT in 2,121 instances, or 100%; WHERE in 1,821 instances, or 85.86%; = in 1,032 instances, or 48.66%; IN in 682 instances, or 32.15%; OR in 616 instances, or 29.04%; ORDER BY in 591 instances, or 27.86%; < in 476 instances, or 22.44%; > in 450 instances, or 21.22%; COUNT in 370 instances, or 17.44%; DISTINCT in 280 instances, or 13.20%; MAX in 213 instances, or 10.04%; AND in 195 instances, or 9.19%; GROUP BY in 175 instances, or 8.25%; NOT in 154 instances, or 7.26%; AVG in 135 instances, or 6.36%; MIN in 133 instances, or 6.27%; SUM in 102 instances, or 4.81%; and LIKE in 50 instances, or 2.36%. This distribution is consistent with the benchmark’s emphasis on composition rather than isolated lookup.
Question-answer quality is assessed through a crowd study on 120 randomly sampled instances, with 20 per skill and 3 annotators per instance. Annotators answer using the database tables rather than 1M-token full texts. The reported overall annotator agreement is 88.3%, and alignment of SQL-executed answers with human-majority answers is 83.3%. By skill, alignment ranges from 73.7% for Relational Filtering to 89.5% for Filtering + Sorting. This suggests that the automatic generation pipeline generally preserves the intended SQL semantics, while also indicating that the mapping from SQL to natural language is not flawless.
5. Experimental evaluation and empirical findings
The evaluation covers open-weight models—Qwen2.5-7B-Instruct-1M, Qwen2.5-14B-Instruct-1M, Gemma-3-27B-IT, Llama-4-Scout-17Bx16E-Instruct, Llama-3.3-70B-Instruct, and DeepSeek-R1-Distill-Llama-70B—and proprietary models—Gemini 2.5 Pro, GPT-4.1, and o4-mini. All models are instruction-tuned, the same prompting protocol is used across models, each model generates 3 answers per question, and results are reported as average Exact Match (EM) and F1 (Li et al., 25 Sep 2025). Two context settings are evaluated: full-text scientific articles, which is the main benchmark setting, and database tables, a diagnostic setting whose contexts average about 2K tokens.
On full-text contexts, zero-shot EM is low and declines sharply with context length for nearly all models. Qwen2.5-7B-Instruct-1M scores 4.5 at 64K, 2.8 at 128K, 0.3 at 512K, and 0.0 at 1M. Qwen2.5-14B-Instruct-1M scores 8.3, 6.5, 1.6, and 0.1. Llama-4-Scout-17Bx16E-Instruct scores 5.4, 2.8, 1.3, and 1.1. GPT-4.1 scores 21.1, 11.7, 3.9, and 2.5. Gemini 2.5 Pro scores 41.7 at 64K and 26.0 at 128K, while o4-mini scores 61.0 at 64K and 46.5 at 128K. The central empirical trend is collapse with length: examples explicitly noted in the paper include 8.3 to 6.5 to 1.6 to 0.1 for Qwen2.5-14B, 21.1 to 11.7 to 3.9 to 2.5 for GPT-4.1, and 4.5 to 2.8 to 0.3 to 0.0 for Qwen2.5-7B.
When the same problems are presented as database tables, models perform much better, but performance still degrades as the underlying article collection grows. On table contexts, Qwen2.5-14B-Instruct-1M reaches 33.3, 27.2, 11.0, and 5.9 across 64K, 128K, 512K, and 1M. Llama-3.3-70B-Instruct reaches 47.0, 36.2, 14.6, and 8.1. DeepSeek-R1-Distill-Llama-70B reaches 83.3, 74.2, 56.8, and 42.2. Gemini 2.5 Pro reaches 91.7, 83.5, 55.4, and 31.5. o4-mini reaches 95.2, 87.8, 79.4, and 72.6. F1 follows the same pattern: on full text, Gemini 2.5 Pro reaches 58.1 at 64K and 48.8 at 128K, while GPT-4.1 reaches 36.0, 29.7, 22.3, and 19.6; on tables, o4-mini reaches 97.2, 89.2, 87.0, and 86.5, and Gemini 2.5 Pro reaches 95.3, 88.3, 79.0, and 69.8.
The contrast between full-text and table settings indicates that failure is not reducible to one factor. A substantial part of the difficulty lies in locating sparse metadata in long article text, maintaining accurate counting across many documents, and preserving exact output format. At the same time, table-context results remain far from perfect for many models, which implies that the benchmark is not purely a retrieval test. Fine-grained analysis at 128K full-text contexts shows little variation by subject, Sorting as especially hard, Aggregation as somewhat easier, and citation-related tasks as the hardest. By topic, worst performance occurs on Citation Relation and Reference Count, while author-related and title-related questions are comparatively easier.
6. Post-training behavior, failure modes, and limitations
The paper examines whether post-training can mitigate the observed weaknesses by training Qwen2.5-7B-Instruct-1M on SciTrek data using 7,703 training instances from the 19,543-instance training set, with contexts up to 128K (Li et al., 25 Sep 2025). The supervised fine-tuning setup uses 500 steps, batch size 32, learning rate , and warm-up rate 0.05. Reinforcement learning uses GRPO with a mixed reward based on EM + F1. The reasoning prompt for RL asks the model to "Think step by step, and place your final answer within \textbackslash{boxed{}." Training lasts one epoch, about 5 days.
The reported results show substantial gains on in-distribution settings but weak out-of-distribution generalization. Qwen2.5 zero-shot reaches 3.1 EM on Length ID and 0.2 on Length OOD; after SFT these become 16.3 and 2.3; after GRPO, 22.5 and 2.0. For Topic ID and Topic OOD, the scores are 3.9 and 1.5 in zero-shot, 20.9 and 10.0 after SFT, and 30.6 and 7.5 after GRPO. For Skills ID and Skills OOD, the scores are 1.4 and 5.8 in zero-shot, 10.8 and 19.2 after SFT, and 20.0 and 26.8 after GRPO. The authors’ interpretation is that SFT and GRPO improve local behavior, but do not solve long-context scaling, especially for unseen 512K and 1M contexts.
SciTrek’s SQL backbone enables detailed error analysis. At 128K full-text contexts, model accuracy correlates most consistently with question length. The reported Pearson correlations between EM and question length are approximately -0.15 for Qwen2.5-7B, -0.16 for Qwen2.5-14B, -0.14 for Gemma-3, -0.17 for Llama-4-Scout, -0.21 for GPT-4.1, and -0.22 for Gemini 2.5 Pro. For o4-mini, SQL length matters more strongly, with LenSQL at -0.25. This suggests that linguistic and logical complexity of the question contributes materially to failure.
Several systematic failure modes are highlighted. Weaker models often produce spurious NULL outputs rather than attempting the task; in a hard 128K sample, Qwen2.5-7B outputs NULL on 70% of aggregation questions, 90% of sorting, and 80–90% of several filtering types, while Qwen2.5-14B still outputs NULL on 50–60% of some filtering or relational tasks. Models also make format errors, such as providing an author list when a count is requested, returning unsorted lists, or outputting a list where an aggregate is required; GPT-4.1 is reported as especially prone to such errors in that analysis, including 50% incorrect format on Filtering+Aggregation in the hard sampled set. Partial answers are common when multiple outputs are required. Negation is a particularly severe weakness: among 132 negation instances in the filtering-related test subset, average EM is 0 for Qwen2.5-7B, 3.8 for Qwen2.5-14B, 0.8 for Gemma-3, 3.8 for Llama-4-Scout, 0.8 for Llama-3.3-70B, 1.5 for DeepSeek-R1-Distill-Llama-70B, 9.1 for Gemini 2.5 Pro, 3.8 for GPT-4.1, 22.0 for o4-mini, 9.1 for Qwen2.5-7B SFT, and 8.3 for Qwen2.5-7B GRPO. Inspection of GRPO chain-of-thought traces shows that reasoning structure can remain coherent while low-level counting steps remain incorrect, especially for references; a concrete example involves a 512K query with condition "more than 84 references" where the model incorrectly includes articles with 84 references.
SciTrek is positioned against prior long-context benchmarks such as NeedleBench, Ada-LEval, BABILong, HELMET, LIFBench, RULER, OpenScholar, LongBench v2, LongMemEval, L-Eval, HoloBench, MathHay, and Loong. The paper’s claim is that SciTrek is unusual in combining natural contexts, multiple information-processing skills, explicit reasoning structure, and scalability to 1M tokens. Relative to retrieval-heavy benchmarks, it requires integration and synthesis; relative to synthetic structured-reasoning benchmarks, it uses natural scientific documents; relative to expert-annotated scientific benchmarks such as OpenScholar, CURIE, and LongBench v2, it is easier to scale automatically.
The benchmark also has explicit limitations. It tests mostly metadata rather than deep scientific claims, method comparison, experimental interpretation, or causal scientific reasoning. It depends on Semantic Scholar metadata and preprocessing quality, and the authors note that manual corrections were sometimes needed to align metadata with full-text markdown articles. Human validation covers only 120 instances rather than the full dataset. Automatic SQL-to-question generation may still yield awkward or imperfectly natural questions despite round-trip validation. Coverage is broad but not exhaustive, since the benchmark is built from 662 articles in eight broad subject areas and citation neighborhoods around highly cited seed articles. Finally, the large gap between full-text and table performance indicates that SciTrek measures a mixture of context navigation, metadata extraction, symbolic reasoning, and output control rather than "reasoning" in isolation. A plausible implication is that the benchmark is best understood as a diagnostic test for scientific long-context pipelines rather than a direct measure of deep scientific understanding alone.