Theory-Informed Annotation Framework
- Theory-informed annotation frameworks are systems that ground labeling practices in explicit theoretical constructs rather than ad hoc methods.
- They integrate formal schemas with multi-stage human or human-AI workflows, ensuring reliable, traceable, and context-aware data annotations.
- The approach emphasizes clear definitions, iterative guideline revisions, and validation techniques that operationalize complex theoretical constructs into practical annotation practices.
Searching arXiv for the cited papers to ground the article in current literature. arXiv search: "(Li et al., 31 Mar 2025)" A theory-informed annotation framework is an annotation design in which labels, units, workflows, interfaces, and validation procedures are anchored in explicit theories rather than ad hoc conventions. In the recent literature, this orientation appears in human-AI collaborative genome annotation, iterative guideline refinement for LLM annotation, typologically informed construction annotation, DSM-5-TR–aligned depression symptom annotation, graph-based linguistic annotation, and socially situated critiques of labeling practice (Li et al., 31 Mar 2025, Kim et al., 20 May 2026, Weissweiler et al., 2024, Cao et al., 16 Jul 2026) 0010033. Taken together, these works suggest that annotation is not merely a technical act of attaching labels to data; it is also a procedure for operationalizing theoretical constructs, preserving provenance, and constraining interpretation.
1. Conceptual scope
The literature defines annotation broadly. In genome annotation, it is “the process of identifying and interpreting the functional elements encoded within a genome,” and the central problem is how to combine automated prediction with manual curation at scale (Li et al., 31 Mar 2025). In economic intelligence, annotation has a dual nature: it is both an object attached to a document and an action through which an annotator enriches information, records intellectual activity, and creates added value for decision-making (0812.1394). In formal linguistic work, annotation covers “any descriptive or analytic notations applied to raw language data,” whether the primary data are audio, video, physiological recordings, or text, and the emphasis shifts from file format to logical structure [0010033].
This breadth matters because theory-informed frameworks do not assume a single annotation substrate. They appear in sentence-level narrative labeling, token- and span-level biomedical named entity recognition, graph-based semantic annotation, document-centered economic intelligence systems, multimodal video coding, structured extraction from historical directories, and case-level mental-health evidence synthesis (Li et al., 2017, Kim et al., 20 May 2026, Abend et al., 2020, Paekivi et al., 14 Jun 2026, Ren, 25 May 2026, Cao et al., 16 Jul 2026). The commonality is not the medium, but the insistence that the annotation scheme be tied to an explicit account of what the labels mean, how they should be applied, and why disagreement occurs.
A recurrent distinction is between annotation as representation and annotation as procedure. Formal models such as annotation graphs, UCCA graphs, UCxn labels in CoNLL-U MISC, and attribute–value objects specify how annotations are stored and linked to source material 0010033. Procedural frameworks specify how humans, models, or both produce those annotations, often through pilot rounds, moderation, adjudication, conflict resolution, or multi-stage review (Boudoua et al., 19 Jan 2026, Kim et al., 20 May 2026, Ren, 25 May 2026).
2. Theoretical foundations
Theory-informed annotation is plural in its intellectual sources. In human-AI collaborative genome annotation, the explicit foundations include human-AI collaboration and teaming, mixed-initiative systems, human-in-the-loop machine learning, active learning, uncertainty-aware decision support, cognitive workload and trust calibration, socio-technical systems, and knowledge representation through ontologies and knowledge graphs (Li et al., 31 Mar 2025). In typologically informed construction annotation, the relevant theories are Construction Grammar and linguistic typology, with constructions defined primarily by function and then realized through language-particular morphosyntactic strategies (Weissweiler et al., 2024). In UCCA, the Foundational Layer is grounded in cognitive-semantic and typological principles, especially the distinction between Scenes and non-Scene units, and categories such as Process, State, Participant, Elaborator, and Ground (Abend et al., 2020).
Other frameworks derive their categories from domain theory rather than from general annotation methodology. The misogyny annotation scheme is grounded in ambivalent sexism theory, gender essentialism, toxic masculinity, gendered racism, post-feminism, backlash or “sexism shift,” and internalized misogyny (Deligianni et al., 24 Jan 2026). The depression-symptom framework operationalizes DSM-5-TR Criterion A for Major Depressive Disorder through criteria A1–A9 and a case-level rule requiring at least five symptoms during the same two-week period, with at least one being A1 or A2 (Cao et al., 16 Jul 2026). The narrative-structure framework consolidates Freytag’s dramatic arc with Labov and Waletzky’s functional categories, alongside Prince and Todorov, to define labels such as Orientation, Complicating Action, Most Reportable Event, Resolution, Aftermath, Evaluation, and Direct Comment to Audience (Li et al., 2017).
A separate line of work emphasizes social theory and critique. “Discipline and Label” argues that annotation is a form of automated social categorization and that labels often encode WEIRD values, state categories, and institutional assumptions that are then imposed through instructions onto non-WEIRD annotators (Smart et al., 2024). “On Anomalies in Annotation Systems” frames the persistence of pen-and-paper assumptions in digital systems as a paradigm problem, and proposes the artefact–context model to capture time-binding, gestural activity, and cross-document workspaces (0706.1087). These interventions broaden the meaning of “theory-informed”: theory can justify a codebook, but it can also expose the power relations and validity limits of the categories being coded.
3. Representational and architectural elements
A theory-informed framework typically specifies not only labels but also a representational architecture. In HAICoGA, the seven key elements are humans, AI systems/tools, data, goals/tasks, human-machine interface, environment, and collaboration (Li et al., 31 Mar 2025). In the economic-intelligence model, annotation objects are built through a semantic function linking document content to explicit or implicit attribute–value structures, with inference used to operationalize missing attributes or values (0812.1394). In the annotation graph formalism, the common core is a labeled, directed acyclic graph whose nodes may be anchored to timelines or textual positions, thereby providing a substrate for heterogeneous linguistic layers [0010033]. UCCA likewise uses a labeled DAG with primary and remote edges, while UCxn adds construction labels to UD treebanks by storing Cxn and optional CxnElt metadata in CoNLL-U MISC without altering the underlying dependency graph (Abend et al., 2020, Weissweiler et al., 2024).
Guidelines and codebooks are the most common operational bridge between theory and representation. The generic corpus-creation methodology insists on explicit research objectives, clear category definitions, inclusion and exclusion criteria, boundary rules, negative examples, and an “I don’t know” option (Boudoua et al., 19 Jan 2026). The LLM moderation framework reuses human-written guidelines, entity schema definitions, and prompt templates as alignment mechanisms; it stores outputs in PubAnnotation JSON and refines the guideline document iteratively rather than changing the label ontology wholesale (Kim et al., 20 May 2026). In depression annotation, the exported object includes evidence spans, criterion-level analyses, case-level synthesis, reasoning traces, and edit histories, making auditability a first-class part of the schema (Cao et al., 16 Jul 2026).
The same principle appears in multimodal and document-extraction settings. The TikTok study uses a theory-ensemble codebook with 77 coded content variables across 11 theory blocks, plus 2 controls and 20 audit fields, all represented as forced-choice categorical or binary variables with “Not Applicable” for dependent fields (Paekivi et al., 14 Jun 2026). Double Triangle Annotation treats each extracted field as an independent atomic unit, compares exact field-level outputs, and records adjudicated corrections through a custom platform that shows the source image, both model outputs, and character-level differences (Ren, 25 May 2026). This suggests that in theory-informed annotation, data structures are not neutral containers: they encode what counts as evidence, what constitutes a unit, and how a theoretical construct becomes machine-readable.
4. Workflow patterns and human-machine organization
The workflows proposed in this literature are iterative and strongly structured. A generic corpus-creation process begins with scoping, literature review, corpus selection, pilot sampling, preliminary guideline drafting, blind annotation by at least two experts, inter-annotator agreement assessment, disagreement analysis, revision, and finalization when agreement is satisfactory (Boudoua et al., 19 Jan 2026). The narrative-structure project used three tutorial rounds over five weeks, with 25 stories per session, before computing reliability on a shared set of 71 stories (Li et al., 2017). These are classical human-centered procedures in which theory is stabilized through repeated application and adjudication.
Recent machine-assisted frameworks preserve this logic but redistribute labor. In “Refining and Reusing Annotation Guidelines for LLM Annotation,” the moderation loop has four phases: LLM-based annotation, evaluation, discrepancy analysis, and moderation. Moderation itself proceeds through Pattern Explanation, Principle Generation, and Guideline Refinement, after which the updated guideline is re-used for a new iteration (Kim et al., 20 May 2026). The framework targets the dominant discrepancy group, produces exactly one generalized IF/THEN rule with an EXCEPT clause per iteration, and discards the last refinement if the cycle fails to improve the strict F1 used as an inter-annotator-agreement proxy (Kim et al., 20 May 2026).
Human-AI collaboration frameworks make the distribution of initiative explicit. HAICoGA organizes genome annotation into data ingestion, automated prediction, uncertainty estimation and triage, human curation, revision and validation, provenance and documentation, feedback loops, and continuous improvement (Li et al., 31 Mar 2025). Its multi-agent vision includes a user, a manager agent, an automated genome-annotation agent, specialized manual-curation agents, and a critique agent, with bi-directional GUI/CUI interactions and human approval authority (Li et al., 31 Mar 2025). The depression framework uses a three-stage pipeline—candidate evidence selection, criterion-level DSM-5-TR analysis, and case-level synthesis—then updates Example Memory and Reflection Memory through expert feedback without retraining (Cao et al., 16 Jul 2026).
A stricter consensus architecture appears in “Double Triangle Annotation.” Layer 1 pairs two architecturally independent multimodal LLMs and escalates disagreements to a human jury; Layer 2 compares the outputs of two such Layer-1 systems and escalates residual conflicts to a final reviewer (Ren, 25 May 2026). MAFA applies a related specialization principle to FAQ annotation through a Query Planning Agent, four specialized ranker agents using Attentive Reasoning Queries, and a Judge Agent Reranker, with structured JSON outputs and fallback average scoring when the judge fails (Hegazy et al., 19 May 2025). Across these systems, a stable design principle emerges: theory-informed annotation increasingly relies on orchestrated disagreement, rather than on a single annotator or single model.
5. Reliability, evaluation, and explainability
The literature treats reliability as necessary but task-dependent. For corpus construction, recommended agreement metrics include Cohen’s kappa for two annotators,
with the guide citing as minimum acceptable, $0.6$ as acceptable, and $0.9$ as perfect (Boudoua et al., 19 Jan 2026). In theory-grounded misogyny annotation, the resulting dataset achieved after excluding the first training week, which the paper interprets as substantial agreement (Deligianni et al., 24 Jan 2026). By contrast, high-level narrative annotation produced fair pairwise kappas of $0.39$, $0.41$, and $0.42$, improving only modestly after category merging (Li et al., 2017). The TikTok study reports a reliability gradient: directly observable audiovisual variables can be coded fairly reliably, whereas deeper semiotic and archetypal constructs are difficult for both humans and machines (Paekivi et al., 14 Jun 2026).
Task-specific metrics vary with the annotation object. The LLM guideline-moderation paper uses strict span-and-type precision, recall, and F1,
and treats strict F1 as an inter-annotator-agreement proxy during moderation (Kim et al., 20 May 2026). Double Triangle Annotation evaluates structured extraction with Word Error Rate and Character Error Rate,
achieving a final WER of 0 and CER of 1 in controlled evaluation (Ren, 25 May 2026). The depression framework evaluates sentence-level evidence selection, criterion-level classification, evidence-pair linking, case-level diagnosis accuracy, and expert-effort measures such as total edits and time saved (Cao et al., 16 Jul 2026).
Explainability and traceability are treated as part of quality, not as optional extras. HAICoGA emphasizes citation grounding, verification reports, action logs, and documented workflows through systems such as VarChat, GeneAgent, BKGAgent, and Virtual Lab (Li et al., 31 Mar 2025). The depression framework exports clinical evidence, reasoning traces, and edit histories (Cao et al., 16 Jul 2026). The moderation framework preserves guideline versions 2 and identifies discrepancy patterns by category and label pair (Kim et al., 20 May 2026). Formal frameworks such as annotation graphs and UCCA make provenance and reanalysis possible by separating the logical structure of annotation from any particular surface serialization 0010033. A plausible implication is that theory-informed annotation evaluates not only whether a label is correct, but also whether its justification remains inspectable.
6. Applications, controversies, and open directions
Theory-informed annotation now spans a wide range of domains. The surveyed cases include genome annotation, biomedical named entity recognition, construction annotation atop Universal Dependencies, economic intelligence, FAQ mapping, historical-document extraction, single-cell annotation, narrative macro-structure, misogynistic language detection, depression symptom annotation, and multimodal short-video analysis (Li et al., 31 Mar 2025, Kim et al., 20 May 2026, Weissweiler et al., 2024, 0812.1394, Hegazy et al., 19 May 2025, Ren, 25 May 2026, Huang et al., 2 Dec 2025, Li et al., 2017, Deligianni et al., 24 Jan 2026, Cao et al., 16 Jul 2026, Paekivi et al., 14 Jun 2026). In each case, the framework is justified by a mismatch between raw predictive convenience and the need for interpretable, domain-aligned, or governance-ready labels.
The controversies are equally consistent across domains. One concerns external validity and the social construction of labels: “Discipline and Label” argues that disagreement and cultural variation should not be treated merely as noise, because platform annotations often generalize WEIRD category schemes beyond the settings in which they were produced (Smart et al., 2024). Another concerns system design: “On Anomalies in Annotation Systems” argues that digital annotation systems still inherit the limitations of paper-and-pencil paradigms, especially in cross-format interoperability, gestural capture, and time-binding of collaborative work (0706.1087). In sensitive domains, theory can also reveal what simpler codebooks omit: the misogyny framework shows that hostile-only taxonomies miss benevolent sexism, post-feminist denial, backlash, toxic masculinity, and intersectional harms (Deligianni et al., 24 Jan 2026).
Open research questions are increasingly about scale without epistemic collapse. HAICoGA asks how to build scalable, adaptive multi-agent systems that maintain alignment while enabling autonomy and specialization, how to mitigate hallucinations and support continual learning, and how to evaluate human-AI teams across quality, explainability, trust, and process dynamics (Li et al., 31 Mar 2025). The guideline-moderation paper identifies intrinsic guideline-quality assessment, regression tracking, larger refinement sets, and extension beyond biomedical NER as unresolved issues (Kim et al., 20 May 2026). Double Triangle Annotation highlights independence violations, source-induced correlated errors, and selection bias introduced by filtering on agreement (Ren, 25 May 2026). The depression framework explicitly leaves evaluation across multiple feedback cycles to future work (Cao et al., 16 Jul 2026).
Taken together, these works suggest that a theory-informed annotation framework is best understood as an annotation regime with five persistent properties: explicit theoretical grounding, operational codebooks or formal schemas, structured human or human-AI workflows, multi-dimensional validation, and strong provenance. What varies is the theory being operationalized and the medium being annotated. What remains stable is the claim that annotation quality depends on making the underlying concepts, assumptions, and revision mechanisms visible.