- The paper develops a provision-level model that records linked instruments, relevant provisions, legal qualifications, conditions, and documentary evidence rather than treating cross-references as simple hyperlinks.
- The paper applies bidirectional inversion to a bilingual corpus of 14 EU and Italian instruments, uncovering six incorrect article references, three qualification errors, and one divergent duplicate description.
- The paper shows that textual citation frequency can misrepresent legal importance, while substantive interactions may exist without shared definitions or explicit citations, supporting more auditable legal knowledge bases.
Overview and contribution
This paper develops a provision-level model and a construction protocol for qualified legal cross-references, instantiated in a bilingual Italian/English corpus of fourteen EU and Italian instruments surrounding Regulation (EU) 2024/1689 (the AI Act). The central claim is that a curated cross-reference is not a pointer but a compact legal proposition—identifying connected provisions, characterizing their relationship, and stating conditions—that can be false in ways that a correctly extracted citation cannot. The paper's four contributions are: an interaction-level record model; a separation between applicative interaction and definitional overlap; bidirectional inversion as a construction-time verification protocol; and an empirical correction record generated by applying the protocol (2608.19194).
The scope is deliberately bounded. The paper does not claim to represent the full EU digital acquis, to automate legal interpretation, or to estimate error rates in other corpora. It shows only that qualified legal relationships can be represented as inspectable records whose construction itself functions as quality control.
Positioning relative to prior work
The contribution is complementary to three established strands. Network analysis of EU legal sources establishes that legislation forms structured systems of citations, amendments, and legal bases; semantic edge labeling classifies existing citations by purpose rather than leaving them untyped. Neither strand represents two relationship classes recorded here: substantive interaction, where provisions operate on connected obligations without either passage naming the other act, and mediated intersection, where the connection is carried by a third instrument invoked independently by both sides. In neither case does an arc exist to label.
The paper also distinguishes extraction from curation on epistemic grounds: extraction answers questions about a text, whereas a curated qualified entry asserts a legal consequence with conditions—an assertion that can be false even when every cited document exists and resolves. The EUR-Lex ELI infrastructure records a related limit from below: its interface relates whole acts, and finer resolution would require ELIs at article and clause level. The provision-level model adopted here addresses that granularity directly.
Record model
Each qualified interaction is represented as a tuple R=⟨A,B,PA​,PB​,Q,E⟩, where A and B are instruments, PA​ and PB​ are relevant provision sets, Q is the legal qualification, and E identifies documentary evidence. The evidence element includes an evidence locator, rendering language, corpus revision, and a per-language content fingerprint, so that moved or edited renderings can be detected against the record. Directional explanation is permitted: the account read from act A need not be a textual transposition of the account read from B.
Five core qualification categories organize the corpus:
| Category |
Operational criterion |
| Direct textual reference |
At least one act names the other in the normative text |
| Bounded presumption of conformity |
Stated condition supports presumption of compliance within express bounds |
| Substantive interaction |
Provisions operate on connected facts/obligations without naming each other |
| Mediated intersection |
Connection made legible by a third act, recorded explicitly as mediator |
| Institutional analogy |
Parallel structures with no asserted legal reference or consequence |
The categories guard against symmetric errors: calling every material interaction a textual reference, and discarding legally relevant interactions for lack of a citation. Conditions travel with the relationship—a presumption limited to certificate-covered requirements is not recorded as unconditional equivalence. Separately, applicative interaction (Aij​) and definitional overlap (Dij​) are treated as independent axes; a negative definitional assessment is recorded as "examined, not established," not conflated with an unexamined case.
Bidirectional inversion protocol
The protocol's rationale is that copying a fact-sheet entry from act A's page into B's page reproduces the assertion without testing it. Inversion instead reconstructs the interaction from B's perspective against both acts' provisions: column order changes, deictic expressions ("this Regulation") change referent, and B's engaged provision must be identified independently. The seven-step sequence proceeds from selecting the source entry, through opening both acts and reconstructing from B, to testing qualification and bounds, reconciling the two descriptions, and deciding whether B receives autonomous analysis or a reciprocal pointer.
Outcomes are confirmed, corrected, supplemented, or pointer-only. Crucially, verification and publication are separated: reciprocity of discoverability is required (every interaction reachable from both fact sheets), but duplication is avoided because two full tables describing one legal fact can drift apart—the divergence documented later in this corpus illustrates exactly that failure mode. The structural remedy is a non-directional registry in which the interaction is the unit of record, rendered by fact sheets anchored via language-independent provision addresses supporting mechanical containment checks against each act's actual structure.
Corpus and verification results
The corpus comprises fourteen instruments (AI Act, GDPR, Data Act, DGA, DSA, CRA, NIS2, Machinery Regulation, Product Liability Directive, Market Surveillance Regulation, Cybersecurity Act, Regulation (EU) 2025/37, the Digital Omnibus on AI, Italian Law 132/2025), yielding 84 cross-reference sections per language, 315 table rows per language, 60 documented unordered pairs, of which 28 are documented from both sides and 32 from one side. Eighty-six distinct qualifications occur in the Italian vocabulary, presented as editorial refinements rather than an exhaustive ontology.
Twenty relationships were inverted and twenty-two candidate definitional columns assessed. The protocol surfaced ten defects:
- Six incorrect article references, including Art. 15 NIS2 (corrected to Art. 18, the biennial cybersecurity report rather than CSIRTs network), Art. 66 AI Act (corrected to Art. 67(5)), Art. 24 Machinery Regulation (corrected to Art. 20(9)), Art. 3(3) Regulation 2019/1020 (corrected to Art. 11(5)), Art. 30 Data Act (corrected to Art. 35(1)(d)), and a nonexistent definition at Art. 2(24) Cybersecurity Act ("vulnerability" appears at Art. 51 but is undefined there).
- Three inaccurate qualifications: a direct reference mislabeled as mediated; a certification obligation stated without its delegated-act, scheme-availability, and assurance-level preconditions; and Article 51 Cybersecurity Act objectives reduced to a different security formulation.
- One inter-sheet divergence: two descriptions of the CRA–Cybersecurity Act relationship naming different provisions and qualifications.
Because one defect arose outside inversion, no rate such as "six out of twenty" is reported. The consequential result is limited but firm: document existence and link resolution are insufficient tests for curated legal cross-references, since the errors concerned the legal proposition carried by the link, not its technical target.
Substantive findings from the map
Three presumptions, one divergent precondition. Regulation (EU) 2019/881 supplies the certification framework for three product instruments, each attaching a bounded conformity presumption: AI Act Art. 42(2), Machinery Regulation Art. 20(9), and CRA Art. 27(8). They share trigger, effect, and bound, but differ on one precondition: the AI Act and Machinery Regulation expressly require publication of the scheme's references in the Official Journal; Article 27(8) CRA does not state that condition. A representation recording all three uniformly would lose the distinction. Beyond the presumption, the CRA depends on the framework in four qualitatively distinct ways total: definitional borrowing (Art. 3 points 3 and 46), the bounded presumption, a delegated specification power (Art. 27(9)), and a conditional conformity route for critical product categories requiring assurance level at least "substantial" (Arts. 8(1), 32(4)).
Textual frequency versus legal load (RQ1). The literal string "2019/881" occurs once in enacting terms in both the AI Act and the Machinery Regulation—each deriving a bounded presumption from that single provision—and nine to sixteen times across NIS2 and the CRA. Textual footprint differs by an order of magnitude while the legal effect at issue is comparable in kind. Citation-count-based ranking therefore fails to capture legal dependency weight in this case.
Mediated intersection. The Cybersecurity Act and Market Surveillance Regulation name neither each other nor share textual citations, yet connect through Regulation (EC) No 765/2008: Art. 60(1) of the former relies on it for accreditation of conformity assessment bodies; Art. 11(5) of the latter directs authorities to take due account of reports from accredited bodies. Naming the mediator makes both connection and limits inspectable; the entry does not assert automatic conformity equivalence.
Definitional sparsity (RQ3). Nineteen of twenty-two candidate columns admitted under a one-concept threshold produced sparse tables by design (e.g., six populated cells out of twenty-four in the Cybersecurity Act comparison). Three pairs showed applicative interaction without shared defined concepts, including DGA versus the Market Surveillance Regulation. The claim is narrower than asserting shared vocabulary absence—it is testable under the stated protocol—but it establishes that operational proximity and terminological overlap diverge.
Institution-specific routing. An initial entry recorded Art. 41 Cybersecurity Act (concerning ENISA, a Union body invoking Regulation 2018/1725) as a GDPR reference, corrected during construction—illustrating why definition provenance must be preserved rather than inferred from labels.
Structural asymmetry as provenance. Of thirty-two one-sided pairs, fifteen are accounted for by recorded design choices (seven reflecting the AI Act sheet's self-text-only rule, four directional relations, four other decisions); seventeen reflect publication order. Asymmetry is explicitly interpreted as corpus construction history, not a property of EU law, preventing graph measures from being mistaken for claims about regulatory importance.
Limitations
The paper concedes several boundaries plainly. The graph represents one selected corpus, and missing edges are not evidence of legal absence. Node degree and direction depend on fact-sheet ordering and inclusion rules. The same curator authored and inverted entries; changing modes exposed defects, but errors invisible in both modes may remain, and no inter-annotator agreement was measured. No population-level error estimate is offered, and there is no independent ground truth against which aggregate accuracy could be computed. Finally, while the five core categories structure the corpus, finer qualifications remain open-ended, which preserves legal detail but does not yield a closed ontology suitable for automatic inference. Notably, the corpus repository itself is not public; reproducibility rests on inspection of the published site plus an archived source package released under CC BY-SA 4.0, not on a clonable repository.
Reproducibility
All numerical results derive from eight read-only instruments sharing one parser: section/table readers, a graph census with JSON export, a qualification-coverage census, a bilingual structural validator, an article-structure index, a row-fingerprint utility, an anchor-proposal tool with collision checks, and a column-insertion tool enforcing the definitional threshold. Measurements refer to corpus commit dated 19 August 2026. A candid methodological note records that every parser convention was added after a parse failure revealed it—"a parser that appears complete has usually not yet met the corpus"—making tooling history part of the reproducibility record.
Conclusion
The paper treats legal cross-references as verifiable assertions with identifiable provisions, qualifications, conditions, and evidence, and demonstrates that reconstructing each assertion from the opposite instrument's perspective functions as a productive verification operation: in this corpus it surfaced six incorrect references, three qualification errors, and one divergent duplicate description. The substantive examples—divergent preconditions among three formally similar presumptions, an order-of-magnitude gap between textual frequency and legal load, and operational interaction without definitional overlap—show why the additional representational structure carries information that links alone do not. The combined artifact-and-method contribution offers a basis for curated legal knowledge bases whose assertions can be traced, tested, and maintained, with validity bounded to the corpus studied.