OpenConstruction: AEC Open Data Platform
- OpenConstruction is an open-source platform that aggregates, organizes, and contextualizes digital resources in the AEC domain.
- It employs standardized metadata and federated indexing to enhance discoverability, interoperability, and adherence to FAIR principles.
- The platform supports manual curation and automated API access, enabling reproducible research and data-centric AI in construction.
OpenConstruction is an open-source, community-driven infrastructure for aggregating, organizing, and contextualizing openly accessible digital resources in the Architecture, Engineering, and Construction (AEC) domain. In the literature, the name denotes both a systematic synthesis of open visual datasets for data-centric AI in construction monitoring and a broader ecosystem that interlinks datasets, computational models, validated use cases, and educational resources. Across these formulations, the stated objective is to mitigate fragmentation and inconsistent documentation, improve discoverability and reuse, and operationalize the FAIR principles—Findable, Accessible, Interoperable, and Reusable—through standardized metadata, federated indexing, curator-led validation, and interoperable access mechanisms (Xiong et al., 15 Aug 2025, Xiong et al., 2 Jan 2026).
1. Origins, problem setting, and scope
OpenConstruction emerged in response to a specific infrastructural deficit in AEC research: digital assets such as datasets, models, use cases, and educational materials were dispersed across repositories and described inconsistently, which limited interpretability, comparability, and cumulative reuse. One formulation of the project, focused on construction monitoring, conducted an extensive search of academic databases and open-data platforms and identified 51 publicly available visual datasets spanning the 2005–2024 period. A later formulation generalized the effort into an AEC-wide open-science ecosystem structured around multiple resource types rather than datasets alone (Xiong et al., 15 Aug 2025, Xiong et al., 2 Jan 2026).
The scope therefore extends beyond repository aggregation. In the dataset-centric formulation, OpenConstruction functions as a systematic synthesis of visual resources for AI and ML in construction monitoring, emphasizing data fundamentals, modalities, annotation frameworks, and downstream application domains. In the broader ecosystem formulation, it serves the built environment lifecycle more generally by organizing datasets, models, use cases, and open educational resources into coordinated catalogs. This progression suggests a shift from a review-driven catalog toward a sustainable cyberinfrastructure for digital knowledge sharing in AEC, although the underlying continuity remains the use of structured descriptors, open access, and FAIR-oriented design (Xiong et al., 15 Aug 2025, Xiong et al., 2 Jan 2026).
2. Catalog design and metadata model
The broader OpenConstruction platform is organized into four interlinked catalogs: a Dataset Catalog, a Model Catalog, a Use Case Catalog, and an Educational Resource catalog. The Dataset Catalog aggregates metadata for multimodal AEC datasets, including resources such as images and point clouds, from open-access repositories and archives such as Zenodo, Roboflow, and Google Dataset. The Model Catalog documents computational models and code for domain-specific AEC tasks, drawing from GitHub, Hugging Face, and academic databases. The Use Case Catalog describes validated real-world AEC applications of digital and AI/ML techniques and explicitly connects them to datasets and models. The Educational Resource catalog curates open textbooks, tutorials, and training materials, including materials from platforms such as Pressbooks (Xiong et al., 2 Jan 2026).
The metadata model combines general descriptors and domain-specific descriptors. General fields include identifier, title, contributors, license, and access link. Domain-specific fields include project phase, tasks, applications, stakeholders, modality, and technologies. In the dataset-synthesis formulation, this schema is presented through four pillars: data fundamentals, data modalities, annotation framework, and application domains. The data-fundamentals pillar covers properties such as storage size, quantity, resolution, format, collection method, year, geographic scope, and licensing or access barriers. The modality pillar covers RGB imagery, thermal images, point clouds, synthetic images, video, and temporal or spatial attributes. The annotation pillar covers bounding boxes, segmentation, keypoints, captions, and 3D point-level labels. The application-domain pillar covers vision tasks and industry uses such as safety monitoring, progress tracking, machinery monitoring, site understanding, and navigation or mapping (Xiong et al., 15 Aug 2025).
This metadata emphasis is central to OpenConstruction’s claim to interoperability. Rather than treating resources as isolated uploads, the platform formalizes descriptors that support both human browsing and machine interpretation, with controlled vocabularies, cross-linking between assets, and consistent description of access conditions (Xiong et al., 2 Jan 2026).
3. Infrastructure, indexing, and access mechanisms
OpenConstruction is architected as a federated metadata system rather than a storage host for the underlying resources. The infrastructure layer stores federated metadata, assigns persistent identifiers, and supports schema-based records together with semantic and provenance validation. The resources themselves remain at their original hosts. This design is paired with a service layer that organizes content into the four catalogs, an access layer that exposes both a web portal and a machine-actionable MCP API, an application layer oriented toward research, education, and practice, and a governance layer for validation, contributor recognition, community engagement, and sustainability (Xiong et al., 2 Jan 2026).
The access model combines interactive and programmatic use. The web portal supports browsing, search, and exploration by human users. The MCP API, described as “Model Context Protocol,” is intended for automated access, resource integration, interoperability, and incorporation into research pipelines. The dataset-focused formulation likewise emphasizes machine-readable JSON metadata for automated search, benchmarking, and systematic meta-analysis. Taken together, these features define OpenConstruction not merely as a directory but as an interoperable interface layer for reproducibility and resource reuse (Xiong et al., 2 Jan 2026, Xiong et al., 15 Aug 2025).
A common misconception is that OpenConstruction operates as a monolithic repository. The published descriptions instead specify federated indexing: only metadata is stored centrally, while provenance and access remain linked to original repositories. Another misconception is that openness implies a minimally curated submission model. The platform description explicitly rejects that interpretation through manual validation and quality checks prior to publication (Xiong et al., 2 Jan 2026).
4. Curation, validation, and governance
The validation workflow comprises five stated stages. First, metadata is automatically extracted from original repositories using scripts and documentation. Second, manual alignment ensures consistency with the platform schema and controlled vocabularies. Third, human curation performs quality and ethical checks for completeness, license compliance, accessibility, correct terminology, and privacy, ethical, and legal concerns. Fourth, deduplication is carried out using matching of titles, authors, and repositories. Fifth, approved resources are curated, indexed, and made public (Xiong et al., 2 Jan 2026).
Governance is described as transparent and community-centric. Submissions undergo manual validation for completeness, accessibility, ethics, and license. Contributors receive public recognition through visible profiles and institutional affiliations. Community engagement is pursued through workshops, forums, and regular updates, while long-term sustainability is envisioned through endorsement by professional societies and the development of contributor recognition metrics. These features position governance as part of the infrastructure rather than an external administrative layer (Xiong et al., 2 Jan 2026).
The dataset-synthesis paper reinforces this orientation by identifying systemic problems that justify such curation: scattered documentation standards, inconsistent taxonomies, barriers to access, and missing FAIR-aware metadata schemas. OpenConstruction’s validation and governance mechanisms are therefore designed as responses to structural fragmentation rather than merely editorial preferences (Xiong et al., 15 Aug 2025).
5. Empirical coverage and landscape characterization
As of December 2025, the broader platform hosts 204 total entries: 94 datasets, 65 models, 28 use cases, and 17 educational resources. Within the Dataset Catalog, the reported modality composition is ground-level RGB (64), aerial RGB (10), point clouds (9), and synthetic data (8), with some thermal and video datasets. Within the Model Catalog, reported categories include object detection (20), segmentation (12), tracking (2), pose estimation (2), SLAM (3), image captioning (2), and 3D reconstruction (4). The Use Case Catalog is distributed across construction (18), preconstruction (2), operations/maintenance (2), and design (6). The Educational Resource catalog contains 16 open textbooks and 1 set of slides on AEC, computing, and data-intensive topics (Xiong et al., 2 Jan 2026).
| Catalog | Entries | Brief characterization |
|---|---|---|
| Datasets | 94 | Ground-level RGB, aerial RGB, point clouds, synthetic data, with some thermal and video |
| Models | 65 | Object detection, segmentation, tracking, pose estimation, SLAM, image captioning, 3D reconstruction |
| Use cases | 28 | Construction, preconstruction, operations/maintenance, design |
| Educational resources | 17 | 16 open textbooks and 1 slide set |
The earlier systematic synthesis provides a diagnostic view of the construction-dataset landscape that motivated this broader platformization. It reports that RGB imagery is dominant, with ground-level RGB accounting for more than 80% of datasets, while thermal images represent less than 3%, LiDAR/point clouds about 6%, synthetic images about 8%, and videos or temporal data less than 3%. Bounding boxes are the dominant annotation form at about 78%. Safety monitoring is the majority application domain, and only about 20% of datasets include activity or action labels. The review also notes that approximately one-third of datasets require requests or are otherwise restricted, and that inconsistent taxonomies complicate cross-dataset transfer and benchmarking (Xiong et al., 15 Aug 2025).
These findings are significant because they show that OpenConstruction is not only an index of what exists but also an empirical map of what is missing. The platform’s emphasis on multimodality, consistent descriptors, and cross-resource associations directly reflects the review’s identification of underrepresented modalities, sparse temporal data, and interoperability deficits (Xiong et al., 15 Aug 2025).
6. Benchmarking, education, and future data infrastructure
Two case studies illustrate the operational role of OpenConstruction. In the first, a researcher identifies and benchmarks AEC models for pose estimation by using the web portal to filter models by task, modality, and application, examining standardized descriptors such as task, input/output, license, and data links, and then retrieving model information through the MCP API for automated aggregation into benchmarking pipelines. The stated outcome is more consistent cross-study comparison of independently developed models and support for standardized benchmarking practices. In the second, an educator uses the catalogs to integrate curated datasets, models, and metadata into teaching so that students can compare sensing modalities, annotation formats, and AEC tasks while learning reproducibility, data literacy, and computational methods (Xiong et al., 2 Jan 2026).
The dataset-focused study frames these functions more broadly as part of a FAIR-anchored roadmap for construction data infrastructure. Its recommendations are grouped into data acquisition and quality assurance, data integration and semantic management, governance and community engagement, and performance monitoring and continuous improvement. Specific priorities include broadening sensors and data types, standardizing curation and validation, developing unified metadata and communication protocols, building construction-domain ontologies, linking captured data to project-management context such as BIM, adopting privacy and anonymization practices for worker-centric imagery, and implementing regular benchmarking of coverage, accuracy, and completeness (Xiong et al., 15 Aug 2025).
Within this framing, OpenConstruction functions as both a current platform and a template for future AEC knowledge infrastructures. Its declared contribution is to transform fragmented digital assets into FAIR, reusable, and openly accessible knowledge objects. A plausible implication is that its long-term importance will depend not only on the number of indexed resources but also on whether standardized descriptors, machine-actionable access, and transparent governance become durable conventions across the AEC research community (Xiong et al., 2 Jan 2026, Xiong et al., 15 Aug 2025).