AI City Challenge: Urban AI Benchmark
- AI City Challenge is an annual benchmark evaluating intelligent video analysis and multimodal AI across urban environments, including transportation, retail, and public safety.
- It integrates large-scale real and synthetic datasets with diverse tasks such as vehicle tracking, anomaly detection, and indoor occupancy analytics.
- The challenge employs rigorous evaluation methods with standardized metrics, anti-overfitting protocols, and code-release rules to ensure reproducibility and deployment readiness.
Searching arXiv for the specified AI City Challenge papers to ground the article in published sources. {"query":"AI City Challenge 2025 9th AI City Challenge arXiv (Tang et al., 19 Aug 2025)", "max_results": 5} The AI City Challenge is an annual benchmark for intelligent video analysis and multimodal AI in smart-city settings. It was created to accelerate intelligent video analysis that helps make cities smarter and safer, with a persistent emphasis on transportation and, in later editions, retail business automation, warehouse reasoning, industrial automation, and public safety. Across editions, the challenge has combined large-scale real traffic video, synthetic data, natural-language annotations, and standardized evaluation protocols to assess both research progress and readiness for deployment-oriented use cases such as vehicle counting, multi-camera tracking, anomaly detection, traffic-safety description, driver-behavior analysis, retail checkout, helmet-rule violation detection, and edge-efficient fisheye perception (Naphade et al., 2020, Naphade et al., 2022, Tang et al., 19 Aug 2025).
1. Historical development and scope
The AI City Challenge began in 2017 as a smart-city vision benchmark, and by the 4th edition it had become a CVPR workshop competition organized around richly annotated urban video datasets and standardized evaluation systems. The early editions centered on intelligent transportation systems (ITS), including video-based automatic vehicle counting, city-scale vehicle re-identification, city-scale multi-target multi-camera vehicle tracking, and traffic anomaly detection. The 5th edition added natural-language-based vehicle retrieval; the 6th edition expanded into retail checkout and naturalistic driving analysis; the 7th edition introduced multi-target multi-camera people tracking in indoor environments and motorcycle helmet-rule violation detection; the 8th edition emphasized retail, warehouse settings, and ITS; and the 9th edition focused on transportation, industrial automation, and public safety through four tracks spanning 3D tracking, traffic-safety VQA, warehouse spatial reasoning, and fisheye road-object detection (Naphade et al., 2020, Naphade et al., 2021, Naphade et al., 2022, Naphade et al., 2023, Wang et al., 2024, Tang et al., 19 Aug 2025).
Participation statistics show both the scale of the benchmark and the fact that reports use different denominators such as participation requests, registered teams, evaluation-system sign-ups, and public-leaderboard entrants. The 4th edition attracted 315 participating teams across 37 countries; the 5th edition attracted 305 participating teams across 38 countries; the 6th edition received participation requests from 254 teams across 27 countries; the 7th edition drew participation requests from 508 teams across 46 countries; the 8th edition featured 726 registered teams in 47 countries and regions; and the 9th edition reported 245 teams from 15 countries registered on the evaluation server, with a 17% increase in participation and more than 30,000 dataset downloads to date (Naphade et al., 2020, Naphade et al., 2021, Naphade et al., 2022, Naphade et al., 2023, Wang et al., 2024, Tang et al., 19 Aug 2025).
A recurring characteristic of the challenge is its movement beyond road scenes alone. The published editions show a transition from predominantly traffic-centric tasks toward a broader urban-AI agenda that includes warehouse layouts, retail checkout lanes, indoor occupancy analytics, and multimodal reasoning. This suggests that “AI City” functions less as a narrow ITS benchmark than as a family of deployable perception-and-reasoning problems situated in urban and operational infrastructure.
2. Competition architecture and governance
The challenge is organized as a set of independent tracks, each with its own dataset, metric, and submission protocol. Earlier editions consistently exposed two leaderboards: a Public leaderboard restricted to released data and contest rules, and a General leaderboard allowing broader methods or external data. The Public leaderboard was intended to reflect more realistic deployment conditions where annotated data are limited, while the General leaderboard captured broader state-of-the-art exploration. In the 9th edition, the evaluation framework continued the emphasis on fairness and reproducibility by prohibiting external private data on the public leaderboard and requiring code and models to be released (Naphade et al., 2020, Naphade et al., 2021, Naphade et al., 2023, Tang et al., 19 Aug 2025).
A standard anti-overfitting mechanism recurs across editions. During the competition, scores are computed on a partial test subset—typically 50% of the test data—while final rankings are revealed after the deadline on the full test set. In the 9th edition, the leaderboard was computed on 50% of test frames and only the top 3 ranks were shown during the contest; after the deadline, full test-set evaluation including held-out splits produced the final rankings. Submission caps also constrained leaderboard probing. For 2025, daily limits were 5 submissions per track, with total caps of 30 for Track 1, 20 for Tracks 2 and 3, and 50 for Track 4 (Wang et al., 2024, Tang et al., 19 Aug 2025).
This evaluation design addresses a recurring benchmark concern: public-server adaptation can inflate apparent progress through repeated tuning to visible scores. The challenge’s partially held-out test regime, code-release rules, and delayed disclosure of final rankings were explicitly introduced to foster reproducibility and mitigate overfitting. A plausible implication is that the challenge has increasingly treated benchmark governance as part of the research problem, not merely an administrative detail.
3. Task families and data resources
The benchmark spans a heterogeneous but coherent set of task families: tracking, retrieval, captioning and question answering, action recognition, anomaly detection, object detection under distortion, retail checkout, and spatial reasoning. Some datasets are city-scale real video corpora, while others are synthetic or hybrid resources constructed to support controlled variation, rare events, or precise geometry.
| Domain | Representative tracks | Representative datasets |
|---|---|---|
| ITS tracking and retrieval | vehicle counting, vehicle ReID, MTMC tracking, NL retrieval | CityFlow, CityFlowV2, CityFlow-NL, VehicleX |
| Traffic safety and road perception | anomaly detection, dense captioning, VQA, fisheye detection, helmet violation detection | WTS, BDD100K clips, FishEye8K, FishEye1K_eval, Iowa DOT anomaly dataset |
| Indoor and operational AI | people tracking, automated retail checkout, warehouse spatial reasoning | Omniverse MTMC datasets, ARC, “Physical AI Smart Spaces” |
CityFlow and its successors underpin several ITS tracks. CityFlow in the 4th edition provided approximately 215 minutes of 10 fps video from 46 cameras over 16 intersections, with 880 annotated IDs and approximately 300 K bounding boxes for MTMC vehicle tracking. CityFlowV2 in the 5th and 6th editions retained the 46-camera, 16-intersection structure and supplied 313,931 manually refined bounding boxes for 880 distinct vehicle identities. CityFlow-NL supported natural-language retrieval by pairing vehicle tracks with crowd-sourced descriptions, while VehicleX supplied 1,362 synthetic 3D vehicle identities and more than 190,000 rendered images for synthetic augmentation and domain adaptation (Naphade et al., 2020, Naphade et al., 2021, Naphade et al., 2022).
Traffic-safety and fisheye tasks introduced more specialized resources. The Woven Traffic Safety dataset in the 9th edition included 1,200+ staged pedestrian–vehicle events and 4,800 filtered real-world BDD100K clips, with synchronized overhead, vehicle-mounted, and ego-centric views segmented into five semantic phases from Pre-recognition to Avoidance. The same track incorporated 3D gaze and head-pose labels collected via Tobii Pro Glasses and projected into a 3D scene model for each frame. For fisheye analytics, FishEye8K contains 5,288 training images and 2,712 validation images, with 157K bounding boxes in VOC, COCO, and YOLO formats, while FishEye1K_eval provides 1,000 test images from novel cameras (Wang et al., 2024, Tang et al., 19 Aug 2025).
The indoor and warehouse tracks show the strongest move toward synthetic scene generation. The 9th-edition Track 1 dataset, “Physical AI Smart Spaces,” comprises 42 h video, 504 cameras, 19 scene layouts, and approximately 360 object instances, with detailed calibration, 3D bounding boxes parameterized as , and optional depth maps. Both Tracks 1 and 3 in 2025 were generated in NVIDIA Omniverse, using Replicator Agent (IRA) for agent trajectories and Replicator Object (IRO) for procedural scene composition; Track 3 added RGB-D inputs, 3D boxes, 2D masks, YAML-driven scene configurations, and approximately 500K automatically generated VQA samples (Tang et al., 19 Aug 2025).
Synthetic data has therefore been central rather than peripheral. In earlier ITS editions it was used to enlarge scarce real training data and reduce domain gap; in later indoor editions it became the substrate for calibrated 3D annotation, controlled motion, and spatial-language question generation. This suggests a shift from synthetic data as augmentation toward synthetic environments as benchmark infrastructure.
4. Evaluation methodology and metrics
The challenge uses task-specific metrics chosen to reflect deployment objectives rather than a single universal score. For MTMC and MOT-style tasks, editions have employed IDF, MOTA, and, more recently, HOTA and 3D HOTA. In the 9th edition, 3D multi-object tracking accuracy was defined as
while 3D HOTA was used as the primary ranking measure for Track 1, with a 10% bonus for online methods. The 8th edition had already upgraded MTMC people tracking from IDF to HOTA operating on 3D distances and explicitly encouraged online methods through the same 10% bonus (Wang et al., 2024, Tang et al., 19 Aug 2025).
Retrieval and detection tracks rely on standard ranking and detection metrics. Vehicle ReID was ranked by mean Average Precision over the top matches, with CMC@K reported for analysis in earlier editions. Natural-language-based vehicle retrieval used Mean Reciprocal Rank,
Fisheye and helmet-detection tracks used F or mAP, depending on the edition. The 9th-edition fisheye track introduced an explicit speed–accuracy trade-off through the harmonic mean of F1 and normalized FPS:
where
Benchmarking was performed using Dockerized TensorRT containers on Jetson AGX Orin (64 GB), and the task required at least 10 FPS on Jetson AGX Orin (Naphade et al., 2021, Wang et al., 2024, Tang et al., 19 Aug 2025).
Language-centered tracks employ conventional NLP metrics but also expose their limitations. Dense traffic-safety captioning in the 8th edition used the equal average of BLEU-4, METEOR, ROUGE-L, and CIDEr; the 9th edition Track 2 combined QA accuracy with BLEU-4, METEOR, ROUGE-L, and CIDEr for captions and VQA. Spatial reasoning in Track 3 of 2025 used weighted QA accuracy. Driver-action recognition used the Average Activity Overlap Score, defined over matched predicted and ground-truth intervals within a s window (Naphade et al., 2023, Wang et al., 2024, Tang et al., 19 Aug 2025).
A notable methodological conclusion appears repeatedly: metric choice influences what is optimized. The challenge reports explicitly that conventional NLP scores are insufficient to capture semantic fidelity in safety descriptions, indicating that captioning and VQA tracks remain partially misaligned with the causal and grounding demands of incident analysis.
5. Technical approaches and benchmark results
The published leaderboards show strong continuity in system design: high-performing entries often combine strong detectors, single-camera trackers or temporal encoders, explicit geometric or spatio-temporal constraints, synthetic-data augmentation, post-processing, and, where useful, ensembling. In the 4th edition, the top vehicle-counting system from Baidu achieved 0 with Faster R-CNN, DeepSORT, and movement-specific trajectory clustering. In the 5th edition, Alibaba ranked first in Track 2 with mAP@100 of 0.7445, while Alibaba-UCAS ranked first in MTMC vehicle tracking with IDF1. The dedicated first-place Track 2 study “An Empirical Study of Vehicle Re-Identification on the AI City Challenge” details a pipeline built from CityFlowV2 and VehicleX, weakly supervised cropping, SPGAN style transfer, unsupervised domain adaptation via DBSCAN pseudo-labels, k-reciprocal re-ranking, image-to-tracklet retrieval, camera and orientation bias fusion, and an ensemble of 16 models (Naphade et al., 2020, Naphade et al., 2021, Luo et al., 2021).
The later editions reveal domain-specific specialization. In the 7th edition, UW-ETRI won indoor MTMC people tracking with IDF2 using YOLOv5, OSNet, BoT-SORT, and anchor-guided Hungarian clustering with spatio-temporal consistency; HCMIU won natural-language vehicle retrieval with MRR of 0.8263 via a CLIP-based vision-language encoder, semi-supervised domain adaptation, and multi-contextual pruning; and CTC won helmet-rule violation detection with mAP of 0.8340 using DETA with a Swin-L backbone, SORT tracking, and trajectory grouping for rider-to-helmet assignment (Naphade et al., 2023).
The 8th and 9th editions further expanded the role of multimodal and geometry-aware systems. In 2024, AliOpenTrek won dense video captioning for traffic safety with an average 4-metric score of 33.43 using LLaVA-1.6-34B, Qwen-VL, a Vicuna decoder, a novel visual prompt schema, and global plus local views. In 2025, Team 65 (ZV) led Track 1 with a 3D HOTA of 69.912 through geometry-centric offline fusion and global association; Team 145 (CHTTLIOT) led Track 2 with a caption × VQA mean score of 60.039 through joint QA–caption training and role-aware prompts; Team 16 (UWIPL_ETRI) led Track 3 with weighted QA accuracy of 95.864% using a modular LLM+tool API approach; and Team 33 (UIT-OpenCubee) led Track 4 with a score of 0.6493 using distortion-aware augmentations and a dual-model YOLOv11/D-FINE ensemble (Wang et al., 2024, Tang et al., 19 Aug 2025).
Several benchmark-level patterns recur. Offline geometry fusion still outperforms online tracking in the 2025 3D MTMC task, where online methods trailed the leading offline system by roughly 36–44 points. For traffic-safety captioning and QA, marginal gaps of less than 2 points among the top five systems indicate a more mature VLM-based pipeline regime. For warehouse spatial reasoning, the approximately 4% gap between first and second place suggests the benefit of explicit geometry APIs. For efficient fisheye detection, sub-1% differences among the top five teams indicate that balancing accuracy and speed on edge hardware is now the principal bottleneck rather than raw detector quality alone (Tang et al., 19 Aug 2025).
6. Recurring themes, misconceptions, and open problems
One common misconception is that the AI City Challenge is only a traffic benchmark. The documented editions show otherwise: beyond ITS, it has included automated retail checkout, indoor people tracking, warehouse spatial reasoning, and public-safety VQA. Another misconception is that the benchmark is purely perception-oriented. The more recent tracks incorporate language, question answering, gaze, causal incident phases, RGB-D geometry, and explicit tool-augmented reasoning, indicating a progression from object-centric vision toward operational multimodal intelligence (Naphade et al., 2022, Naphade et al., 2023, Tang et al., 19 Aug 2025).
Several scientific themes persist across editions. Multimodal fusion is repeatedly identified as central: depth plus RGB for tracking, gaze plus video for safety analysis, RGB-D for reasoning, and spatial-plus-language fusion for QA. Synthetic-to-real transfer is another enduring axis, visible in VehicleX plus SPGAN for vehicle ReID, Unity and Omniverse generation for retail and indoor tasks, and pseudo-label or domain-adaptation strategies throughout the benchmark literature. A third theme is the tension between offline accuracy and streaming constraints: the challenge has repeatedly rewarded or encouraged online methods, yet the strongest systems in several tracks remain offline or heavily post-processed (Luo et al., 2021, Wang et al., 2024, Tang et al., 19 Aug 2025).
The benchmark literature also records explicit limitations. Persistent challenges include occlusion and ID switches in crowded indoor and outdoor scenes, peripheral object detection under extreme fisheye distortion, semantic grounding for “why” questions in traffic incidents, domain adaptation, and lightweight inference for edge deployment. For language-centered traffic-safety tasks, conventional syntactic metrics do not adequately capture semantic fidelity, grounding, or causal reasoning. For MTMC systems, dependence on manual rules, zones, or camera-specific heuristics has repeatedly been identified as a scalability limitation. For naturalistic driving analysis and helmet-rule detection, cross-environment and cross-weather generalization remain open (Naphade et al., 2022, Naphade et al., 2023, Wang et al., 2024, Tang et al., 19 Aug 2025).
Published recommendations for future editions follow directly from these constraints. The 9th edition proposes counterfactual QA in the safety track, expanded held-out splits to stress generalization to new cities, unseen warehouse layouts to test reasoning transfer, and automated low-light and weather variability for fisheye detection. Earlier editions call for automatic multi-camera calibration, larger and more varied real driving datasets, end-to-end trainable multimodal retrieval systems, integrated distortion-aware backbones, and real-time edge-deployable solutions. Taken together, these recommendations indicate that the challenge’s trajectory is toward benchmarks that are simultaneously more realistic, more multimodal, more geometry-aware, and more constrained by deployment requirements than by laboratory-only accuracy (Naphade et al., 2022, Naphade et al., 2023, Wang et al., 2024, Tang et al., 19 Aug 2025).