UrbanFeel: Computational Urban Perception
- UrbanFeel is a computational framework that captures urban perceptions through multimodal data like street imagery, sensor readings, and digital narratives.
- It leverages methods such as spatial tessellation, metric learning, and semantic mobility modeling to quantify attributes like beauty, safety, and vibrancy.
- Its actionable insights assist urban planning by identifying stress hotspots and guiding interventions for more livable urban environments.
UrbanFeel denotes the computational characterization of how urban places are perceived, experienced, and inhabited. In the current literature, the term covers both a specific benchmark for multimodal LLMs and a broader research program that models the “feel,” “vibe,” or wellbeing-related qualities of neighborhoods from street-view imagery, satellite imagery, point-of-interest data, mobility traces, environmental signals, physiological sensing, and real-time digital narratives (He et al., 26 Sep 2025, Wang et al., 2020, Johnson et al., 2020). Across these strands, UrbanFeel is not restricted to a single scalar notion of quality: it includes subjective perception attributes such as beauty, safety, wealth, and liveliness, personal preference, collective emotional response, accessibility to centralities and nature, and microclimatic comfort (Xi et al., 14 Jun 2026, Alvarez-Marin et al., 2020, D'Acci, 2014).
1. Conceptual foundations
UrbanFeel is grounded in the view that urban environments are experienced through layered social, perceptual, and material signals rather than only through administrative boundaries or static land-use classes. ConnectiCity formulates the city as a “membrane of information” and a “read/write, ubiquitous publishing surface,” where urban space is continuously rewritten by mobile devices, social media, wearables, location based services, and mixed or augmented reality (Iaconesi et al., 2012). This framing shifts urban landscape from a purely administrative definition to “a lossless sum of their perceptions,” making felt urban experience a legitimate computational object.
A second foundational strand treats UrbanFeel as semantically structured collective behavior. MobInsight argues that collective urban mobility embodies residents’ local insights, because destination choices reflect atmosphere, distance, past experiences, and preferences rather than only generic mobility laws (Park et al., 2017). In that formulation, nightlife-centric, cultural, family-oriented, professional, and recreational neighborhoods become measurable through semantic neighborhood features and all-pairs mobility modeling.
A third strand emphasizes subjectivity. “Indexical Cities” explicitly rejects an objective measure of urban quality and instead models a single observer’s likeability of places, rendering a personal “index” onto georeferenced cities (Alvarez-Marin et al., 2020). This position is important because it establishes that UrbanFeel can be observer-specific rather than crowd-averaged.
A fourth strand is normative and spatial. “Urban DNA for cities evolutions” frames experiential urban quality as an outcome of dynamic equilibria between competitive and cooperative forces and proposes Isobenefit Urbanism, where equal-compensative urban quality, proximity to centralities, and proximity to “real nature” are central design constraints (D'Acci, 2014). This suggests that UrbanFeel is not only descriptive but also prescriptive: it can be used to articulate what cities should avoid in order not to become “unideal.”
2. Data modalities and representational scope
UrbanFeel research is multimodal by construction. The operational signal may be street-level appearance, semantic amenity mix, physiological arousal, self-reported valence, mobility flows, or annual environmental indicators. The main families of inputs are summarized below.
| Paradigm | Primary inputs | UrbanFeel target |
|---|---|---|
| In-situ wellbeing sensing | GPS, heart rate, self-reported valence | Stress-inducing and restorative places |
| Neighborhood embedding | Street view, POIs, ratings, reviews, geospatial context | Joint visual–textual–geospatial “feel” |
| Benchmarking perception and change | Multi-temporal street view, panoramas, satellite imagery | Static perception, temporal change, subjective perception |
| Personalized preference mapping | Satellite imagery, Street View, pairwise ratings | Probability of like for a specific observer |
In Nottingham, sensor-based UrbanFeel was instantiated through continuous, multi-modal sensor data, GPS coordinates, duration metadata, and a simplified 5-step Self-Assessment Manikin valence scale collected from participants, producing 550,432 lines of location traces and 5,345 self-report responses (Johnson et al., 2020). The core analyzed modalities were heart rate and self-reported valence, with GPS-tagged timestamps used for spatial aggregation.
Urban2Vec operationalizes UrbanFeel at neighborhood scale by combining Google Street View Static API imagery with Yelp Fusion API POIs, where each POI is textualized as a bag-of-words containing category, rating, price, and reviews (Wang et al., 2020). The goal is a compact vector for each census tract that captures streetscape appearance, amenity ecosystem, and local sentiment.
UrbanWell extends the scope to annual, grid-aligned urban wellbeing analytics. It spans 38 cities from 2012 to 2024 and aligns satellite imagery, street-view imagery, and 19 indicators covering environmental conditions, spatial accessibility, urban form, urban vitality, and subjective perception at grid-year level (Xi et al., 14 Jun 2026). Its perception attributes include Safe, Beautiful, Lively, Boring, Depressing, Wealthy, and Quietness Suitability Index.
The UrbanFeel benchmark is narrower in input type but broader in cognitive design. It contains over 14,300 visual question-answer pairs built from multi-temporal single-view and panoramic street-view images from 11 cities over 2007–2024, organized around Static Scene Perception, Temporal Change Understanding, and Subjective Environmental Perception (He et al., 26 Sep 2025).
3. Core computational methods
A central methodological divide in UrbanFeel is between explicit spatial aggregation and latent representation learning. In the sensor-analytic tradition, the defining operation is spatial tessellation. “Sensor Data and the City” uses Voronoi diagrams to aggregate GPS-tagged heart rate and self-reported affect, with cells defined by
and
The motivation is that heatmaps suffer from cell-size selection and density heterogeneity issues, whereas Voronoi cell size implicitly conveys local sampling density and adjacency suggests local similarity (Johnson et al., 2020).
In multimodal embedding work, the emphasis is metric learning. Urban2Vec first trains street-view embeddings with geospatial regularization using triplets of nearby and far images, then initializes neighborhood vectors by averaging image embeddings,
and finally updates neighborhood vectors and word embeddings with neighborhood–POI triplet loss (Wang et al., 2020). Its image context size is nearest images, the embedding dimension is , Euclidean distance is the metric, and the staged optimization was empirically superior to simultaneous training.
Semantic mobility modeling provides a third computational regime. MobInsight collects more than 128,000 places from 15 heterogeneous online sources in Barcelona, maps source metadata to word vectors, reduces dimensionality with Latent Semantic Analysis to 100 dimensions explaining 77% of variance, clusters with chosen by silhouette saturation around 0.7, and manually merges clusters into 17 interpretable categories plus “Total place count” (Park et al., 2017). Neighborhood features are then standardized as
and passed to a softmax neural model, with evaluation by Kullback-Leibler divergence over destination distributions.
Entropy-based mobility profiling constitutes another distinct UrbanFeel formalism. UrbanFACET derives Fluidity, vibrAncy, Commutation, divErsity, and densiTy from billions of mobile-device records by computing user-level and record-level entropies over POI classes and administrative divisions (Shi et al., 2017). Vibrancy is the Shannon entropy of a user’s POI-class distribution; Commutation is the entropy of a user’s administrative-division distribution; Diversity and Fluidity redistribute those entropies to records, thereby highlighting atypicality and occasional inflow at place level. This produces a citywide behavioral profile that is not reducible to density alone.
Personalized UrbanFeel adopts a different pipeline. “Indexical Cities” extracts 4,096-D VGG features from satellite and Street View images, uses t-SNE and Self-Organizing Maps to produce an “alphabet” of spatialities, then learns an individual like/dislike classifier from pairwise comparisons (Alvarez-Marin et al., 2020). The compact labeling protocol uses 513 representative centroids and 1,500 image pairs; the trained Street View classifier reports recall of approximately 87% and precision of approximately 90%, after which taste is transferred to satellite tiles via geolocation.
4. Benchmarks and multimodal model evaluation
UrbanFeel has become a benchmarked task family for MLLMs. The 2025 UrbanFeel benchmark defines 11 tasks across three dimensions—Static Scene Perception, Temporal Change Understanding, and Subjective Environmental Perception—and evaluates 20 state-of-the-art MLLMs in zero-shot settings (He et al., 26 Sep 2025). Its four QA formats are binary judgment, multiple choice, open-ended reasoning, and temporal sorting. The benchmark reports human overall accuracy of 67.4%, Gemini-2.5 Pro overall accuracy of 65.9%, and an average human–model gap of about 1.5%.
The benchmark’s most important result is asymmetry across cognitive demands. Models are strong on scene understanding and Time-Consistent Recognition, and some models outperform humans in Panoramic Change Recognition, where GPT-4o reaches 40.5% and Qwen2.5-VL-72B reaches 40.9% against a human score of 21.2% (He et al., 26 Sep 2025). By contrast, Temporal Sorting Reasoning remains difficult: most models score below 10%, while Gemini-2.5 Pro reaches 52.1% against human 70.0%. Single-view inputs also outperform panoramic inputs on average by 11.7%, indicating that panorama-induced geometric distortion and dense context remain a nontrivial bottleneck.
UrbanWell complements that benchmark by aligning perception with annual urban indicators rather than only visual question answering. It evaluates 15 representative MLLMs on single-year estimation, multi-year forecasting, and temporal trend analysis over 38 cities and 19 indicators, using one 256×256 satellite image and 1–4 street-view images per grid-year (Xi et al., 14 Jun 2026). The benchmark shows that perception attributes are more visually grounded than many environmental quantities, while temporal context dramatically stabilizes prediction. In multi-year forecasting, Llama4-Scout achieves NDVI RMSE of 0.35, NO RMSE of 2.42, and PM0 RMSE of 1.92; perception forecasting also improves, with BEA RMSE of 1.35 and BOR RMSE of 0.99 for top models (Xi et al., 14 Jun 2026).
UrbanWell also shows that trend analysis remains substantially harder than value estimation. Overall accuracies in temporal trend classification are often in the 0.2–0.4 range, and models tend to over-smooth temporally by predicting stable when subtle changes are present (Xi et al., 14 Jun 2026). This closely matches the failure mode observed in the UrbanFeel benchmark: current multimodal models can read salient spatial and perceptual cues, but longitudinal urban reasoning is still comparatively weak.
5. Collective sensing, real-time perception, and personalized UrbanFeel
One important lineage of UrbanFeel is direct sensing of embodied urban response. In Nottingham, Voronoi overlays of heart rate and valence revealed that stressful hotspots were scattered along the path rather than confined to one area (Johnson et al., 2020). In a fragrance and beauty shop, clusters of red polygons aligned with moments when participants responded to shop assistants, suggesting negative impact on well-being; in clothes shops, darker polygons indicated increased heart rate, with a reported trend of heart rate increasing when encountering discounted items. At street level, heart-rate Voronoi overlays highlighted areas of increased heart rate and inferred stress on the shopping street. The study is qualitative rather than inferential: no correlation coefficients, 1-values, or effect sizes are reported.
Another lineage treats UrbanFeel as a real-time narrative layer. ConnectiCity’s deployments harvest RSS feeds, APIs, microformats, websites, blogs, and institutional sources, then apply geo-parsing, geo-coding, NLP, emotion analysis, and relevancy thresholding to render urban narratives in public space (Iaconesi et al., 2012). Atlas of Rome used eight synchronized projections across a 35-meter surface; geo-extraction correctness was reported at about 97%. In VersuS Rome, more than 92,000 relevant elements were selected during the protest timeframe, with post-event simulations reporting around 30,000 geo-referenced elements for the police app and approximately 12,000 violence messages, approximately 10,000 law infringement mentions, and approximately 3,000 injury messages.
Personalized UrbanFeel follows a different epistemology. “Indexical Cities” defines city liking probabilistically for a specific observer rather than a collective average, using 20 cities over five continents, 50,000 geolocated satellite tiles, and 32,000 corresponding Street View pictures (Alvarez-Marin et al., 2020). Its specific pixel maps use a warm/cold spectrum in which cold tones indicate preference and warm tones indicate aversion. This formulation makes UrbanFeel explicitly non-universal: the same urban fabric can generate different “Cities of Indexes” for different observers.
Collective mobility can also be read as an aggregate expression of felt urban difference. MobInsight’s Barcelona case studies show that Raval, Barceloneta, Pedralbes, Sarrià, and Poblenou/22@ exhibit distinct semantic profiles and mobility effects, including nightlife and shopping in Raval, club and eating dominance in Barceloneta, service-sector and education signals in Pedralbes, and office-centric dynamics around 22@ (Park et al., 2017). UrbanFACET reaches a related conclusion from massive movement data, showing that entropy-based Fluidity, Vibrancy, Commutation, Diversity, and Density differentiate tourist hotspots, commuting structures, diversified high-value areas, and low-density mountain regions across Beijing, Tianjin, Tangshan, and Zhangjiakou (Shi et al., 2017).
6. Planning, simulation, and design applications
UrbanFeel has increasingly been tied to intervention rather than only diagnosis. The Nottingham sensing study explicitly frames its outputs as actionable for city planners and retailers: store interactions, queueing, signage, crowd flows, promotional layouts, and micro-restorative design elements such as green or calming spaces, seating, shade, acoustic buffering, and wayfinding can be redesigned where negative sentiment or elevated arousal clusters are observed (Johnson et al., 2020).
At neighborhood scale, Urban2Vec supports similarity search, clustering, and transfer to downstream tasks such as socio-economic and real estate prediction (Wang et al., 2020). On Bay Area real estate prediction, Urban2Vec reports 2 for average apartment sale price and 3 for average office sale price, outperforming AE+POI and SL+POISTATS baselines. Its interpretability analyses further associate the first principal component with street enclosure and the second with vegetation, income, and education.
Mobility semantics allow “what-if” planning through feature perturbation rather than population-based proxies. MobInsight reports that NF_Dist achieves approximately 35% relative improvement versus the Average-based model and approximately 30% versus the Gravity model, while NF_Dist outperforming NF_woPub_Dist by approximately 15% demonstrates the importance of including government open directory data (Park et al., 2017). This supports planning applications in which adding or removing categories of places can be used to simulate mobility changes.
Outdoor comfort extends UrbanFeel into physics-based environmental design. UrbanFlow targets windy areas and heat pockets with a RANS-based Eulerian simulator using a unified porosity model for buildings and trees (Liu et al., 2022). In the Nuremberg competition site, the optimization objective pushed region-average pedestrian-level speeds toward 4, used 11 degrees of freedom, 6 target regions, 14 iterations, and 12 forward simulations per iteration, with total runtime of approximately 2 hours on a single workstation. The reported outcome was mitigation of heat pockets without creating wind discomfort.
UrbanGraph generalizes comfort prediction to a physics-informed, heterogeneous, dynamic spatio-temporal graph. It encodes shading, vegetation evapotranspiration, and convective diffusion as relation types and reports 5 and RMSE 6 on UTCI, improving 7 by up to 10.8% and reducing FLOPs by 17.0% over baselines (Xin et al., 1 Oct 2025). This indicates that UrbanFeel can be coupled not only to perceptual analytics but also to physically grounded microclimate forecasting.
Generative visualization adds a further design layer. UrbanGIRAFFE represents urban scenes as compositional generative neural feature fields conditioned on a semantic voxel grid 8 and object layout prior 9, enabling large camera movement, stuff editing, and object manipulation (Yang et al., 2023). On KITTI-360, it reports image-level FID0 of 39.6 and KID1 of 0.036, substantially outperforming GIRAFFE and GSN on that dataset. This suggests a role for UrbanFeel not only in measurement but also in controllable simulation of alternative streetscape configurations.
7. Limitations, biases, and unresolved problems
A persistent limitation is that UrbanFeel is partly subjective and partly proxy-based. The UrbanFeel benchmark reports strong model consistency for beauty and safety, but wealth judgments lag human performance by approximately 10.1% on average, and city identity interventions show that some models shift their judgments when labels such as Paris, Cape Town, or “Global North” are provided (He et al., 26 Sep 2025). UrbanWell likewise notes that subjective perception labels are proxy scores from pretrained Place Pulse-derived models and may inherit training-data bias (Xi et al., 14 Jun 2026).
Sampling bias is equally pervasive in sensing and mobility-based approaches. The Nottingham study used an all-female cohort with average age 28, under similar weather around 11am, which limits demographic and temporal generalizability (Johnson et al., 2020). Urban2Vec depends on Google Street View and Yelp coverage, both of which vary spatially and demographically (Wang et al., 2020). MobInsight mitigates tourist and commercial skew by incorporating government directory data, but its CDR-derived mobility still reflects operator-specific bias and a one-month temporal window (Park et al., 2017). UrbanFACET, despite its enormous scale, is conditioned on app usage patterns and device ownership, and retains only GPS and Wi-Fi records (Shi et al., 2017).
Privacy and governance remain unresolved. ConnectiCity explicitly references k-anonymity and privacy-preserving data mining while also acknowledging digital-divide issues and interpretation errors (Iaconesi et al., 2012). The Nottingham paper does not detail privacy procedures and states that an UrbanFeel deployment should include consent, anonymization of GPS traces, and aggregation thresholds to avoid re-identification (Johnson et al., 2020). Personalized approaches such as Indexical Cities intensify these concerns because they model a single observer’s taste rather than an aggregate tendency (Alvarez-Marin et al., 2020).
Methodologically, temporal reasoning is a major open problem. The UrbanFeel benchmark shows a large deficit in Temporal Sorting Reasoning, and UrbanWell reports modest accuracies in temporal trend classification with systematic over-smoothing (He et al., 26 Sep 2025, Xi et al., 14 Jun 2026). Cross-cultural calibration, panorama-aware modeling, standardized justification evaluation, network-based accessibility, noise-sensitive quietness metrics, and tighter integration between perceptual, environmental, and behavioral signals are all identified as future directions in the surveyed literature.
UrbanFeel therefore remains a heterogeneous but increasingly coherent domain: one part benchmark for multimodal urban reasoning, one part sensing-and-visualization pipeline for urban wellbeing, one part embedding framework for neighborhood character, and one part planning instrument for comfort, accessibility, and experiential equity. Its unifying premise is that urban environments can be studied not only as infrastructures or administrative units, but as places that are felt.