MoveFM-R: Advancing Mobility via Language
- The paper introduces MoveFM-R, a framework that bridges the vocabulary gap between continuous GPS data and discrete language tokens through semantically enhanced location encoding and bidirectional instruction-tuning.
- The framework employs a progressive curriculum and interactive self-reflection mechanism to align MFM latent sequences with LLM reasoning, leading to more accurate trajectory prediction and generation.
- MoveFM-R demonstrates significant improvements in prediction and generation metrics, achieving 12%–30% gains over state-of-the-art models by integrating mobility models with language-driven semantic reasoning.
MoveFM-R is a framework for Mobility Foundation Models (MFMs) that integrates language-driven semantic reasoning with spatio-temporal mobility modeling. It is introduced in the paper "MoveFM-R: Advancing Mobility Foundation Models via Language-driven Semantic Reasoning" (Meng et al., 26 Sep 2025). The framework is motivated by a specific asymmetry: MFMs learn the statistical and geometric regularities of GPS trajectories but lack high-level semantic grounding, whereas LLMs provide strong semantic reasoning but do not possess an innate understanding of continuous geography or spatio-temporal statistics. MoveFM-R addresses this asymmetry through three linked components: a semantically enhanced location encoding, a progressive curriculum for aligning MFM latent sequences with LLM reasoning, and an interactive self-reflection mechanism for conditional trajectory generation.
1. Problem setting and architectural objective
MoveFM-R is designed around two stated bottlenecks in mobility modeling. The first is a vocabulary mismatch between continuous geographic coordinates and discrete language tokens. The second is a representation gap between MFM latent vectors and the semantic world of LLMs. The framework therefore treats the interface between mobility data and language not as a direct coordinate-to-token conversion problem, but as an alignment problem across heterogeneous representational spaces (Meng et al., 26 Sep 2025).
The model rationale is explicit. Rather than asking an LLM to infer mobility structure directly from raw coordinates, MoveFM-R first constructs semantically meaningful discrete location tokens, then teaches the LLM how to consume mobility representations through a curriculum. This organization also clarifies a common misconception in LLM-based mobility modeling: semantic capacity alone is not sufficient for physically plausible trajectory synthesis. The paper states that LLMs lack the built-in understanding of spatio-temporal statistics required for generating physically plausible mobility trajectories, while MFMs face a ceiling due to limitations in data scale and semantic understanding.
The three core innovations are presented as complementary. The semantically enhanced location encoding bridges the geography-language gap. The progressive curriculum alignment teaches the LLM "how to read" MFM latent sequences. The interactive self-reflection loop supports conditional generation under scenario constraints by editing a baseline trajectory with minimal changes. A plausible implication is that MoveFM-R is not merely an MFM augmented with prompts, but a multimodule alignment scheme in which symbolic location tokens, latent trajectory embeddings, and natural-language instructions are jointly coordinated.
2. Semantically enhanced location encoding
The location encoding pipeline begins with a large multi-city dataset of semantic location profiles. Each profile concatenates address strings and counts of 34 POI categories, and a pre-trained text encoder maps the profile to an embedding . A Residual-Quantized VAE (RQ-VAE) then discretizes this embedding into a sequence of codewords drawn from codebooks (Meng et al., 26 Sep 2025).
At layer , the quantization step is defined as
The training loss combines residual quantization and reconstruction:
with total objective
This design matters because the paper explicitly states that the vocabulary mismatch is solved by quantizing a high-dimensional semantic space rather than raw coordinates. That choice distinguishes semantic tokenization from naive discretization of latitude-longitude pairs.
After discretization, the codebook is aligned to the LLM in two stages. In static embedding alignment, each new token embedding is initialized as the average of its subword pieces and refined with
0
where
1
2
In bidirectional instruction-tuning, the LLM is fine-tuned so that it can map Location-ID 3 Description and Description 4 Location-ID using a standard sequence-to-sequence negative log-likelihood objective. This second stage gives the new location tokens a role inside the LLM’s generative and interpretive routines, not merely inside its embedding table.
3. Curriculum alignment between mobility latents and language reasoning
MoveFM-R does not assume that an LLM can directly interpret an MFM latent sequence once a token vocabulary exists. Instead, it introduces a two-phase curriculum that is interleaved during training (Meng et al., 26 Sep 2025). The mobility representation is formed as
5
where 6 is the frozen MFM encoder.
In Phase 1: Low-level description, the LLM is prompted to translate a latent sequence into factual text such as "At time 7, visited location 8." In Phase 2: High-level summarization, the LLM is prompted to infer more abstract patterns, including "Most frequent locations" and "Probability of visits by time window," before performing prediction or generation. The three subtasks—description, summarization, and next-token prediction or generation—are trained jointly via
9
The paper describes this as imposing an "understand 0 predict/generate" inductive bias. The key point is that high-level mobility reasoning is not treated as an emergent side effect of sequence modeling; it is directly trained as an intermediate behavior. This addresses the stated representation gap through two mechanisms: a lightweight MLP that projects the MFM trajectory embedding 1 into the LLM embedding space, and LLM fine-tuning that conditions on these projected inputs in context.
This suggests a particular interpretation of MoveFM-R’s alignment strategy. The curriculum is not only a training schedule but also a representation regularizer: factual description constrains the model to remain grounded in trajectory content, while summarization encourages compression into semantically meaningful mobility abstractions.
4. Interactive self-reflection for conditional generation
For conditional trajectory generation, MoveFM-R does not generate a scenario-constrained trajectory end-to-end. The procedure is explicitly iterative (Meng et al., 26 Sep 2025). First, the model generates an unconditional baseline trajectory 2 from the user’s history. It then enters a reflection loop in which, at each iteration 3, it compares 4 against the scenario constraints and, if a violation is detected, proposes a single edit from the action set {add point, delete point, modify point} to produce 5. The process terminates when all constraints are met.
The reward function prioritizes consistency with a true trajectory 6:
7
8
The policy is trained with Group Relative Policy Optimization (GRPO) and is initialized by supervised fine-tuning on a small synthetic corpus to stabilize the format. The conditional-generation mechanism is therefore organized around edit-based repair rather than direct constrained decoding.
A recurring misunderstanding in scenario-based mobility generation is that conditioning can be handled simply by prepending an instruction. MoveFM-R formalizes a different view: instructions are checked against a generated trajectory, and violations are corrected through targeted edits. A plausible implication is that the self-reflection stage functions as a structured feasibility operator over natural-language constraints, rather than as a generic stylistic prompt.
5. Experimental protocol and reported performance
The empirical study uses four real-world U.S. cities: Atlanta, Chicago, Seattle, and Washington D.C. Each city is discretized into 500 m grid cells and half-hour time bins. Training instances use 3-day sliding windows, with sequences under 5 points or over 145 points discarded, keeping the most recent points. Semantic POI features are drawn from OpenStreetMap (Meng et al., 26 Sep 2025).
The reported metrics are divided by task. For prediction, the metric is Hit Rate@1 (HR@1). For unconditional generation, the metrics are BLEU, Total Variation Distance (TVD), and Jensen–Shannon Divergence (JSD), each computed separately over time and location distributions. Conditional generation uses the same BLEU/TVD/JSD suite under three scenarios: late-night commuters, temporary travel plans, and weekend users. The baselines include MFM-based prediction models (DeepMove, GETNext, TrajBert, TrajFM, Unitraj, TrajMoE), LLM-based prediction models (Mobility-LLM, QT-Mob), and generation models comprising diffusion methods (DiffTraj, Marionette) and LLM-chain methods (COPB, LLMob). The zero/few-shot protocol trains on three cities and tests on the fourth in zero-shot mode, then fine-tunes on 500 examples for few-shot evaluation.
For next-location prediction, MoveFM-R reaches HR@1 = 0.281 on Atlanta, 0.334 on Chicago, 0.368 on Seattle, and 0.328 on Washington. The paper states that this improves over the strongest MFM baseline, TrajMoE, by 14.7%–16.8%, and over the best LLM baseline, QT-Mob, by 9%–17%. In zero-shot and few-shot transfer, MoveFM-R records 0.164/0.264 on Atlanta and 0.280/0.309 on Chicago, compared with QT-Mob at 0.132/0.203 and 0.242/0.255, and TrajMoE at 0.121/0.151 and 0.085/0.098. The paper summarizes these gains as 12%–30% over the best baseline in zero-shot and 7%–30% in few-shot settings.
For unconditional generation, averaged over the four cities, MoveFM-R reports the best values on both temporal and spatial distributions. On time, it achieves BLEU 0.628, TVD 0.064, and JSD 0.006. On location, it achieves BLEU 0.136, TVD 0.250, and JSD 0.062. These are better than the listed baselines, including LLMob with time metrics 0.605/0.085/0.007 and location metrics 0.095/0.323/0.095, and Marionette with time metrics 0.582/0.082/0.008 and location metrics 0.092/0.346/0.102.
For conditional generation, the paper contrasts scenario-blind generation with self-reflective reasoning. The w/o SR (baseline) setting yields time metrics BLEU 0.387, TVD 0.117, JSD 0.009 and location metrics BLEU 0.076, TVD 0.494, JSD 0.220. Under Scenario-i (late-night), the metrics are 0.532/0.109/0.010 on time and 0.128/0.339/0.124 on location. Under Scenario-ii (tour plan), they are 0.506/0.121/0.011 on time and 0.148/0.243/0.080 on location. Under Scenario-iii (weekend), they are 0.414/0.153/0.019 on time and 0.080/0.560/0.323 on location. These figures show that self-reflective conditioning can improve scenario adherence, but they also indicate variation across scenarios rather than uniform gains on every metric.
6. Component contributions, qualitative behavior, and implementation
The ablation study isolates three components: CB = codebook, RU = representation understanding (curriculum), and FM = base MFM. In prediction, every removal reduces HR@1 relative to the full model (Meng et al., 26 Sep 2025). For example, in Atlanta the full model obtains 0.281, while w/o CB falls to 0.243, w/o RU to 0.270, and w/o FM to 0.259. Similar drops appear in Chicago (0.334 \rightarrow 0.310/0.328/0.318), Seattle (0.368 \rightarrow 0.326/0.350/0.337), and Washington (0.328 \rightarrow 0.306/0.314/0.304).
In unconditional generation, the same pattern holds. The full model’s time metrics are 0.628/0.064/0.006, while w/o CB gives 0.598/0.090/0.007, w/o RU gives 0.613/0.072/0.006, and w/o FM gives 0.594/0.087/0.007. On location, the full model’s 0.136/0.250/0.062 degrades to 0.112/0.273/0.072, 0.108/0.265/0.068, and 0.108/0.278/0.074 respectively. The paper’s conclusion from these results is that every component contributes meaningfully.
The qualitative examples illustrate natural-language controllability. In the Weekend Planner example, the instruction is: "Plan a Saturday itinerary based on my Thursday–Friday history, emphasizing coffee shops and parks." The generated trajectory begins with entries such as At 08:30, loc [124], At 10:00, loc [217], and At 12:00, loc [341]. The accompanying commentary states that the trajectory alternates between high-density café cells and nearby green-space cells. In the Late-Night Commuter example, the instruction is: "Generate a next-day schedule where 80% of visits occur after 10 PM." The generated trajectory includes At 22:15, loc [412], At 23:00, loc [389], and At 23:45, loc [412]; the commentary states that 4 out of 5 points lie after 22:00, with movement between home and entertainment POIs.
The implementation uses four NVIDIA A800 (40 GB) GPUs, Qwen2.5-7B as the backbone, and TrajMoE as the frozen MFM. Fine-tuning is performed with LoRA on AdamW using cosine annealing, peak LR 1e-4, warmup 2e-5, batch 96, for up to 5 epochs. The codebook RQ-VAE uses an MLP encoder with layers [2048,1024,512,256,128,64], four codebooks of 512×64, and is trained with LR 1e-3 and batch 1024. For self-reflection, the study additionally uses two A100 80 GB GPUs with Qwen3-4B to stabilize the GRPO stage. Code, data splits, and prompt templates are provided at https://anonymous.4open.science/r/MoveFM-R-CDE7/.
Taken together, the reported results define MoveFM-R as a language-conditioned mobility framework in which semantic tokenization, curriculum-based alignment, and self-reflective editing are treated as distinct but interdependent mechanisms. This suggests that, within the paper’s formulation, language-driven reasoning is most effective when it is anchored to an MFM that already models spatio-temporal regularities, rather than used as a standalone replacement for mobility modeling.