---
title: 'MoveFM-R: Advancing Mobility via Language'
url: https://www.emergentmind.com/topics/movefm-r
type: topic
---

# MoveFM-R: Advancing Mobility via Language

MoveFM-R is a framework for Mobility Foundation Models (MFMs) that integrates language-driven semantic reasoning with spatio-temporal mobility modeling. It is introduced in the paper "MoveFM-R: Advancing Mobility Foundation Models via Language-driven Semantic Reasoning" [2509.22403]. The framework is motivated by a specific asymmetry: MFMs learn the statistical and geometric regularities of GPS trajectories but lack high-level semantic grounding, whereas Large Language Models (LLMs) provide strong semantic reasoning but do not possess an innate understanding of continuous geography or spatio-temporal statistics. MoveFM-R addresses this asymmetry through three linked components: a semantically enhanced location encoding, a progressive curriculum for aligning MFM latent sequences with LLM reasoning, and an interactive self-reflection mechanism for conditional trajectory generation.

## 1. Problem setting and architectural objective

MoveFM-R is designed around two stated bottlenecks in mobility modeling. The first is a **vocabulary mismatch** between continuous geographic coordinates and discrete language tokens. The second is a **representation gap** between MFM latent vectors and the semantic world of LLMs. The framework therefore treats the interface between mobility data and language not as a direct coordinate-to-token conversion problem, but as an alignment problem across heterogeneous representational spaces [2509.22403].

The model rationale is explicit. Rather than asking an LLM to infer mobility structure directly from raw coordinates, MoveFM-R first constructs semantically meaningful discrete location tokens, then teaches the LLM how to consume mobility representations through a curriculum. This organization also clarifies a common misconception in LLM-based mobility modeling: semantic capacity alone is not sufficient for physically plausible trajectory synthesis. The paper states that LLMs lack the built-in understanding of spatio-temporal statistics required for generating physically plausible mobility trajectories, while MFMs face a ceiling due to limitations in data scale and semantic understanding.

The three core innovations are presented as complementary. The semantically enhanced location encoding bridges the geography-language gap. The progressive curriculum alignment teaches the LLM "how to read" MFM latent sequences. The interactive self-reflection loop supports conditional generation under scenario constraints by editing a baseline trajectory with minimal changes. A plausible implication is that MoveFM-R is not merely an MFM augmented with prompts, but a multimodule alignment scheme in which symbolic location tokens, latent trajectory embeddings, and natural-language instructions are jointly coordinated.

## 2. Semantically enhanced location encoding

The location encoding pipeline begins with a large multi-city dataset of semantic location profiles. Each profile concatenates address strings and counts of **34 POI categories**, and a pre-trained text encoder maps the profile to an embedding $E \in \mathbb{R}^d$. A Residual-Quantized VAE (RQ-VAE) then discretizes this embedding into a sequence of $N$ codewords $(c_1,\dots,c_N)$ drawn from codebooks $\{\mathcal{C}^n\}_{n=1}^N$ [2509.22403].

At layer $n$, the quantization step is defined as
$$
v^n_{c_n}=\arg\min_{v\in\mathcal C^n}\|r_n-v\|_2^2,\quad r_{n+1}=r_n-v^n_{c_n},\quad r_1=E.
$$

The training loss combines residual quantization and reconstruction:
$$
\mathcal L_{\mathrm{RQ}}=\sum_{n=1}^N\left(\|\,\mathrm{sg}[r_n]-v^n_{c_n}\|_2^2+\alpha\,\|\,r_n-\mathrm{sg}[v^n_{c_n}]\|_2^2\right),
$$
$$
\mathcal L_{\mathrm{rec}}=\bigl\|E-\mathrm{MLP}(\hat E)\bigr\|_2^2,\quad \hat E=\sum_{n=1}^N v^n_{c_n},
$$
with total objective
$$
\mathcal L_{\mathrm{codebook}}=\mathcal L_{\mathrm{rec}}+\mathcal L_{\mathrm{RQ}}.
$$

This design matters because the paper explicitly states that the vocabulary mismatch is solved by quantizing a **high-dimensional semantic space rather than raw coordinates**. That choice distinguishes semantic tokenization from naive discretization of latitude-longitude pairs.

After discretization, the codebook is aligned to the LLM in two stages. In **static embedding alignment**, each new token embedding $e_t^{(0)}$ is initialized as the average of its subword pieces and refined with
$$
\mathcal L_{\mathrm{align}}=\mathcal L_{\mathrm{main}}+\lambda_{\mathrm{prior}}\mathcal L_{\mathrm{prior}}+\lambda_{\mathrm{coh}}\mathcal L_{\mathrm{coh}},
$$
where
$$
\mathcal L_{\mathrm{main}}=\mathbb E\bigl[\max(0,1-\cos(\hat y,y))\bigr],\quad \hat y=\mathrm{Linear}\left(\tfrac1k\sum_{i=1}^k e_{t_i}\right),
$$
$$
\mathcal L_{\mathrm{prior}}=\tfrac1M\sum_{t\in\mathcal N}\|e_t-e_t^{(0)}\|_2^2,\quad
\mathcal L_{\mathrm{coh}}=\tfrac1{|\mathcal E|}\sum_{(t,u)}\mathrm{PMI}(t,u)\,\|e_t-e_u\|_2^2.
$$

In **bidirectional instruction-tuning**, the LLM is fine-tuned so that it can map **Location-ID $\rightarrow$ Description** and **Description $\rightarrow$ Location-ID** using a standard sequence-to-sequence negative log-likelihood objective. This second stage gives the new location tokens a role inside the language model’s generative and interpretive routines, not merely inside its embedding table.

## 3. Curriculum alignment between mobility latents and language reasoning

MoveFM-R does not assume that an LLM can directly interpret an MFM latent sequence once a token vocabulary exists. Instead, it introduces a two-phase curriculum that is interleaved during training [2509.22403]. The mobility representation is formed as
$$
H_{\mathrm{seq}}=\mathrm{MLP}\bigl(g_\phi(X_{\mathrm{seq}})\bigr),
$$
where $g_\phi$ is the frozen MFM encoder.

In **Phase 1: Low-level description**, the LLM is prompted to translate a latent sequence into factual text such as "At time $t$, visited location $l$." In **Phase 2: High-level summarization**, the LLM is prompted to infer more abstract patterns, including "Most frequent locations" and "Probability of visits by time window," before performing prediction or generation. The three subtasks—description, summarization, and next-token prediction or generation—are trained jointly via
$$
\mathcal L_{\mathrm{NLL}}=-\frac1N\sum_{t=1}^N \log P_\theta\bigl(y_t \mid y_{<t}, X_{\mathrm{Ins}}, X_{\mathrm{seq}}\bigr).
$$

The paper describes this as imposing an **"understand $\rightarrow$ predict/generate" inductive bias**. The key point is that high-level mobility reasoning is not treated as an emergent side effect of sequence modeling; it is directly trained as an intermediate behavior. This addresses the stated representation gap through two mechanisms: a lightweight MLP that projects the MFM trajectory embedding $g_\phi(X_{\mathrm{seq}})\in\mathbb R^{d'}$ into the LLM embedding space, and LLM fine-tuning that conditions on these projected inputs in context.

This suggests a particular interpretation of MoveFM-R’s alignment strategy. The curriculum is not only a training schedule but also a representation regularizer: factual description constrains the model to remain grounded in trajectory content, while summarization encourages compression into semantically meaningful mobility abstractions.

## 4. Interactive self-reflection for conditional generation

For conditional trajectory generation, MoveFM-R does not generate a scenario-constrained trajectory end-to-end. The procedure is explicitly iterative [2509.22403]. First, the model generates an unconditional baseline trajectory $\tau^{(0)}$ from the user’s history. It then enters a reflection loop in which, at each iteration $k$, it compares $\tau^{(k-1)}$ against the scenario constraints and, if a violation is detected, proposes a single edit from the action set **\{add point, delete point, modify point\}** to produce $\tau^{(k)}$. The process terminates when all constraints are met.

The reward function prioritizes consistency with a true trajectory $\tau^*$:
$$
R_{\mathrm{dist}}(\tau)=\sum_{i=1}^K \mathbf 1\bigl[\phi_i(\tau)=\phi_i(\tau^*)\bigr],\quad
R_{\mathrm{len}}(\tau)=-\frac{|\;|\tau|-|\tau^*|\;|}{|\tau^*|},
$$
$$
R(\tau)=R_{\mathrm{dist}}(\tau)+R_{\mathrm{len}}(\tau).
$$

The policy is trained with **Group Relative Policy Optimization (GRPO)** and is initialized by supervised fine-tuning on a small synthetic corpus to stabilize the format. The conditional-generation mechanism is therefore organized around edit-based repair rather than direct constrained decoding.

A recurring misunderstanding in scenario-based mobility generation is that conditioning can be handled simply by prepending an instruction. MoveFM-R formalizes a different view: instructions are checked against a generated trajectory, and violations are corrected through targeted edits. A plausible implication is that the self-reflection stage functions as a structured feasibility operator over natural-language constraints, rather than as a generic stylistic prompt.

## 5. Experimental protocol and reported performance

The empirical study uses four real-world U.S. cities: **Atlanta, Chicago, Seattle, and Washington D.C.** Each city is discretized into **500 m grid cells** and **half-hour time bins**. Training instances use **3-day sliding windows**, with sequences under **5 points** or over **145 points** discarded, keeping the most recent points. Semantic POI features are drawn from **OpenStreetMap** [2509.22403].

The reported metrics are divided by task. For prediction, the metric is **Hit Rate@1 (HR@1)**. For unconditional generation, the metrics are **BLEU**, **Total Variation Distance (TVD)**, and **Jensen–Shannon Divergence (JSD)**, each computed separately over time and location distributions. Conditional generation uses the same BLEU/TVD/JSD suite under three scenarios: **late-night commuters**, **temporary travel plans**, and **weekend users**. The baselines include MFM-based prediction models (**DeepMove, GETNext, TrajBert, TrajFM, Unitraj, TrajMoE**), LLM-based prediction models (**Mobility-LLM, QT-Mob**), and generation models comprising diffusion methods (**DiffTraj, Marionette**) and LLM-chain methods (**COPB, LLMob**). The zero/few-shot protocol trains on three cities and tests on the fourth in zero-shot mode, then fine-tunes on **500 examples** for few-shot evaluation.

For next-location prediction, MoveFM-R reaches **HR@1 = 0.281** on Atlanta, **0.334** on Chicago, **0.368** on Seattle, and **0.328** on Washington. The paper states that this improves over the strongest MFM baseline, **TrajMoE**, by **14.7%–16.8%**, and over the best LLM baseline, **QT-Mob**, by **9%–17%**. In zero-shot and few-shot transfer, MoveFM-R records **0.164/0.264** on Atlanta and **0.280/0.309** on Chicago, compared with **QT-Mob** at **0.132/0.203** and **0.242/0.255**, and **TrajMoE** at **0.121/0.151** and **0.085/0.098**. The paper summarizes these gains as **12%–30%** over the best baseline in zero-shot and **7%–30%** in few-shot settings.

For unconditional generation, averaged over the four cities, MoveFM-R reports the best values on both temporal and spatial distributions. On **time**, it achieves **BLEU 0.628**, **TVD 0.064**, and **JSD 0.006**. On **location**, it achieves **BLEU 0.136**, **TVD 0.250**, and **JSD 0.062**. These are better than the listed baselines, including **LLMob** with time metrics **0.605/0.085/0.007** and location metrics **0.095/0.323/0.095**, and **Marionette** with time metrics **0.582/0.082/0.008** and location metrics **0.092/0.346/0.102**.

For conditional generation, the paper contrasts scenario-blind generation with self-reflective reasoning. The **w/o SR (baseline)** setting yields time metrics **BLEU 0.387, TVD 0.117, JSD 0.009** and location metrics **BLEU 0.076, TVD 0.494, JSD 0.220**. Under **Scenario-i (late-night)**, the metrics are **0.532/0.109/0.010** on time and **0.128/0.339/0.124** on location. Under **Scenario-ii (tour plan)**, they are **0.506/0.121/0.011** on time and **0.148/0.243/0.080** on location. Under **Scenario-iii (weekend)**, they are **0.414/0.153/0.019** on time and **0.080/0.560/0.323** on location. These figures show that self-reflective conditioning can improve scenario adherence, but they also indicate variation across scenarios rather than uniform gains on every metric.

## 6. Component contributions, qualitative behavior, and implementation

The ablation study isolates three components: **CB = codebook**, **RU = representation understanding (curriculum)**, and **FM = base MFM**. In prediction, every removal reduces HR@1 relative to the full model [2509.22403]. For example, in Atlanta the full model obtains **0.281**, while **w/o CB** falls to **0.243**, **w/o RU** to **0.270**, and **w/o FM** to **0.259**. Similar drops appear in Chicago (**0.334 \rightarrow 0.310/0.328/0.318**), Seattle (**0.368 \rightarrow 0.326/0.350/0.337**), and Washington (**0.328 \rightarrow 0.306/0.314/0.304**).

In unconditional generation, the same pattern holds. The full model’s time metrics are **0.628/0.064/0.006**, while **w/o CB** gives **0.598/0.090/0.007**, **w/o RU** gives **0.613/0.072/0.006**, and **w/o FM** gives **0.594/0.087/0.007**. On location, the full model’s **0.136/0.250/0.062** degrades to **0.112/0.273/0.072**, **0.108/0.265/0.068**, and **0.108/0.278/0.074** respectively. The paper’s conclusion from these results is that every component contributes meaningfully.

The qualitative examples illustrate natural-language controllability. In the **Weekend Planner** example, the instruction is: *"Plan a Saturday itinerary based on my Thursday–Friday history, emphasizing coffee shops and parks."* The generated trajectory begins with entries such as **At 08:30, loc [124]**, **At 10:00, loc [217]**, and **At 12:00, loc [341]**. The accompanying commentary states that the trajectory alternates between high-density café cells and nearby green-space cells. In the **Late-Night Commuter** example, the instruction is: *"Generate a next-day schedule where 80% of visits occur after 10 PM."* The generated trajectory includes **At 22:15, loc [412]**, **At 23:00, loc [389]**, and **At 23:45, loc [412]**; the commentary states that **4 out of 5 points lie after 22:00**, with movement between home and entertainment POIs.

The implementation uses **four NVIDIA A800 (40 GB) GPUs**, **Qwen2.5-7B** as the backbone, and **TrajMoE** as the frozen MFM. Fine-tuning is performed with **LoRA** on **AdamW** using **cosine annealing**, **peak LR 1e-4**, **warmup 2e-5**, **batch 96**, for **up to 5 epochs**. The codebook RQ-VAE uses an **MLP encoder with layers [2048,1024,512,256,128,64]**, **four codebooks of 512×64**, and is trained with **LR 1e-3** and **batch 1024**. For self-reflection, the study additionally uses **two A100 80 GB GPUs** with **Qwen3-4B** to stabilize the GRPO stage. Code, data splits, and prompt templates are provided at **https://anonymous.4open.science/r/MoveFM-R-CDE7/**.

Taken together, the reported results define MoveFM-R as a language-conditioned mobility framework in which semantic tokenization, curriculum-based alignment, and self-reflective editing are treated as distinct but interdependent mechanisms. This suggests that, within the paper’s formulation, language-driven reasoning is most effective when it is anchored to an MFM that already models spatio-temporal regularities, rather than used as a standalone replacement for mobility modeling.

Source: https://www.emergentmind.com/topics/movefm-r