LAD-VF: Fine-Tuning-Free LLM Alignment
- The paper introduces LAD-VF, a framework that uses formal verification feedback to iteratively optimize LLM prompts without fine-tuning.
- The method converts verification failures into textual feedback, guiding prompt refinements for improved safety and regulatory compliance.
- Experimental results show LAD-VF boosts specification compliance from 60% to over 90% while significantly reducing data and compute costs.
Searching arXiv for the target paper and closely related work on LLM-AutoDiff / formal-feedback planning. LAD-VF is a fine-tuning-free framework for aligning LLM planners with formal safety and regulatory constraints by iteratively optimizing prompts rather than model parameters. Introduced in the paper titled "AD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback," the method combines formal verification, a formal-verification-informed text loss, and LLM-AutoDiff so that verification failures are translated into textual feedback for prompt refinement (Yang et al., 22 Sep 2025). In the reported formulation, LAD-VF targets LLM-driven planning for robotics, autonomous driving, and related physical-world settings where hallucination or weak alignment can produce plans that violate temporal logic specifications. Its stated benefits are scalable adaptation without fine-tuning, compatibility with modular LLM architectures, and interpretable refinement via auditable prompts (Yang et al., 22 Sep 2025).
1. Concept and problem setting
LAD-VF is presented for settings in which an LLM receives a task instruction and produces an executable or formalized plan, but must do so under strict safety or regulatory requirements. The motivating problem is that conventional LLM planning may violate such requirements, while traditional data-driven alignment methods such as Direct Preference Optimization require costly human labeling, and recent formal-feedback approaches still depend on resource-intensive fine-tuning (Yang et al., 22 Sep 2025).
The framework therefore shifts the alignment target from parameter adaptation to prompt adaptation. In the reported formulation, the planner is not retrained; instead, prompts are iteratively edited in response to failures detected by a formal verifier. This makes prompt engineering the locus of optimization. A plausible implication is that the method treats prompts as a reusable interface layer between the underlying model and domain-specific constraints, rather than embedding those constraints into model weights.
A central terminological point is that the paper title uses "AD-VF," whereas the abstract and technical exposition describe the proposed framework as "LAD-VF" (Yang et al., 22 Sep 2025). Within that exposition, LAD-VF denotes the verification-feedback-driven prompt optimization procedure built on LLM-AutoDiff.
2. Closed-loop architecture and execution pipeline
The workflow is a closed loop with four principal stages: task-conditioned plan generation, formal verification, loss computation, and prompt optimization (Yang et al., 22 Sep 2025). The user provides a task instruction and an initial prompt set . Conditioned on these prompts, the LLM generates a high-level formalized plan, described in the reported implementation as, for example, a plan in NuSMV format (Yang et al., 22 Sep 2025).
That generated plan is then converted to a formal automaton and checked against user-provided temporal logic safety specifications using a model checker (Yang et al., 22 Sep 2025). The verifier records the number of violated specifications, denoted , and this becomes the basis of the optimization signal. Rather than regarding verification as a terminal acceptance test, LAD-VF integrates it into the planning loop as a structured source of supervision.
The reported system then uses LLM-AutoDiff to propagate feedback through the planning pipeline in textual form. In place of numeric gradients over neural weights, the framework uses human-readable refinement suggestions derived from verification failures. These suggestions are fed to an optimizer LLM, which produces revised prompts for the next planning iteration (Yang et al., 22 Sep 2025). The result is an iterative prompt-refinement loop in which formal verification acts as the critic and prompt editing acts as the update mechanism.
This architecture is explicitly described as modular. Nodes in the pipeline can be LLM modules or functional modules such as verifiers, and the framework is said to remain applicable when modules are interchangeable or when input-output formats change, because only the textual interface requires adjustment (Yang et al., 22 Sep 2025).
3. Mathematical formulation and LLM-AutoDiff integration
The formal-verification-informed loss is defined as the specification violation rate,
where is the total number of specifications (Yang et al., 22 Sep 2025). The corresponding safety score is reported as
with higher values indicating greater compliance (Yang et al., 22 Sep 2025). This establishes a direct optimization target grounded in formal methods rather than in human preference labels.
At the system level, the reasoning pipeline is modeled as a directed computation graph , whose nodes may be LLM modules with trainable prompts or non-LLM functional modules (Yang et al., 22 Sep 2025). The prompt set is
and prompt optimization is posed as
0
Here, 1 denotes plan generation for task 2 under prompt set 3 (Yang et al., 22 Sep 2025).
The distinguishing technical mechanism is the backpropagation of textual gradients through the computation graph. For a node 4, the reported propagation rule is
5
where 6 acts as a critic or editor that transforms downstream verification failures into prompt-relevant feedback (Yang et al., 22 Sep 2025). Prompt updating is then written as
7
with 8 synthesizing an improved prompt from the prior prompt and the accumulated textual gradient context (Yang et al., 22 Sep 2025).
This formulation makes the entire planning-and-verification loop legible as a differentiable-style optimization procedure over prompts. The paper’s framing suggests that the essential novelty is not only the use of formal verification as a binary checker, but its conversion into an optimization signal that can be routed through multi-module LLM pipelines.
4. Prompting regimes and refinement behavior
Three prompting regimes are reported: single-iteration, two-iteration, and multi-iteration (Yang et al., 22 Sep 2025). In the single-iteration regime, the LLM generates the whole plan in one pass. In the two-iteration regime, it first decomposes the task into plain-language steps and then converts those steps into formal plan code. The default multi-iteration regime generates each plan step in sequence (Yang et al., 22 Sep 2025).
The reported interpretation is that the multi-iteration regime is most effective for safety compliance because it captures dependencies at finer granularity and gives prompt optimization more localized leverage over the generation process (Yang et al., 22 Sep 2025). This suggests that verification feedback can be more actionable when it attaches to intermediate planning decisions rather than only to a final full-plan output.
The refinement process is described as interpretable and auditable. Examples cited in the paper’s appendices and figures show prompt updates expressed as understandable textual edits, such as adding conditional checks after a verification failure reveals that an action violates a stop-sign rule in a particular case (Yang et al., 22 Sep 2025). The exposition also notes that optimized prompts become sequential, structured, and specific—for example, "Check Stop_Sign, then Traffic Light, then Pedestrian..."—whereas initial prompts are comparatively vague (Yang et al., 22 Sep 2025).
A common misconception would be to view LAD-VF as mere prompting supplemented by static safety instructions. The reported method is more specific: prompt text is revised iteratively as an optimization variable in response to formal counterevidence, rather than being fixed or manually tuned once.
5. Experimental findings
The experimental domains include robot navigation, manipulation, and delivery tasks, including robot navigation with a Jackal robot and table-top manipulation with robotic arms (Yang et al., 22 Sep 2025). The paper reports that LAD-VF substantially enhances specification compliance, improving success rates from 60% to over 90% (Yang et al., 22 Sep 2025).
Using safety score as the main metric, the reported quantitative results are: LAD-VF achieves 0.88 on validation and 0.86 on test; LAD-VF combined with in-context learning (ICL) achieves 0.95 on validation and 0.95 on test; and the fine-tuning baseline RLVF achieves 0.978 on both validation and test, but at much greater data and compute cost (Yang et al., 22 Sep 2025). By contrast, "Prompt+Spec" has a reported test score of 0.013, while ICL alone reaches 0.80 on test (Yang et al., 22 Sep 2025).
The paper further reports that LAD-VF reaches peak performance within 10 prompt refinement steps, and often matches or exceeds fine-tuning-based methods with half the data and a fraction of the compute (Yang et al., 22 Sep 2025). In addition, out-of-domain evaluation indicates that prompts optimized on one robot or navigation task generalize to indoor delivery tasks and table-top manipulation with no human-in-the-loop re-optimization (Yang et al., 22 Sep 2025).
These results position LAD-VF between conventional prompting and full fine-tuning. It does not surpass the reported RLVF fine-tuning score, but it narrows the gap substantially while preserving the advantages associated with prompt-only adaptation. A plausible implication is that the framework is particularly attractive when retraining is operationally expensive or when the underlying LLM stack is subject to frequent modular changes.
6. Interpretability, modularity, and relation to trustworthy planning
The paper identifies three distinguishing properties: fine-tuning-free adaptation, modular compatibility, and interpretable, auditable refinement (Yang et al., 22 Sep 2025). Fine-tuning-free adaptation means that no model parameters are updated; only prompts are changed. Modular compatibility means the framework can be applied to pipelines with multiple LLM and non-LLM components, including altered interfaces and interchangeable modules. Interpretability follows from the fact that each optimization step is encoded as a textual prompt revision rather than an opaque change in neural weights (Yang et al., 22 Sep 2025).
In the context of trustworthy robotic planning, LAD-VF is framed as a method for enforcing temporal logic constraints through a closed interaction between generation and verification. Its significance lies not only in compliance improvement, but in how compliance is achieved: formal methods produce explicit failure signals, and those signals are converted into auditable edits to the planner’s instructions (Yang et al., 22 Sep 2025). This differs from approaches in which alignment is represented only implicitly in a fine-tuned parameter state.
An objective reading of the reported evidence also clarifies what LAD-VF does not claim. It is not presented as eliminating the need for formal specifications; the method depends on a set of user-provided temporal logic safety specifications 9 and a model checker (Yang et al., 22 Sep 2025). It is also not presented as universally superior to fine-tuning on raw benchmark score, since the reported RLVF baseline attains 0.978 on validation and test (Yang et al., 22 Sep 2025). Its contribution is instead the combination of strong compliance gains, few prompt refinement steps, lower data and compute demands, modularity, and prompt-level interpretability.
Taken together, LAD-VF defines a specific paradigm for safety alignment in LLM-based planning systems: optimize prompts with formal verification feedback and LLM-AutoDiff, keep the base model fixed, and retain an auditable record of every refinement (Yang et al., 22 Sep 2025).