Class 8 Truck VLA

Data-Efficient Adaptation of a Driving VLA to Class 8 Trucks

1 University of Southern California2 Stack AV

Off-the-shelf BaseAdapted

Adapted = Targeted SFT + FVS.
Accident I · Kcand = 1 · 6.4 s prediction horizon · Open-loop evaluation

Abstract

CollapseRead the abstract

Class 8 trucks differ from passenger cars in geometry, dynamics, and maneuvering requirements. We propose an adapt-then-steer strategy that uses limited targeted demonstrations to adapt vision-language-action (VLA) models for generating candidate trajectories in long-tail truck-driving scenarios. In the adapt stage, we use NVIDIA’s Alpamayo 1.5 as the base model, fine-tuning only its action-generation stack on a few hundred real-world construction and accident-related highway scenarios. In the steer stage, we introduce Flow Velocity Steering (FVS) to further refine the model’s predictions while holding the adapted VLA fixed. FVS is a compact, flow-time-conditioned residual module that adds learned corrections to the action-space flow velocity used to update the action sequence at each generation step. In open-loop evaluation on a scenario-disjoint held-out set, targeted fine-tuning more than halves single-candidate average displacement error (ADE) and final displacement error (FDE) over the entire 6.4 s horizon compared to the base model. Using the same targeted demonstrations, FVS further reduces the fine-tuned model’s full-horizon ADE and FDE by 13.9% and 16.5%, respectively. At matched data budgets, targeted supervision yields 19–26% lower full-horizon ADE than general truck-driving supervision, while the targeted model remains competitive with one fine-tuned on approximately 65× as many general truck-driving scenarios. These results support adapt-then-steer for data-efficient vehicle-domain transfer to Class 8 trucks.

01 / Motivation

Adapting driving VLAs to Class 8 trucks

We study how to adapt a driving VLA to Class 8 trucks, which differ from passenger cars in sensor viewpoint, articulated geometry, vehicle dynamics, and maneuvering requirements.

Why adapt a driving VLA?

We study whether limited targeted demonstrations can adapt a pretrained driving VLA to better match expert truck trajectories in construction and accident-related highway scenarios, without retraining its vision-language backbone.

The figure illustrates one motivating geometric example: under the same reference path, a passenger car clears a work-zone closure, while trailer off-tracking carries an articulated truck’s swept envelope into it.

Figure 1 panels a and b: a passenger car clears a work-zone closure, while trailer off-tracking carries the articulated truck’s swept envelope into it under the same reference path.
View full paper figure ↗
Figure 1(a–b). Passenger-car and articulated-truck swept envelopes under the same reference path. Illustrative schematics; not to scale. View full paper figure.

02 / Method

Adapt, then steer.

We first adapt the model using truck demonstrations. We then train FVS to refine its predictions while keeping the adapted model frozen.

Two-stage adaptation

We first adapt the action-generation stack, then refine it with FVS.

Stage I: the frozen vision-language backbone conditions a trainable flow-matching Action Expert; Euler action generation feeds a fixed trajectory decoder.

Step 1 of 2

Learning from truck demonstrations

In Stage 1, we fine-tune the model’s action-generation stack on 229 targeted truck scenarios while keeping the vision-language backbone frozen.

Stage II: both the backbone and adapted Action Expert are frozen; the trainable FVS module adds a correction at every flow integration step.

Step 2 of 2

Refining predictions with FVS

FVS is a compact, flow-time-conditioned residual module that corrects action-space flow velocity at each generation step. We train it on the same demonstrations while keeping the adapted model frozen.

Swipe across the diagram, or tap it to enlarge.

Stage I / Adapt

Targeted Truck-SFT

We fine-tune the action-generation stack on 229 targeted truck scenarios. The vision-language backbone stays frozen.

Stage II / Steer

Flow Velocity Steering

FVS is a compact, flow-time-conditioned residual module that corrects action-space flow velocity at each generation step. We train FVS on the same demonstrations while keeping the adapted model frozen. FVS adds 7.31M trainable parameters, less than 0.1% of the base model.

Steered Euler update in normalized action space

xk+1 = xk + Δτk (vkSFT + α Δvθ,k)

vkSFT is the frozen Action Expert’s nominal flow velocity, while Δvθ,k is the learned FVS correction.

View the complete pipeline · Figure 2
Complete adapt-then-steer architecture. View full figure ↗
Figure 2. Stage I adapts the action-generation stack. Stage II freezes the adapted policy and trains FVS throughout the Euler rollout, before fixed trajectory decoding.

03 / Test scenarios

Performance on unseen test scenarios

We compare Base, Targeted SFT, and FVS on four unseen test scenarios that were not used for Targeted SFT or FVS training. Each comparison uses the same observations, 6.4 s prediction horizon, and inference-noise seed.

1 Hz inference10 fps playback27 s per clip6.4 s horizon
GT · logged expertBaseTargeted SFTFVS

04 / Quantitative results

Trajectory prediction results

Full-horizon trajectory error

Kcand = 1 · 6.4 s · Lower is better
0123Base2.6358Targeted Truck-SFT1.1650FVS (Transformer)1.0035ADE@6.4s (metres) · lower is better
ADE@6.4s, Kcand = 1. Means across inference-noise seeds; standard deviations are reported in the table below.
02468Base7.1830Targeted Truck-SFT3.1590FVS (Transformer)2.6383FDE@6.4s (metres) · lower is better
FDE@6.4s, Kcand = 1. Means across inference-noise seeds; standard deviations are reported in the table below.
Mean ± sample standard deviation across inference-noise seeds; metres.
MethodADE@6.4s ↓FDE@6.4s ↓
Base2.6358 ± 0.14937.1830 ± 0.2868
Targeted Truck-SFT1.1650 ± 0.01843.1590 ± 0.0213
FVS (Transformer)1.0035 ± 0.00352.6383 ± 0.0156

SFT versus Base. We find that Targeted SFT reduces scenario-mean ADE@6.4s relative to Base in all 26 held-out test scenarios.

FVS versus SFT. We find that Transformer FVS further reduces scenario-mean ADE@6.4s relative to Targeted SFT in all 26 held-out test scenarios.

Complete trajectory comparison · All methods and candidate settings
Full trajectory accuracy (metres; lower is better)
MethodKcand = 1 · Single candidateKcand = 16 · Oracle best-of-16
ADE@1sADE@3sADE@6.4sFDE@6.4sminADE@1sminADE@3sminADE@6.4sminFDE@6.4s
CTRV0.16430.70032.04305.8412N/AN/AN/AN/A
Base0.1530 ± 0.00180.6836 ± 0.04402.6358 ± 0.14937.1830 ± 0.28680.0621 ± 0.00010.2556 ± 0.00820.9402 ± 0.00382.3849 ± 0.0118
Targeted Truck-SFT0.0600 ± 0.00080.3120 ± 0.00811.1650 ± 0.01843.1590 ± 0.02130.0413 ± 0.00040.1855 ± 0.00290.6968 ± 0.01151.8118 ± 0.0398
General Truck-SFT0.0584 ± 0.00170.3146 ± 0.01941.2868 ± 0.00493.7404 ± 0.08240.0361 ± 0.00010.1659 ± 0.00510.6634 ± 0.00441.7652 ± 0.0240
General → Targeted SFT0.0580 ± 0.00170.3121 ± 0.01681.2539 ± 0.04223.6012 ± 0.07030.0371 ± 0.00080.1716 ± 0.00450.6697 ± 0.00091.7587 ± 0.0313
Final-Action Residual (MLP)0.0595 ± 0.00060.3060 ± 0.00431.1362 ± 0.00923.1071 ± 0.01860.0402 ± 0.00030.1813 ± 0.00120.6918 ± 0.00731.8001 ± 0.0232
Final-Action Residual (Transformer)0.0589 ± 0.00060.3010 ± 0.00321.1276 ± 0.00813.0002 ± 0.01750.0391 ± 0.00030.1801 ± 0.00160.6876 ± 0.00511.7992 ± 0.0210
FVS (MLP)0.0481 ± 0.00030.2659 ± 0.00041.0065 ± 0.00352.6527 ± 0.01710.0363 ± 0.00000.1658 ± 0.00010.6613 ± 0.00271.7419 ± 0.0115
FVS (Transformer)0.0468 ± 0.00030.2656 ± 0.00091.0035 ± 0.00352.6383 ± 0.01560.0354 ± 0.00000.1642 ± 0.00060.6573 ± 0.00121.7003 ± 0.0102

ADE averages Euclidean position error through each horizon; FDE measures error at its endpoint. For Kcand = 16, each metric independently selects its lowest-error candidate before averaging across windows. This oracle best-of-set measure does not assume an online selector with access to GT. CTRV is deterministic; learned-model uncertainty reflects inference-noise variability for fixed checkpoints.

Supporting geometry metrics · Lateral, longitudinal, and heading error
Kcand = 1 trajectory geometry at 6.4 seconds (lower is better)
MethodLateral ADE (m)Longitudinal ADE (m)Heading error (°)
CTRV0.79581.69341.6162
Base1.0204 ± 0.06382.1706 ± 0.23063.2926 ± 0.8549
Targeted Truck-SFT0.4749 ± 0.01300.9323 ± 0.01970.8851 ± 0.0232
General Truck-SFT0.4832 ± 0.01521.0929 ± 0.04021.1377 ± 0.0459
General → Targeted SFT0.4932 ± 0.01401.0385 ± 0.02981.1192 ± 0.0553
Final-Action Residual (MLP)0.4634 ± 0.00700.9115 ± 0.01150.8590 ± 0.0172
Final-Action Residual (Transformer)0.4592 ± 0.00680.9009 ± 0.01170.8512 ± 0.0171
FVS (MLP)0.4059 ± 0.00530.7738 ± 0.01270.7644 ± 0.0120
FVS (Transformer)0.3870 ± 0.00510.7650 ± 0.01100.7169 ± 0.0105

Lateral and longitudinal ADE measure mean absolute VCP-position error along the corresponding axes. Final heading error is the absolute angular difference at 6.4 s. Values are mean ± sample standard deviation across inference-noise seeds, except deterministic CTRV.

Data Efficiency of Targeted Supervision

We compare targeted and general truck-driving supervision at matched training budgets of 2, 10, 100, and 200 scenarios. Targeted supervision yields 19–26% lower full-horizon ADE.

With 229 targeted scenarios, we obtain competitive trajectory accuracy compared with a model fine-tuned on 14,891 general truck-driving scenarios, approximately 65× as much data.

The ordering depends on the metric: general supervision performs better on some oracle best-of-16 measures. Sweep curves summarize three independently trained subset runs; full-data checkpoints are separate references.

Data efficiency of targeted long-tail supervision compared with general truck-driving supervision, including standard-deviation bands. View full figure ↗
Figure 3. ADE@6.4s versus training-set size. Curves show means across three subset runs; bands indicate ±1 standard deviation. Full-data checkpoints and zero-shot Base are separate references.
View mean-only graph · Without standard-deviation bands
Mean ADE versus training-set size for Targeted and General Truck-SFT, without standard-deviation bands.View full figure ↗
Mean-only view. Mean ADE@6.4s across three subset runs. Uncertainty bands are omitted in this optional view; the original Figure 3 with ±1 standard deviation is shown above. Full-data checkpoints and zero-shot Base are separate references.

05 / Reference

Cite this work.

Citation to be updated.

BibTeXTo be updated
% Citation details to be updated.
@misc{citation_to_be_updated,
  title = {To be updated},
  author = {To be updated}
}

Research figure

Open full-resolution image ↗