ACL 2026 Simulator-free VLN MLLM agent diagnosis

VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents

A unified, modular, and simulator-free evaluation framework for probing multimodal large language models as zero-shot embodied navigation agents.

Xunyi Zhao*, Gengze Zhou*†, Qi Wu
Australian Institute for Machine Learning, University of Adelaide
* Equal contribution. † Project Lead.

Overview

Beyond leaderboard success rates

VLN-MME is designed to diagnose how MLLMs behave as language-guided visual navigation agents, not just whether they reach the goal.

Unified evaluation

Models, agents, and tasks are treated as interchangeable components under one evaluation interface.

Simulator-free

Pre-rendered observations and graph metadata remove the need to install Matterport3D or Habitat simulators.

Agent diagnosis

The benchmark compares memory formats, CoT prompting, reflection, and MLLM backbones under consistent tasks.

Failure analysis

Step-level trajectory inspection reveals looping, weak historical grounding, and perception-action gaps.

Framework

Model, Agent, Task, and Runner

The implementation separates model APIs, agent prompting and memory, navigation tasks, environment loading, and evaluation logging.

The Model interface handles open-weight and proprietary MLLMs. The Agent converts state into prompts, stores history, parses model output, and selects actions. The Task encapsulates dataset splits, maps, and evaluation metrics.

A central Runner coordinates loading, rollout, logging, and metric computation, allowing component-level comparisons without rewriting the evaluation loop.

Memory variants include text summary agents and topological text-map agents.
Reasoning variants include baseline prompting, CoT, reflection, and CoT + reflection.
Adding a new model, agent, or task follows the same registry pattern used by the current codebase.
VLN-MME modular framework diagram
Framework overview from the paper.
VLN-MME composition across tasks, agents, models, and datasets
VLN-MME composition over tasks, agents, models, and datasets.

Composition

One benchmark, multiple axes

VLN-MME is organized around the interaction between navigation tasks, agent architectures, and MLLM backbones. This makes it possible to diagnose whether a result comes from model capability, memory design, reasoning prompts, or task granularity.

Tasks: fine-grained, coarse-grained, object-oriented, and extensible navigation settings.
Agents: NavGPT-style memory, MapGPT-style memory, CoT, reflection, and combined variants.
Models: open-weight and proprietary MLLMs can be swapped under the same protocol.

Benchmark

Efficient coverage of navigation granularity

The paper evaluates representative navigation tasks across fine-grained instructions, coarse-grained referring expressions, and object-oriented navigation.

Distribution comparison between the curated subset and original R2R validation split
The curated subset closely follows the original R2R val-unseen distribution.

VLN-MME constructs compact but representative evaluation subsets so large MLLM agents can be compared without exhaustive simulator runs.

Fine-grained: R2R instruction following with step-by-step language guidance.
Coarse-grained: REVERIE navigation to remote, out-of-sight targets.
Object-oriented: ObjectNav-MP3D navigation from a target category.
The released code also supports RXR-EN and CVDN for broader extensions.

Observation

Simulator-free visual grounding

VLN-MME replaces online simulator rendering with pre-rendered, agent-centric observations and semantic metadata.

Four-view panoramic observation with navigable markers
Each step provides four 90-degree perspective views arranged as Left, Front, Right, and Back, with navigable markers overlaid.
4 views

Agent-centric panoramic input avoids dense image sweeps and distorted stitched panoramas.

~1.7GB

Runtime memory compared with about 10GB for simulator-based evaluation.

0.016s

Observation access compared with about 0.14s when rendering online.

~25s

Approximate per-episode time saved by avoiding simulator rendering overhead.

Results

Reasoning prompts do not guarantee better navigation

The results expose a gap between formatted reasoning and usable embodied context awareness.

Proprietary MLLMs such as GPT-5 and Gemini-2.5 Pro set the strongest zero-shot upper bound, while Qwen2.5-VL-7B is a robust open-weight baseline.

Counterintuitively, adding CoT or reflection often hurts navigation. For Qwen2.5-VL-7B on fine-grained navigation, CoT reduces success rate from 27.5% to 21.0%, indicating that explicit reasoning traces can fail to use historical context.

CoT outputs can look structured while still relying on local visual cues.
Text-map memory does not automatically solve spatial grounding.
Object-oriented navigation is generally more tractable than coarse-grained navigation.
Performance comparison under different reasoning strategies
Reasoning strategy comparison for text-summary memory agents.

Main Table

Compact web view of main table

The paper table reports eight metrics for each navigation category. On the website, the same comparison is condensed to the most readable success-efficiency pair: SR and SPL for fine-grained, coarse-grained, and object-oriented navigation.

The pattern is consistent with the paper discussion: proprietary models define the upper bound, Qwen2.5-VL-7B is the strongest open-weight baseline in several settings, and explicit CoT/reflection often reduces navigation success rather than improving it.

Show compact main table: SR / SPL across tasks
Agent / MLLM Fine SR Fine SPL Coarse SR Coarse SPL Object SR Object SPL
Text Summarization Memory Agents
NavGPT
GPT-538.5029.2330.0020.7648.0023.84
Gemini-2.5 Pro41.0032.6733.5024.3851.5027.19
InternVL3-2B13.505.467.332.5021.503.57
InternVL3-8B28.0012.6120.007.1839.007.69
LLaVA-OV-7B11.504.9414.675.1927.504.51
Qwen2.5-VL-7B27.5017.1118.679.0037.5013.18
NavGPT w/ CoT
GPT-536.0027.3128.5020.4246.5023.67
Gemini-2.5 Pro32.5023.8824.0016.9342.0019.27
InternVL3-2B8.004.475.333.2425.006.66
InternVL3-8B19.0010.9515.339.3134.5012.67
LLaVA-OV-7B12.505.4114.005.9033.507.31
Qwen2.5-VL-7B21.0011.4115.678.1033.0013.25
NavGPT w/ Reflection
GPT-533.0024.1825.0017.3443.0020.62
Gemini-2.5 Pro37.5028.7329.0021.6847.5025.14
InternVL3-2B8.005.208.504.8028.007.83
InternVL3-8B12.009.5311.005.9732.509.12
LLaVA-OV-7B10.509.449.335.8234.007.90
Qwen2.5-VL-7B24.0014.9512.007.9735.5014.67
NavGPT w/ CoT & Reflection
GPT-538.5029.8130.0022.1748.5025.92
Gemini-2.5 Pro34.0025.7425.5018.3643.5021.28
InternVL3-2B4.501.709.334.6324.507.26
InternVL3-8B22.0015.3317.3310.0732.508.14
LLaVA-OV-7B10.005.8314.006.7828.507.25
Qwen2.5-VL-7B25.5017.6811.677.8936.0013.67
Text Map Memory Agents
MapGPT
GPT-534.0025.8326.0018.2944.0020.91
Gemini-2.5 Pro39.5030.7231.5023.1649.5026.24
InternVL3-2B11.003.7112.004.4127.504.41
InternVL3-8B18.0012.4613.677.8731.5011.61
LLaVA-OV-7B8.505.5914.676.4822.504.28
Qwen2.5-VL-7B26.0017.3121.678.9636.5011.05
MapGPT w/ CoT
GPT-532.5027.1424.5019.6844.5023.21
Gemini-2.5 Pro31.0023.3723.0015.8439.5018.76
InternVL3-2B6.504.124.001.5820.004.76
InternVL3-8B12.009.6613.338.7734.0012.33
LLaVA-OV-7B8.502.838.673.3816.004.14
Qwen2.5-VL-7B17.0010.4716.3310.6132.009.95
MapGPT w/ Reflection
GPT-536.5028.1925.5020.7345.5023.65
Gemini-2.5 Pro32.0024.3124.0016.8841.0019.94
InternVL3-2B4.003.263.673.3125.004.50
InternVL3-8B16.5010.8912.676.6430.0010.36
LLaVA-OV-7B10.005.5011.006.0015.502.02
Qwen2.5-VL-7B26.5010.1215.676.0033.507.41
MapGPT w/ CoT & Reflection
GPT-534.5026.1324.5018.6742.5021.79
Gemini-2.5 Pro30.0022.2822.0015.4138.0017.86
InternVL3-2B9.004.884.671.6118.004.74
InternVL3-8B18.0010.8413.006.0633.508.01
LLaVA-OV-7B13.005.819.004.5324.505.45
Qwen2.5-VL-7B14.007.1210.676.6825.506.78

Compact rendering of `Tab/main_table.tex`: SR and SPL are retained for readability; the paper PDF contains TL, NE, nDTW, SDTW, CLS, and OSR as well.

Diagnostics

Looping reveals the perception-action gap

VLN-MME turns aggregate failure into interpretable behavioral patterns at the trajectory level.

Sankey diagram of successes and failures in VLN-MME
Failure and success flow over 200 analyzed trajectories.

Among 200 analyzed trajectories, 148 fail and 52 succeed. Of the failures, 131 are incorrect navigation rather than generation-format errors, and 106 involve looping.

Even successful episodes are often inefficient: 42 of 52 successes include looping behavior before reaching or stopping near the target. This points to weak state tracking, poor historical grounding, and incomplete translation from perception to action.

Common failure causes include region recognition, vertical movement, and object grounding.
The agent can recognize visual evidence but still choose actions that repeat prior mistakes.
Failure analysis pie chart
Failure sources: incorrect navigation dominates malformed generation.
Success behavior analysis pie chart
Most successful trajectories are still inefficient.
106 loops

Looping is the dominant incorrect-navigation pattern, reflecting unstable spatial memory and weak long-horizon correction.

Diagnostic Study

What fixes the hard negatives?

The paper revisits 25 fine-grained hard-negative trajectories where all open-weight models failed with the standard text-summary agent.

Oracle-guided navigation

A stronger Qwen3VL oracle assistant provides high-level reasoning guidance when the navigator loops, enters a wrong region, or moves in a critically wrong direction. The jump from 0% SR to 52% SR suggests the base navigator has some visual grounding, but lacks strategic planning and error-correction logic.

MethodSROSRSPL
Baseline Qwen2.5VL-7B0.000.000.00
+ Oracle Assistant52.0068.0041.28

Failure-aware in-context learning

Instead of direct oracle support, the model receives examples of common failure cases in the prompt. This improves performance modestly, but remains far below the oracle-guided setting, indicating that awareness of failure modes helps less than active reasoning assistance.

MethodSROSRSPL
Zero-shot baseline0.000.000.00
1-shot failure example12.0024.009.42
2-shot failure examples16.0028.0011.25
3-shot failure examples16.0032.0014.29

Case Studies

Step-level behavior exposes where agents fail

The paper case studies show successful but inefficient paths, vertical-movement confusion, invalid outputs, and region-understanding failures.

Successful but inefficient treadmill navigation trajectory

Successful but inefficient

The agent sees the treadmill, loops through nearby rooms, and only later stops by the target.

Successful trajectory with vertical movement looping on stairs

Recovery with vertical confusion

The agent eventually succeeds, but staircase positioning causes repeated local corrections.

Failure case with oracle success but no final stop

Oracle success, final failure

The agent reaches the correct general area but fails to stop and loops on the staircase.

Navigation failure caused by directional misunderstanding and invalid action

Instruction and generation failure

A directional grounding error is compounded by malformed output that yields an invalid action.

Failure caused by misunderstanding the target sauna region

Region-understanding failure

The agent cannot correctly interpret and enter the specified target region, the sauna.

Quick Start

Run evaluation without installing a simulator

The framework downloads processed VLN annotations, marked observations, connectivity graphs, captions, and environment metadata.

conda create --name VLNMME python=3.10
conda activate VLNMME
pip install -r requirements.txt

cd src
python main.py --config_dir configs/experiment.yaml

Citation

BibTeX

If VLN-MME is useful for your research, please cite the project.

@inproceedings{zhao-etal-2026-vln,
  title     = {{VLN}-{MME}: Diagnosing {MLLM}s as Language-guided Visual Navigation Agents},
  author    = {Zhao, Xunyi and Zhou, Gengze and Wu, Qi},
  booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
  year      = {2026},
  address   = {San Diego, California, United States},
  publisher = {Association for Computational Linguistics},
  url       = {https://aclanthology.org/2026.acl-long.1300/},
  doi       = {10.18653/v1/2026.acl-long.1300},
  pages     = {28207--28231}
}