DORA Explorer:

Improving the Exploration Ability of LLMs Without Training

Anonymous Authors

DORA Explorer (Diversity-Oriented Ranking of Actions) is a training-free framework for improving exploration in LLM agents. LLM agents struggle with exploration, often converging to sub-optimal solutions or getting stuck in loops. DORA addresses this by generating diverse action candidates, scoring them via token log-probabilities, and selecting actions with a tunable exploration parameter. DORA achieves UCB-competitive performance on Multi-Armed Bandits and consistent gains across the Text Adventure Learning Environment Suite (TALES).

Method

DORA Explorer samples diverse action candidates, scores them by log-probability, and selects the best action via a tunable exploration parameter.

Algorithm 1 DORA Explorer

Input: History H, observation ot, used-action set Aot

Hyperparams: nC, {τd, τλ, τC}, α

Output: Selected action a*t


 1: c ← [H; o_t]
 2: d ~ π_d(c, τ_d)                           ▷ Decide greedy or explore
 3: if d = EXPLORE then
 4:   λ ← λ_exp(t)  or  π_λ(c, τ_λ)        ▷ Schedule or policy
 5:   C ← ∅
 6:   C' ~ π_C(c, τ_C, n_C)                  ▷ Sample n_C candidates
 7-11: Filter previously used actions from C' → C
12:   s = (Score(a, α))_{a∈C}
13:   p_λ = Softmax(λ · s)
14:   a*_t ~ Categorical(p_λ)
15: else
16:   a*_t ~ π_g(c, τ=0)                     ▷ Greedy action
17: end if
18: A_{o_t} ← A_{o_t} ∪ {a*_t}
19: return a*_t

Results on Multi-Armed Bandits

Arm-selection dynamics on a 5-armed bandit (the best arm is purple). On the left, DORA Explorer with a λ-scheduler progressively converges to the optimal arm, while on the right the LLM at temperature τ = 0 stays stuck on a suboptimal arm.

t = 0 / 199
slow fast
1.0x
DORA Explorer
total reward 0
chooses -
reward 0
LLM at τ=0
total reward 0
chooses -
reward 0

Results on TALES

A TextWorld cooking task from TALES. DORA Explorer systematically explores rooms, finds the cookbook, gathers ingredients, and completes the recipe (3/3), while the zero-shot agent wanders in loops across the Shed, Backyard, and Garden without ever reaching the Kitchen (0/3).

Step 0
slow fast
1.0x
DORA Explorer
Score 0 / 3
Shed
Zero-shot
Score 0 / 3
Shed

Findings

DORA explores early in MAB and converges quickly to the optimal arm.

Arms selected by DORA over a single episode. In the early timesteps the λ policy samples all arms, and as t grows λ increases and the agent shifts to exploiting the optimal arm (red).

Cumulative arm selection by DORA over a single episode

Cumulative regret across methods over a horizon of T = 500. DORA’s λexp schedule approaches classical baselines (UCB, TS), while a greedy LLM (τ = 0) accumulates substantially higher regret.

Cumulative regret across baselines and DORA
DORA achieves near-optimal performance on Multi-Armed Bandits.

Multi-Armed Bandits (MAB) provide a classical testbed for the exploration–exploitation trade-off. We study the Bernoulli instance with K=5 arms, gap Δ=0.2, and horizon T=500. For τ ≤ 1, the model behaves overly greedily, often committing early to a suboptimal arm (suffix failure up to 90%). For τ > 1, increased randomness prevents consistent exploitation. DORA explores effectively in early stages and rapidly converges to the optimal arm, achieving zero Suffix Failure Frequency and approaching the near-optimal UCB baseline.

Metric UCBTSGreedyε-Greedy τ=0τ=0.7τ=1τ=1.5τexp DORA λexp
Mean Avg Reward 0.5640.5480.5030.509 0.4140.4140.4220.5180.484 0.538
SuffFailFreq(T/2) 0.0110.0000.4650.230 0.9000.9000.8500.0000.000 0.000
Best Arm Fraction 0.8240.7430.5150.550 0.1000.1000.1310.6750.698 0.703
Cum Regret 17.63725.68648.50244.998 90.00090.00086.90034.64037.160 28.730

Table 1. Performance on the Bernoulli MAB instance using Llama-3.1 8B with horizon T=500. UCB, TS, Greedy, and ε-greedy are averaged over N=1000 runs; LLM-based methods are averaged over N=20 runs.

DORA improves overall task success on TALES.

DORA visits over 2× more unique states in TextWorld and ScienceWorld, enabling discovery of hidden objects and task-relevant signals. By enabling agents to discover new information and avoid repetitive loops, DORA leads to consistent performance gains across all TALES environments, including measurable progress in AlfWorld where feedback is sparse and unhelpful.

Model Setting TextWorld TW Express AlfWorld ScienceWorld Jericho
Qwen-2.5 7B Zero-shot 29.20±1.92 48.93±1.69 0.00±0.00 13.49±0.12 0.66±0.04
Chain of Thought 24.89±0.91 37.44±0.00 0.00±0.00 9.81±0.18 1.58±0.03
Tree of Thought 11.39±0.83 45.90±0.21 0.00±0.00 10.10±0.00 1.40±0.00
Prompt Explore 23.12±1.29 47.81±0.00 0.00±0.00 13.21±0.34 1.27±0.04
ReAct 31.43±3.16 39.50±0.62 2.78±2.78 9.92±0.62 1.87±0.25
DORA (Ours) 45.40±1.68 51.17±3.38 2.78±2.78 19.01±0.62 1.42±0.24
Llama-3.1 8B Zero-shot 40.74±0.67 48.29±1.04 0.00±0.00 15.53±0.50 2.21±0.14
Chain of Thought 37.72±1.16 24.72±2.50 0.00±0.00 7.81±0.87 2.22±0.22
Tree of Thought 15.06±0.69 15.52±0.21 0.00±0.00 8.64±1.64 2.05±0.15
Prompt Explore 43.07±0.72 52.77±2.08 0.00±0.00 13.70±0.59 1.60±0.21
ReAct 22.92±5.04 60.38±1.95 0.00±0.00 15.94±2.34 2.88±0.51
DORA (Ours) 50.15±6.88 52.89±3.51 2.78±2.78 16.85±0.58 2.17±0.21
Mistral Small 22B Zero-shot 62.25±1.08 31.70±2.08 0.00±0.00 26.47±0.57 1.35±0.03
Chain of Thought 31.43±2.22 25.16±0.28 0.00±0.00 9.63±0.35 1.43±0.13
Tree of Thought 32.54±0.83 26.40±1.71 0.00±0.00 10.88±0.85 1.27±0.03
Prompt Explore 50.60±0.69 42.46±0.79 0.00±0.00 21.48±0.48 1.70±0.05
ReAct 57.78±5.47 23.54±2.07 0.00±0.00 26.70±4.19 3.37±0.02
DORA (Ours) 74.28±2.99 43.12±1.21 2.78±2.78 32.43±1.15 2.92±0.66
Llama-3.1 70B Zero-shot 63.95±2.00 82.25±0.75 11.11±5.56 54.35±1.29 5.17±0.10
Chain of Thought 61.29±0.00 80.23±1.88 5.56±5.56 45.23±2.05 5.37±0.17
Tree of Thought 67.03±4.55 48.73±3.70 0.00±0.00 28.32±3.47 3.31±0.44
Prompt Explore 61.95±1.15 83.87±3.12 0.00±0.00 52.51±3.61 5.57±0.19
ReAct 57.03±4.42 70.60±3.37 16.67±0.00 44.79±1.47 4.72±0.37
DORA (Ours) 71.12±2.24 85.33±0.73 16.67±0.00 63.40±2.62 5.22±0.09

Table 2. Performance of DORA compared to baseline prompting strategies on TALES environments. Results report mean normalised scores across tasks ± standard error over 3 runs.

DORA enables loop recovery with 5× to 20× higher rates.

Baseline methods frequently enter loops and rarely recover (below 1%). DORA substantially reduces loops encountered and achieves 5× to 20× higher recovery rates across all TALES environments.

Method TextWorld ScienceWorld AlfWorld
LoopsRec.% LoopsRec.% LoopsRec.%
Zero-shot61230.5255190.3581100.0
Prompt Explore66350.82321100.4372110.1
CoT673243.6262160.2384610.1
ToT78381.0280120.0785700.0
DORA (Auto) 1853619.5 1234453.6 1111210.8

Table 3. Loops encountered, loops recovered, and recovery rate (%) across TALES environments.

DORA covers more states compared to baselines.

DORA visits significantly more unique states per task across all TALES environments, indicating greater coverage and information gathering.

Zero-shot Prompt Explore CoT ToT ReAct DORA
TextWorld
14.0
12.1
17.5
12.1
22.0
32.0
ScienceWorld
6.5
8.9
8.23
4.6
13.77
20.23
AlfWorld
2.67
3.67
3.25
3.67
3.33
27.92

Figure 1. Average unique states visited per task in TALES environments for Llama-3.1 8B. DORA explores substantially more of the environment than prompting methods.

DORA consistently outperforms all temperature strategies.
Model τ=0τ=0.3τ=0.7τ=1τ=1.5τ=2τ-policy DORA
Llama-3.1 8B 39.4134.0830.7414.330.000.000.00 44.74
Qwen-2.5 7B 27.6523.6526.7417.6521.8312.8310.33 44.62
Mistral Small 22B 60.1654.6239.835.835.830.000.00 73.59

Table 4. Comparison of τ-sampling strategies against DORA on TextWorld.

α=0.8 yields the best confidence–consistency trade-off.
Metric (1.0, 0.0)(0.8, 0.2)(0.6, 0.4)(0.4, 0.6)(0.2, 0.8)(0.0, 1.0)
Mean Avg Reward 0.4960.5210.5090.5060.5000.511
Suff. Fail Freq (T/2) 0.050.000.050.050.050.00
Best Arm Frac 0.5160.6180.5840.5400.5280.577
Cum. Regret 19.3415.2616.6318.4018.9016.92

Table 5. Ablation over explore-score weights (α, 1−α) on the hard MAB instance using Llama-3.1 8B. (0.8, 0.2) provides the best overall trade-off.

Auto-Explorer outperforms the λ-scheduler on complex tasks.
Modelλ-SchedAuto
Qwen-2.5 7B5.8342.95
Llama-3.1 8B21.0041.89
Mistral Small 22B9.1769.57

Table 6. λ-schedule versus Auto-Explore on TextWorld.

DORA uses more tokens but achieves substantially better outcomes.

DORA uses approximately 2× more tokens than Prompt Explore while remaining less expensive than ReAct; this additional compute translates into significantly better task performance.

Zero-shot
0.20M
Prompt Explore
0.19M
CoT
0.28M
ToT
0.30M
DORA
0.36M
ReAct
0.39M

Figure 2. Average token usage per TextWorld task for Llama-3.1 8B. DORA uses approximately 2× more tokens than Prompt Explore while remaining less expensive than ReAct.