JARVISBENCH

Benchmarking where human attention matters.

Long-horizon agents can work continuously. Human attention is intermittent and scarce. JarvisBench measures whether an always-on intermediary can coordinate both sides without changing the agents underneath.

45
agentic task instances
20
single-agent tasks
10
multi-agent projects
19
distinct domains
2K+
public candidates reviewed

THE COORDINATION PROBLEM

Execution is continuous. Attention is not.

JarvisBench separates task execution from attention coordination. Agents do the work; the user owns judgment; Jarvis decides when information should move between them.

USER Owns judgment

Intent, preferences, authority, private context, and acceptance.

JARVIS Coordinates attention

Answers the user, observes bounded events, and routes scoped guidance.

AGENTS Execute the work

Plan, use tools, produce artifacts, and continue in their native runtime.

USER → JARVIS Reach ongoing work at any moment.

Ask about progress or provide guidance while the worker keeps running.

AGENT → JARVIS → USER Bring judgment in before it is too late.

Surface one consequential decision when the unfolding work makes it concrete.

THE TASK SUITE

The need for attention emerges during execution.

Every episode contains enough public information for substantial work to begin. The benchmark does not manufacture interaction by leaving an obvious hole in the initial request. A user-owned decision becomes relevant only after the agents expose a real tradeoff, anomaly, authorization boundary, or acceptance judgment.

20

SINGLE-AGENT

One trajectory, one emerging decision.

Multi-step tasks across 15 domains and seven forms of attention need. Jarvis may pause at an action boundary, cancel the pending action, inject scoped soft guidance, and continue without discarding completed work.

25

WORKSTREAMS · 10 PROJECTS

Parallel evidence, one shared judgment.

Five projects contain two workstreams and five contain three. The streams contribute to one coupled outcome, exposing a project-level decision that no worker can resolve from its local view alone.

  1. 01Work beginsThe prompt is actionable.
  2. 02Evidence unfoldsAgents make objective progress.
  3. 03A decision maturesUser judgment changes the outcome.
  4. 04Guidance returnsExecution resumes with minimal disruption.

TWO DIRECTIONS, TWO TRACKS

Outcome improvement and user access stay separate.

JarvisBench does not collapse the two directions into one score. One track asks whether attention improves the work; the other asks whether Jarvis is useful whenever the user reaches out.

TRACK 01 · AGENT → USER

Agent-Collaboration

Was human attention used effectively?

The same worker, OpenClaw harness, prompt, tools, and environment run with and without Jarvis. The worker remains frozen; only the external attention layer changes.

Outcome
Weighted task score on a 0–100 scale.
Requests
User-attention turns requested per task.
Efficiency
Fraction of the full score gained per requested turn.
TRACK 02 · USER → AGENT

User-Interaction

Was Jarvis useful when the user reached out?

Recorded trajectories are replayed causally at early and late checkpoints. Jarvis answers a fixed progress question and a conversation-grounded follow-up using only the state visible at that moment.

General
Direct progress answer grounded in current state.
Follow-up
Contextual answer to the user's next question.
Latency
End of user speech to first audible output.

RESULTS

The attention layer transfers across workers.

With GPT-5.6-Sol as Jarvis, every completed worker configuration improves without changing the underlying agent loop. The size and attention cost of the gain still depend strongly on both the worker and the Jarvis brain.

+24.7
best single-agent gain
+28.2
best multi-agent gain
96.3
top user-interaction score
FIXED JARVIS BRAIN · GPT-5.6-SOL

Worker-agent compatibility

Task Outcome Score · higher is better

Worker agent Single-Agent Multi-Agent
BaseJarvisGain BaseJarvisGain
Claude Opus 5.058.177.7+19.655.278.6+23.4
Claude Opus 4.859.083.1+24.152.680.8+28.2
GPT-5.6-Sol57.262.1+4.951.6
GPT-5.555.167.1+12.052.965.4+12.5
DeepSeek V4-Pro51.676.3+24.753.369.4+16.1
GLM 5.251.075.6+24.652.875.3+22.5
USER-INTERACTION TRACK

Jarvis brain comparison

Overall score · latency to first audio

Jarvis brainOverallLatency
GPT-5.6-Sol96.31.8s
Claude Opus 4.890.82.1s
DeepSeek V4-Pro85.82.0s
Qwen235B82.51.6s
GPT-OSS-120B84.61.3s