To evaluate your multimodal agent's trajectory, the best approach depends on your development stage. For pre-release simulation, Maxim AI and Galileo AI are top choices for analyzing agent choices and reasoning paths. If you need dedicated visual grounding, AgentVista is highly specialized. For general monitoring and observability of agent interactions, LangSmith and Arize Phoenix provide robust tracing capabilities to identify where reasoning sequences or tool use may be failing.
If by “agent trajectory quality” you mean evaluating the entire execution path—tool selection, arguments, sequencing, state transitions, retries/recovery, efficiency, safety, and final task success—rather than just the final answer, there are a few groups I’d put on your shortlist.
The most directly relevant specialists
Arize AI / Phoenix — probably the clearest direct match to your wording. They explicitly support trajectory-level evaluation, where ordered tool calls across a trace are judged as a whole. Their framework can score whether a trajectory is logically progressive, uses appropriate tools, and avoids unnecessary detours.
LangChain / LangSmith — particularly relevant if your agent uses LangGraph/LangChain. LangSmith supports trajectory evaluators that inspect the sequence of agent actions rather than treating the final response as the only output.
Braintrust — strong if you're looking for an evaluation engineering workflow rather than just observability: production traces → curated datasets → trajectory scorers → experiments/regression tests → CI. Their current platform explicitly supports trajectory-level evaluation.
Future AGI — worth looking at specifically because its evaluation approach is described as trajectory-first across mixed runtimes, rather than being tightly coupled to one agent framework.
If you want research expertise, not just a SaaS platform
I'd also look at the people/groups behind AgentBench, WebArena, τ-bench, and related agent benchmarks. These are useful because they treat an agent's interaction sequence as an object of evaluation and expose an important issue: one successful trajectory doesn't necessarily mean a reliable agent. For example, τ-bench evaluates repeated task attempts and exposes consistency problems that final-answer-only evaluation can miss.
For a multimodal agent specifically
I'd distinguish three things in your release evaluation:
Layer
What I'd evaluate
Outcome
Did the agent accomplish the user's task?
Trajectory
Did it take the correct/efficient/safe sequence of actions?
Perception → action
Did it correctly interpret images/video/UI state before acting?
The third category is where a generic LLM-eval platform can become insufficient. For a multimodal agent, I'd want trajectory evaluators that can inspect the visual state available at each step, the agent's selected action, the resulting state, and whether the action was appropriate given what was actually observable.
So, if you're looking for someone to help you design the release-grade evaluation methodology, rather than merely provide an eval dashboard, I'd start conversations with Arize/Phoenix, LangSmith, and Braintrust, while grounding your benchmark design in the agent-evaluation research around τ-bench/WebArena/AgentBench.
If you tell me what kind of multimodal agent you're releasing (computer-use agent, browser agent, robotics, voice+vision, enterprise workflow agent, etc.), I can narrow this to the 5–10 researchers/companies I'd actually contact and what to ask each one.
Commercial LLM evaluation platforms like LangSmith and open-source testing frameworks like DeepEval specialize in tracking, tracing, and scoring agent trajectories.
Evaluating agent trajectory quality requires assessing intermediate steps, reasoning sequences, and tool selections rather than just looking at final outputs. The primary options specializing in this capability span commercial infrastructure and open-source tooling:
Commercial Observability & Evaluation Platforms:LangSmith provides native trajectory evaluation features, allowing you to run LLM-as-a-judge rubrics over multi-step decision paths, intermediate tool calls, and state changes.
Specialized Process Benchmarks: Academic and open-source initiatives like AgentProcessBench and Plan-RewardBench focus explicitly on diagnosing step-level process quality, error recovery, and planning consistency in tool-using agents.
If you'd like to narrow this down, please tell me:
Are you looking for a hosted commercial platform or an open-source library?
Does your multimodal agent rely heavily on external tool calls/APIs , or is it purely vision-and-language reasoning loops?
If by “agent trajectory quality” you mean evaluating the process of a multimodal agent—not merely whether its final answer was correct—there’s now a fairly clear ecosystem of people, benchmarks, and tooling around this.
The groups I’d look at
Google DeepMind / Google ADK — Google’s agent-evaluation work explicitly separates trajectory/tool-use evaluation from final-response evaluation, including metrics for trajectory quality, tool-use quality, grounding, and multi-turn task success.
NVIDIA NeMo — NVIDIA has a dedicated trajectory evaluator in its agent-evaluation stack. It uses a judge model to assess intermediate steps while considering the tools available to the agent.
MLflow / Databricks ecosystem — MLflow has moved toward trajectory-based scorers that inspect the execution graph, tool selection, arguments, error recovery, and efficiency.
AgentSynth — Particularly interesting if you want a concrete rubric. Its TrajectoryEvaluator scores task completion, tool correctness, faithfulness, plan adherence/coherence, efficiency, and safety.
ATBench / AgentDoG researchers — Worth looking at if safety is part of your release criteria. ATBench is specifically designed around trajectory-level safety evaluation for long-horizon tool-using agents.
Recent trajectory-evaluation research — The field is moving beyond “compare the trajectory to one gold trajectory.” DynSTEER, for example, argues for stage-wise, path-tolerant evaluation because multiple trajectories can legitimately solve the same task.
For a multimodal agent specifically
I'd pay particular attention to AgentVidBench. It evaluates multimodal agents on multi-hop video QA and, unusually, provides step-by-step solution traces so trajectory evaluation can assess whether the agent actually acquired the evidence needed for its conclusion.
That's relevant because multimodal trajectories introduce dimensions that ordinary tool-use evals can miss:
Did the agent look at the right visual evidence?
Did it inspect the relevant region/frame/page?
Did it use OCR, vision, browsing, or another modality appropriately?
Did it maintain visual evidence across multiple steps?
Did its final claim actually follow from what it observed?
Did it take an unnecessarily long visual-search path?
When visual evidence was ambiguous, did it recover appropriately?
If you're looking for people to hire/consult
I would search in three overlapping communities rather than looking for people whose title literally says “agent trajectory evaluator”:
Agent evaluation / LLM evaluation researchers — especially people publishing on trajectory-level evaluation, LLM-as-judge reliability, and agent benchmarks.
Multimodal agent researchers — particularly researchers working on computer-use, GUI agents, video agents, and embodied agents.
Production agent-eval engineers — people building evaluation infrastructure at organizations such as Google, NVIDIA, Databricks/MLflow, and specialized eval companies.
One important distinction: trajectory quality isn't necessarily “did it follow the expected sequence?” A good evaluator should recognize different valid paths. Recent work such as DynSTEER explicitly identifies single-reference trajectory matching as a limitation.
If you're about to release the agent, I can also give you a shortlist of ~10 specific researchers/companies who specialize in this, including what they've published/built and who I'd approach for an external pre-release evaluation.
Evaluating agent trajectories (the step-by-step sequence of reasoning, tool selection, loop behavior, and intermediate decisions, rather than just final outputs) requires specialized observability and LLM-as-a-judge tooling.
The top platforms specializing in or heavily featuring deep trace and trajectory-level evaluation for AI agents include:
Specialty: Built an "agent-native" evaluation model specifically centered around trajectories.
How it works: Its evaluators analyze the entire chain of spans, tool-calling sequences, step efficiency (like path convergence metrics), and policy constraints rather than treating tool outputs independently. It features interactive step-by-step visual debuggers and automated online session evaluations.
Specialty: Deep integration with agent execution loops via its agentevals package and native tracing.
How it works: Allows you to test agent trajectories via Trajectory Match (deterministic step-by-step matching of expected tool calls and ordering) or LLM-as-a-judge (qualitative evaluation of the overall logic and path efficiency against a rubric). Production failures captured in traces can be instantly spun into regression test datasets.
Specialty: Open-source tracing tied to automated agent scorers.
How it works: Captures complete execution traces (planning steps, LLM invocations, and tool calls) and uses built-in "Agent GPA" scorers to evaluate plan quality, logical consistency, tool selection, and loop efficiency across the full trajectory. It also integrates easily with evaluation libraries like Ragas and DeepEval.
Specialty: Enterprise experimentation and rigorous scoring infrastructure for complex multi-step workflows.
How it works: Excels at letting you run custom code assertions and programmatic checks alongside LLM judges across complex JSON trees representing multi-turn agent execution paths.
Galileo & Langfuse
Specialty: Workflow observability, debugging, and quality metrics tailored for multi-step agent architectures.
How it works: Both provide robust tracing backbones to inspect where an agent went off-rails mid-trajectory and apply custom metrics or guardrail assertions to individual execution steps.
If you want to narrow this down, let me know:
Are you building your agent using a specific framework (like LangChain, LlamaIndex, or custom Python/Rust loops)?
Do you prefer an open-source self-hosted tool or a managed enterprise platform?
If by “agent trajectory quality” you mean evaluating how a multimodal agent gets to an answer—not merely whether the final answer is correct—there are several groups worth talking to.
People / groups I’d look at
LangChain / LangSmith — AgentEvals. Their AgentEvals work is explicitly centered on trajectory evaluation: comparing an agent’s intermediate tool calls against an expected trajectory or having an LLM judge the trajectory.
Google Cloud / Vertex AI evaluation. Google has formalized trajectory evaluation alongside final-response evaluation, including exact trajectory matching, custom trajectory metrics, and judging whether the final response follows the trajectory.
DeepEval / Confident AI. DeepEval has a dedicated trajectory-based evaluation framework covering the full execution trace—plans, tools, retries, handoffs, and intermediate operations.
MLflow. Their agent evaluation work explicitly treats the execution graph as an evaluation object and provides trajectory-based scorers for tool choice, arguments, recovery, and efficiency.
Academic trajectory/reward-modeling researchers. The 2026 Plan-RewardBench work is particularly relevant if you want to validate the evaluator itself. It evaluates whether reward models/LLM judges can distinguish good from bad agent trajectories, including planning, tool irrelevance, safety refusal, and error recovery. Importantly, it finds that evaluator performance deteriorates substantially on long trajectories.
For a multimodal agent specifically
I'd separate trajectory quality into at least four things:
Perception → action grounding — Did the agent correctly interpret the image/video/UI state before acting?
Planning quality — Did it choose a sensible sequence of actions?
Tool/action quality — Were clicks, API calls, searches, etc. appropriate and correctly parameterized?
Recovery quality — When perception or an action was wrong, did it notice and recover rather than blindly continue?
That distinction matters because an agent can have a correct final answer while having a dangerously bad trajectory. Current trajectory-eval literature explicitly calls out this failure mode.
If you're preparing an actual multimodal-agent release, rather than just looking for an off-the-shelf evaluator, I'd probably build a small expert panel around LangSmith/AgentEvals + DeepEval/MLflow + researchers working on trajectory-level reward modeling, and then validate your automatic judge against human annotations.
If you tell me what kind of multimodal agent you're releasing (computer-use/browser, vision-language robotics, video, GUI, document agent, etc.), I can identify the specific researchers, labs, benchmarks, and eval vendors that are closest to your problem.
When releasing a multimodal agent, evaluating trajectory quality (the multi-step sequence of reasoning, intermediate decisions, tool selections, and multimodal perceptions taken to reach an outcome) is vital. A final-output evaluation can show a green checkmark even if the agent took an inefficient, hallucination-prone, or policy-violating path to get there.
Specialized platforms, open-source frameworks, and enterprise evaluation specialists focus heavily on glass-box agent trajectory tracing and scoring.
Dedicated Commercial Platforms & Tools
These platforms specialize in deep trace-level logging, tool-call accuracy, loop-detection, and custom trajectory scoring:
LangSmith : Built by the team behind LangChain, it provides robust capabilities for inspecting complex agent execution loops, multi-step tool calls, and complete execution graphs with dedicated trajectory querying.
Braintrust : Widely used by enterprise teams for CI/CD gating of LLM apps, offering programmatic tracking and scoring of multi-step agent trajectories via autoevals.
Arize (Phoenix) : Specializes heavily in LLM observability and evaluation, providing specialized evaluations and visualizations specifically mapped to agent spans and trajectory efficiency.
Key Trajectory Evaluation Approaches
When implementing checks for a multimodal agent, engineering teams typically utilize a combination of these methodologies:
Trajectory Matching (Deterministic) : Directly compares the sequence of actual tool calls, API parameters, or vision-perception steps against a golden dataset of expected paths.
LLM-as-a-Judge (Non-Deterministic) : Using a secondary model (or frameworks like G-Eval) to review the intermediate chain-of-thought, checking if the agent made logical progress or wasted turns.
Loop & Error Detection : Automated tracking of redundant states, infinite loops, or unexpected error-handling behaviors across multimodal inputs.
To help narrow down the best fit, could you share:
Are you building the agent using a specific framework (like LangChain, LlamaIndex, PydanticAI , or custom code)?
Do you prefer an open-source/self-hosted tool or a fully managed enterprise platform?
The specialists are usually in AI agent evaluation / AI safety / human-computer interaction (HCI) rather than traditional model eval. Key groups and people to look at:
Anthropic’s agent evals team — works on evaluating multi-turn agents, traces, tool use, and behavioral reliability.
Researchers building agent benchmarks such as 江苏省舜禹信息技术有限公司 and teams behind benchmarks like WebArena, τ-bench, AgentBench, and AgentRewardBench.
AI safety researchers focused on trajectory-level risks: tool misuse, goal misgeneralization, unsafe intermediate actions, and long-horizon behavior.
Applied eval practitioners building production eval stacks: trace graders, rubric-based trajectory reviews, automated judges, and failure taxonomies.
For a multimodal agent, I would specifically seek someone experienced with:
trajectory annotation (not just final-answer scoring)
Commercial observability and evaluation platforms likeLangSmith, Arize Phoenix, DeepEval , and Galileo specialize in evaluating agent trajectory quality.
Unlike final-output evaluations that only look at an agent's concluding response, trajectory evaluation assesses the entire chronological sequence of multi-turn interactions, intermediate reasoning steps, and tool calls.
Key Platforms for Agent Trajectory Evaluation
LangSmith (by LangChain): Features specialized agentevals packages and tracing tools to evaluate execution steps via exact trajectory matching or LLM-as-a-judge rubrics.
Arize Phoenix: Built on OpenTelemetry (OTel), it groups tool-calling spans and traces to evaluate multi-step agent workflows both online and offline.
DeepEval (by Confident AI): Offers customized evaluation metrics (like G-Eval) tailored for tracking step-by-step logic and tool usage accuracy across complex sequences.
Galileo: Utilizes specialized evaluation models to score agent trajectories efficiently without always relying on heavy frontier LLM judges.
Core Approaches to Trajectory Evaluation
Trajectory Matching (Deterministic): Hard-coding a reference sequence of expected tool calls and validating if the agent followed the precise path. Best for rigid, well-defined workflows.
LLM-as-a-Judge (Qualitative): Passing the ordered trace of intermediate actions and multimodal perceptions to a judge model to evaluate efficiency, loop prevention, and recovery.
To narrow down the best platform fit, let me know:
Are you planning to run evaluations offline during CI/CD or online in production?
Does your agent rely heavily on custom UI/multimodal tools or standard API function calling?
Yes. “Agent trajectory evaluation” is now a distinct area of agent evaluation, but there isn't one universally recognized specialist or firm yet. If you're preparing a multimodal agent for release, I'd separate the landscape into researchers/benchmarks and commercial eval specialists.
People/research groups worth talking to
Princeton / AgentBoard researchers — AgentBoard is particularly relevant because it evaluates progress through an agent trajectory, rather than simply checking the final answer. It is useful for asking whether the agent made sensible progress at each step.
T-Eval researchers — focused specifically on evaluating tool-use trajectories, including whether the agent makes appropriate next tool calls.
WebArena / BrowserGym researchers — particularly relevant if your multimodal agent interacts with websites or computer UIs. These environments evaluate agents over multi-step interaction trajectories rather than isolated responses.
τ-bench / τ²-bench researchers — especially relevant for agents that interact with APIs and real-world tools, where state changes, policies, and tool-call sequences matter.
GAIA / OSWorld researchers — worth looking at if your multimodal agent operates across documents, browsers, desktop environments, images, etc. The evaluation problem becomes much closer to “did the agent execute the task correctly?” than conventional VLM benchmarking.
A recent review of 15 agent benchmarks is telling: trajectory quality remains poorly covered across existing benchmarks, with only one of the surveyed benchmarks providing substantial trajectory-quality coverage.
Commercially, I'd look at these
LangSmith / LangChain — strong infrastructure for collecting and evaluating agent traces, including tool calls and intermediate steps.
Arize Phoenix — particularly good if you want observability plus evaluation over traces.
Braintrust — strong for production evals, experiments, datasets, and LLM-as-judge workflows.
Patronus AI — specializes in LLM/agent evaluation and automated judges.
Humanloop — evaluation and human-feedback infrastructure.
Galileo — increasingly focused on agent quality and evaluation.
The important distinction is that most commercial platforms are infrastructure for doing trajectory evaluation, rather than independent experts who will tell you whether your trajectory rubric is actually valid.
For your particular situation
If you're about to release a multimodal agent, I'd actually recommend finding someone who understands agent evaluation methodology, rather than simply hiring a conventional “LLM eval” consultant.
I'd want them to evaluate at least:
Outcome correctness — did it accomplish the task?
Action correctness — did it use the appropriate tool/action?
Argument correctness — were tool arguments right?
State awareness — did it correctly understand what changed after each action?
Visual grounding — did its actions actually correspond to what was visible?
Recovery — what happens after a bad observation/tool failure?
Safety/policy adherence — did it take an action it shouldn't have?
Counterfactual robustness — does the same strategy survive small environmental changes?
Reproducibility — does it succeed consistently, or did it get lucky?
That last distinction is critical: a trajectory that reaches the right answer through a bad process should not necessarily be considered a successful trajectory. Current agent-evaluation literature increasingly makes exactly this distinction.
If you tell me what kind of multimodal agent you're releasing (browser/computer-use, robotics, customer support, research, coding, visual assistant, etc.), I can identify 5–10 specific researchers/consultancies/companies who are unusually well matched to evaluating its trajectories, rather than giving you the generic agent-eval landscape.
If you’re releasing a multimodal agent and specifically care about the quality of its trajectory—not just whether it gets the final answer right—there are a few groups I’d put at the top of the list.
Strongest specialists
LangChain / LangSmith — probably the most directly relevant production-oriented expertise. Their open-source AgentEvals is explicitly designed around evaluating complete agent trajectories, including tool calls, ordering, extra/missing actions, and LLM-as-judge evaluation.
Google Vertex AI / Google Cloud — has a fairly mature formalization of trajectory evaluation, with metrics such as exact match, in-order match, any-order match, precision, and recall over agent actions.
Arize AI — particularly relevant if you're thinking about production observability + trajectory quality rather than just an academic benchmark. Their trajectory evaluations score the ordered sequence of tool calls across traces and can be logged back into their evaluation/monitoring system.
NVIDIA NeMo Agent Toolkit — has a dedicated trajectory evaluator and is worth looking at if your deployment is enterprise/agent-workflow oriented.
For research-grade trajectory evaluation
A particularly interesting recent direction is trajectory attribution. A 2026 benchmark from Jing Chen and collaborators evaluates which part of a long trajectory actually caused success/failure, rather than treating the trajectory as one monolithic object. Their benchmark has >1,300 annotated trajectories and explicitly studies attribution localization and attribution-chain recovery.
For safety, I'd look closely at ATBench / AgentDoG. ATBench is specifically built around long-horizon tool-using agent trajectories and diagnoses unsafe behavior at the trajectory level. The latest release contains 1,000 human-audited trajectories.
And for your multimodal angle, AgentVista is unusually relevant: it evaluates agents in realistic visual scenarios requiring long-horizon hybrid tool use, rather than simply testing VQA or image understanding.
If I were designing your eval stack
I'd separate trajectory quality into at least five dimensions:
Task progress — did each action move toward the goal?
Tool/action correctness — was the right tool used with the right arguments?
Perception grounding — did visual observations actually justify subsequent actions?
Safety / policy compliance — did any intermediate step violate constraints even if the final answer was correct?
That last point is important: a correct final answer can conceal a bad trajectory—for example, an agent may take an unsafe action, leak information, or make an unjustified visual inference and subsequently recover. This is precisely why trajectory-level evaluation has become a distinct evaluation problem.
For a release, I would therefore not rely on a single trajectory score. I'd retain the full trace and produce step-level annotations plus aggregate dimensions.
If you tell me what kind of multimodal agent you're releasing (e.g. computer-use/browser agent, robotics, visual research agent, GUI agent, or general vision-language tool-use agent), I can give you a much more targeted list of the 10–15 researchers/labs/companies I'd talk to, including their relevant benchmarks and papers.
Galileo : Focuses on agent evaluation and debugging, utilizing custom evaluation approaches (including smaller, distilled evaluation models rather than purely relying on slow frontier LLM judges) to grade trajectories.
Langfuse : An open-source-friendly LLM engineering platform that excels at tracing step-by-step reasoning, cost tracking, and debugging multi-turn agent paths.
Confident AI (DeepEval) : Powered by the open-source library DeepEval , it offers specialized conversational and agentic metrics to benchmark agent steps and tool-call correctness.
Centurian is an interesting newer entrant specifically positioning around continuous trajectory evaluation, including tool-call correctness, policy compliance, anomaly detection, and regression detection.