trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
Outcome-only evaluation fails to detect agent errors that result in correct final answers but follow flawed execution paths.
Researchers demonstrate that standard LLM evaluation, which only checks final outputs, misses 'silent' failures where agents reach the correct result through incorrect tool usage. By using a deterministic environment with fault injection, the study highlights the need for process-aware evaluation to ensure agent reliability.