Finance EvaluationSaved results · Browser-local edits

LLM AS JUDGE · HUMAN REVIEW · EVALUATION DASHBOARDS

Finance Evaluation

An interactive workbench for evaluating financial agent responses.

Follow factual claims back to their sources, compare model judgments, review summary quality and explore the evidence behind evaluation metrics.

350 company responses9 scoring dimensions79 saved fact-check labels32 saved jobs

From agent output to evaluation evidence

  1. Ground the response. Keep the user query, structured response and retrieved source data together.
  2. Apply a judge rubric. Check factual claims against evidence and score summaries across nine defined dimensions.
  3. Compare model judgments. Inspect per-model reasoning, voted scores and disagreements.
  4. Add human review. Correct labels, record rationales and compare model outputs with a golden dataset.
  5. Measure and diagnose. Use confusion matrices, agreement metrics and error reports to guide the next evaluation iteration.

The dashboards use saved evaluation outputs from the original project. Fact-check metrics retain its original label-merging rules; unreviewed claims use the saved judge label. Summary metrics compare saved model scores with the golden dataset.

Explore the evaluation workbench