LLM AS JUDGE · HUMAN REVIEW · EVALUATION DASHBOARDS
Finance Evaluation
An interactive workbench for evaluating financial agent responses.
Follow factual claims back to their sources, compare model judgments, review summary quality and explore the evidence behind evaluation metrics.
From agent output to evaluation evidence
- Ground the response. Keep the user query, structured response and retrieved source data together.
- Apply a judge rubric. Check factual claims against evidence and score summaries across nine defined dimensions.
- Compare model judgments. Inspect per-model reasoning, voted scores and disagreements.
- Add human review. Correct labels, record rationales and compare model outputs with a golden dataset.
- Measure and diagnose. Use confusion matrices, agreement metrics and error reports to guide the next evaluation iteration.
The dashboards use saved evaluation outputs from the original project. Fact-check metrics retain its original label-merging rules; unreviewed claims use the saved judge label. Summary metrics compare saved model scores with the golden dataset.
Explore the evaluation workbench
Fact Check Dashboard
Browse all 350 financial agent responses and human-review status.
Grounded Fact Review
Trace highlighted claims to source evidence, inspect judge reasoning and label disagreements.
Summary Scoring
Compare judge scores across nine dimensions, inspect rubric criteria and add human scores.
Fact Check Metrics
Explore precision, recall, F1 and confusion-matrix drill-downs at claim and response level.
Model Evaluation Metrics
Compare exact match, weighted kappa, adjacent accuracy, MAE and correlation with the golden dataset.
Fact Check Jobs
Browse saved job inputs and results, or try a browser-local sample evaluation.
Summary Scoring Jobs
Inspect score distributions and open side-by-side response and reasoning drawers.
Error Analysis
Review incorrect-claim reports, classify issues and record optimization notes.
Optimization Summary
Explore error categories and notes collected during evaluation review.
Contributor Ranking
See the original review activity with anonymized contributor identities.
Evaluation Handbook
Read the labeling guide, browse page documentation and inspect saved response examples.