Benchmark Report: How EVA Measures Against the CareerNet Benchmark
A blind, judge-scored comparison of EVA’s live product against real student career questions from CareerVillage.org, measured on the same three quality dimensions used by Renaissance Philanthropy to evaluate frontier models.

Key Findings:
96% Completeness
(Substantive Answers)
EVA substantially outperforms top frontier models (~77%) and human forum volunteers (62%) in providing full, thorough coverage of student questions.
99% Factually Clean Rate
In human-adjudicated reviews for citable factual errors, 99% of EVA answers were completely clean, far surpassing the 74% accuracy rate of human forum answers.
59% Coherency Win Rate
In blind head-to-head pairwise comparisons against unconstrained frontier model prose, judges preferred EVA’s structured, feature-rich answers.
What’s inside the report?
Download the full benchmark report for a deep dive into the technical evaluation architecture, calibration methodology, and empirical results:
Methodology & Calibration:
Details on how Claude Fable judges were calibrated against 200 expert-rated items to achieve human-level agreement (QWK 0.58).
Blinding & Integrity Protocols:
How responses were shuffled, pooled, and assigned opaque IDs to prevent judge bias.
Human-in-the-Loop Verification:
How 100% of factual error flags and the 15 lowest-scoring responses were qualitative reviewed by human experts.
Detailed Dimension Breakdown:
Full performance charts across Completeness, Correctness, and Coherency.
Get the full report today.
Download the report for comprehensive survey insights, statistical pipelines, judge calibration curves, and comparisons against frontier AI baselines.

