Tag
Benchmarks
2essays ·All tags
When Rubrics Become Rewards
A fair agent benchmark is not only a scoreboard. Two-arm runs with expert criteria produce the traces that become SFT, preference, and GRPO training signal — if you design the evaluation so the verdict is trustworthy.
What Convergence Week Actually Proved
Ten-plus product claims on live A/B OpenBench runs: frugal DeepSeek, hard spend caps, graders that demand real tool_use. On scores 1.0, off scores 0.0 — and every passing pack wraps an RTP reasoning trace the training flywheel can consume.
