Benchmarks

3essays ·All tags

Architecture11 min read

Both Sides of Context Compression

Codemode-style tool gateways solved half of context compression. The other half — what happens to a tool's response after it returns — is still an open problem for the whole category. Here's a measured look at why it matters more than people think.

  • Token Efficiency
  • Mcp
  • Benchmarks
  • Agents
Architecture14 min read

When Rubrics Become Rewards

A fair agent benchmark is not only a scoreboard. Two-arm runs with expert criteria produce the traces that become SFT, preference, and GRPO training signal — if you design the evaluation so the verdict is trustworthy.

  • Agents
  • Benchmarks
  • Llm Ops
  • Ontology
  • Legal Tech
Architecture18 min read

What Convergence Week Actually Proved

Ten-plus product claims on live A/B OpenBench runs: frugal DeepSeek, hard spend caps, graders that demand real tool_use. On scores 1.0, off scores 0.0 — and every passing pack wraps an RTP reasoning trace the training flywheel can consume.

  • Benchmarks
  • Agents
  • Mcp
  • Llm Ops
  • Ouroboros
  • Vault