Ouroboros

3essays ·All tags

Agent Safety8 min read

The Model Believed in Itself

Two approaches to pushing AI past its limits — adversarial pressure and encouragement. The Anthropic Riemann-bound result and the desperation vector say something important about what the model’s relationship with its own capability actually is.

  • Agents
  • Ouroboros
  • Llm Ops
  • Trust Boundaries
Agent Safety10 min read

De-Desperation and the Capability Prior

What happens if you ablate desperation vectors, use adversarial probing to find what’s left, ablate again, then switch to full encouragement? A proposed methodology and the open research questions it surfaces.

  • Agents
  • Ouroboros
  • Llm Ops
  • Trust Boundaries
Architecture18 min read

What Convergence Week Actually Proved

Ten-plus product claims on live A/B OpenBench runs: frugal DeepSeek, hard spend caps, graders that demand real tool_use. On scores 1.0, off scores 0.0 — and every passing pack wraps an RTP reasoning trace the training flywheel can consume.

  • Benchmarks
  • Agents
  • Mcp
  • Llm Ops
  • Ouroboros
  • Vault