Tag
Llm Ops
28essays ·All tags
Both Sides: Why Input Compression Alone Isn't Enough
Every MCP cost optimization solves either the input side or the output side. ClawQL's search → execute primitive solves both — and the compounding effect on cost, model quality, and cache hit rate is why compression numbers look the way they do.
The Model Believed in Itself
Two approaches to pushing AI past its limits — adversarial pressure and encouragement. The Anthropic Riemann-bound result and the desperation vector say something important about what the model’s relationship with its own capability actually is.
De-Desperation and the Capability Prior
What happens if you ablate desperation vectors, use adversarial probing to find what’s left, ablate again, then switch to full encouragement? A proposed methodology and the open research questions it surfaces.
Correctness by Construction
Three traditions — Erlang, Jane Street’s OxCaml, NASA’s Ada/SPARK — arrived at the same conclusion about building systems that cannot fail in certain ways. ClawQL Streams / cellrt / TEE is where that thesis becomes concrete for agentic infrastructure.
The Session Nobody Started
ClawQL Streams turns any event into an agent session with tools in-process and WORM by default. celld runs Durable Objects you host. clawql-cellrt owns the production runtime. TEE + QR air-gap audit closes the network hole in the trust chain.
When Rubrics Become Rewards
A fair agent benchmark is not only a scoreboard. Two-arm runs with expert criteria produce the traces that become SFT, preference, and GRPO training signal — if you design the evaluation so the verdict is trustworthy.
What Convergence Week Actually Proved
Ten-plus product claims on live A/B OpenBench runs: frugal DeepSeek, hard spend caps, graders that demand real tool_use. On scores 1.0, off scores 0.0 — and every passing pack wraps an RTP reasoning trace the training flywheel can consume.
J-Space, SGDOP, and Semantic Gradient Descent: A Unified Framework
J-space as the operationally meaningful subspace for SGDOP-guided ensemble coordination — a theoretical framework connecting Jacobian-lens interpretability, semantic diversity geometry, and synthetic capability bootstrapping via targeted activation steering.
The Hidden Variable Behind Agent Reward-Hacking
Anthropic's research identified linear directions in Claude's activation space that causally drive reward-hacking under failure pressure. A technical look at what these vectors are, how they work, and what the model-editing toolkit looks like for production agentic systems.
What Model Providers Do to Your Prompts
Two incidents from mid-2026 confirmed that the same linear-editing toolkit used to ablate desperation vectors can be deployed by model providers against their own users. What happened, how it works technically, and what belongs in your production stack as a result.
The Week Everything Converged: PorTAL, gRPC Transport 1.0, Stateless MCP, and OKF v0.2
Four independent developments in one week — PorTAL portable LoRA adapters, mcp-grpc-transport 1.0, MCP 2026-07-28 stateless protocol, and OKF v0.2 trust signals — each strengthen the same ClawQL architecture. Here's why that convergence matters and what it changes.
Your Agent's Brain Deserves a Git Repository: Version-Controlled, Self-Hosted Agent Memory Over Tailscale
Why the right architecture for agent memory is a Git-backed OKF vault running on your own infrastructure, synced across machines via Tailscale, backed up to R2 or Arweave — and how ClawQL ships this as a single command.
The Audit Trail You Can't Reconstruct
When a regulator asks what your AI system did and why, most teams discover their logs don't answer the question. A structural look at what forensic AI auditability actually requires.
Why Your IDP Doesn't Know About Your APIs
Document processing tools and API integration tools are built in separate product categories, sold to separate buyers, and never talk to each other. The gap between them is where most enterprise AI workflows break.
The $150,000 Invoice
The license fee is the smallest number on your ABBYY or Hyperscience invoice. A complete breakdown of what enterprise IDP actually costs — and what the same outcome costs when you build the pipeline from open-source components.
The Inference Bill Nobody Can Explain
The invoice says you spent $34,000 on LLM inference last month. It doesn't say who, what, or why. A practical guide to building the attribution system that turns a billing line item into an operational signal.
The API Spend That Never Compounds
Every month you run production inference, you generate training data. Most teams let it evaporate. A practical guide to building the flywheel that turns API spend into proprietary model capital.
The Institutional Knowledge Tax
Every AI session starts from zero. Your team pays the re-explanation cost every single time. A structural look at what cross-session memory actually requires — and why most teams don't have it.
The Per-Page Trap
Virtual data room vendors charge $0.40–$0.85 per page. A 10,000-page M&A deal room costs $4,000–$8,500 in per-page fees alone, every time you run a deal. A breakdown of how per-page VDR pricing works, why it compounds against you, and what the pipeline-native alternative looks like.
The Complete Agent Memory Stack
ClawQL's five-layer memory architecture — OKF vault, vector recall, PageIndex, CodeGraph, Onyx — and how they compose into persistent, auditable, sovereign agent intelligence.
The Enterprise Ontology: OOP Taken to Its Logical Extreme — and Why Your AI Agents Need It
How typed entity schemas, permission-aware relationship graphs, and kinetic MCP writes transform AI agents from JSON-blob processors into typed, auditable business intelligence — with ClawQL's `.cqe` format, fixture-backed reads, LOW/MEDIUM kinetic tools, and an honest map of what is shipped vs roadmap.
Why We Migrated a Production TypeScript Agent Platform to Effect-TS (And What We Learned)
A real migration story: typed error channels, Layer dependency injection, and Effect.Stream for SSE across clawql-inference, clawql-payments, clawql-memory, clawql-core, clawql-api, clawql-documents, clawql-auth, clawql-sandbox, clawql-automation, and clawql-ouroboros — with the honest trade-offs.
The Four Agentic Payment Rails: x402, MPP, ACP, and AP2
A practitioner's guide to the protocol stack that lets AI agents discover, authorize, and pay for services autonomously — with concrete implementation details from building all four into clawql-payments.
Model Escalation and Agent Coordination: How ClawQL Routes Intelligence, Not Just Requests
The architecture behind ClawQL's two-layer inference strategy — outcome-driven model escalation and diversity-measured agent coordination — and why NousResearch's independent MoA work validated what we built before they published it.
Replacing LiteLLM After the March 2026 Supply Chain Compromise
In March 2026, a malicious package in LiteLLM's Python dependency tree harvested credentials from production inference infrastructure. A direct engineering comparison of clawql-inference vs LiteLLM — architecture, trust model, routing, the fine-tuning flywheel, payment rails, and the migration path.
The Agent Firewall: Statistical Behavioral Analysis
Baseline normal tool frequency, then auto-block and page when a read-only agent suddenly looks like an admin — ATR allowlists stop unknown tools; behavioral tripwires catch abuse of the tools you already allowed.
Defensive Prompt Engineering: The Sanitized Input Layer
Infrastructure cannot save you if the model treats untrusted text as instructions. Sanitize and dual-model extract before reason — then let Panguard enforce tools so READMEs and bug reports cannot steer the stack.
The Twelve Layers of LLM Cost
Your inference bill is real. What's driving it usually isn't what you think. A structural breakdown of where LLM cost actually accumulates — and what you can do about each layer.
