In March 2026, a malicious package in LiteLLM’s Python dependency tree harvested credentials from production inference infrastructure. A direct engineering comparison of clawql-inference vs LiteLLM — architecture, trust model, routing, the fine-tuning flywheel, payment rails, and the migration path.
This pairs with Mini Shai-Hulud and supply chain controls, immutable releases, and the twelve layers of LLM cost for model escalation and the fine-tuning flywheel. Product reference: docs.clawql.com/inference/clawql-inference.
March 2026
A malicious package in LiteLLM’s Python dependency tree was used as an infection vector. Teams running LiteLLM in production CI/CD pipelines and inference infrastructure were exposed to credential harvesting. The LiteLLM team wasn’t responsible for the upstream package. Supply chain attacks against Python packages are a documented, growing, now open-sourced threat class — the Mini Shai-Hulud campaign that followed in May 2026 proved the tooling for this is freely available and actively copied. What happened in March is a consequence of how Python’s ecosystem works, not a failure specific to the LiteLLM codebase.
Teams evaluating their inference gateway after the incident were asking a question with no clean prior answer: what does it actually mean to trust a piece of infrastructure that sits between your application and every external model call?
This post runs through the comparison directly. Where LiteLLM is strong, that’s stated. Where clawql-inference is ahead, that’s demonstrated with specifics. Where clawql-inference has gaps, those are named.
What LiteLLM Is
LiteLLM is a Python proxy that provides a unified OpenAI-compatible API surface across 100+ LLM providers. You point your existing OpenAI SDK at LiteLLM, configure your providers, and your code reaches Anthropic, Groq, Together, Bedrock, Vertex, and everything else without changing a line of application code.
40,000 GitHub stars. YC-backed. Used in production at Netflix. The 100+ provider coverage represents years of maintenance work across dozens of provider API formats. That’s worth keeping in mind throughout the comparison — the question is whether LiteLLM covers what your team actually needs, not whether it’s a bad product.
Architecture
LiteLLM is a proxy. It sits between your application and LLM providers, normalizes the API surface, and routes calls. The proxy’s job ends when the provider responds. It has no knowledge of what your agent was trying to accomplish, whether the response was useful, what it cost relative to your plan limits, or whether the same response was requested 20 minutes ago.
clawql-inference closes the production loop. Every call passes through a decorator stack:
Traced → Observed → Entitlement → TokenEfficiency → Cache → Fallback → Configured → provider
Each layer is independently toggleable. ObservedInferenceGateway writes to the call store as part of the completion pipeline, before the response returns to the caller. The output of one call informs routing decisions on the next. Production traffic accumulates into training data. Training data produces a fine-tuned model. The model registers back into the routing tier.
LiteLLM logs to Langfuse or a configured callback. If the callback fails, the log is lost. If it’s misconfigured, nothing is recorded. The call store in clawql-inference is the primary record, written synchronously as part of the happy path.
Lesson: observation as a callback is observation when it works. Observation as a pipeline layer is observation always.
Trust Model
LiteLLM:
- Python codebase with hundreds of dependency packages
- The March 2026 compromise came through a dependency, not LiteLLM’s own code
- Packages install from upstream PyPI with no content-addressed pinning
- No SBOM per release
- No startup integrity verification
clawql-inference:
- TypeScript-native, same monorepo as the full ClawQL platform
- Every release generates a CycloneDX SBOM recording every dependency and its hash
- Container images Cosign-signed with SBOM attached
- Layer 0 release manifest records SHA-256 of every artifact, anchored to Arweave
clawql doctor --smokeverifies the installed binary against the manifest hash on startup — a tampered binary fails before handling any requests- Integration catalog mirrors all provider adapters to ClawQL-controlled Harbor/R2 — upstream updates cannot reach users silently
The startup verification is what the March 2026 incident specifically called for. A compromised dependency that modifies the installed package after the fact keeps running. A modified clawql-inference binary is caught at next startup.
See immutable releases for the full Arweave anchoring and content-hash verification pattern.
Lesson: “installed from the official package” stopped being a sufficient trust claim in March 2026. Mirror, sign, pin, verify on boot.
Routing
LiteLLM distributes calls based on load balancing rules and cost configuration. It knows whether the provider returned HTTP 200. It doesn’t know whether the response was actually useful.
Model escalation in clawql-inference is outcome-driven. The router tracks whether tasks succeeded, whether confidence was high, and whether drift signals indicate a more capable tier is needed.
| Tier | Default model | Role |
|---|---|---|
| Frugal | ollama/phi4 | Local Ollama or custom fine-tuned model |
| Standard | groq/llama-3.3-70b | Primary cloud tier |
| Frontier | anthropic/claude-sonnet-4 | Escalation target |
Escalation rules: decomposed child tasks start at Frugal, top-level tasks start at Standard, failure escalates one notch and never skips tiers, combined_drift > 0.3 triggers agent coordination rather than further escalation. Every escalation is WORM-logged with the specific failure signal that caused it.
export CLAWQL_INFERENCE_ROUTING_ENABLED=1
export CLAWQL_INFERENCE_MODEL_FRUGAL=ollama/phi4
export CLAWQL_INFERENCE_MODEL_STANDARD=groq/llama-3.3-70b
export CLAWQL_INFERENCE_MODEL_FRONTIER=anthropic/claude-sonnet-4
# Pin to a single model and bypass escalation entirely
export CLAWQL_INFERENCE_MODEL_PIN=anthropic/claude-sonnet-4
LiteLLM’s retry logic handles transient provider errors by retrying the same model. Escalation to a more capable model on quality failure is a different mechanism entirely.
Semantic Cache
LiteLLM uses Redis for exact-match or near-match caching. The same prompt returns the cached response. Textually different prompts that ask the same question don’t match.
clawql-inference embeds every request and compares against cached embeddings using cosine similarity. “Summarize this document” and “give me a summary of this doc” hit the cache even though the text differs.
export CLAWQL_INFERENCE_SEMANTIC_CACHE=1
export CLAWQL_INFERENCE_CACHE_THRESHOLD=0.92
export CLAWQL_INFERENCE_CACHE_TTL=24h
export CLAWQL_EMBEDDING_MODEL=text-embedding-3-small
The cache is backed by pgvector with a Postgres inference database configured, falling back to in-memory for single-instance deployments.
Three design decisions worth naming:
Embedding failures fail open. If the embedding service is unavailable, the request proceeds to live inference. The cache is an optimization, never a gate.
Write operations are never cached. Only read-safe operations are eligible for cache hits. This is enforced in the gateway stack, not delegated to caller discipline.
Cache hits are flagged with cache_hit: true in the call store. The --exclude-cache-hits flag on export removes them from fine-tuning datasets, because training on cached responses instead of model outputs is a data quality problem.
The Fine-Tuning Flywheel
LiteLLM routes inference. The output of your agents today has no connection to what your agents cost or produce tomorrow. Every call is independent.
clawql-inference captures production traffic as training data, filters it by verified outcome, scrubs PII, and closes the loop back into the model tier that handles that task type.
# Export verdict-filtered training data
clawql inference export \
--verdict passed \
--tier frugal \
--format openai-jsonl \
--min-score 0.8 \
--output ./training-data/$(date +%Y-%m).jsonl
# Submit fine-tuning job
clawql inference finetune \
--dataset ./training-data/2026-07.jsonl \
--base-model gpt-4o-mini \
--provider openai
# Check status
clawql inference finetune status --job-id ftjob_abc123
# Register the resulting model as the new Frugal tier
clawql inference finetune register \
--job-id ftjob_abc123 \
--tier frugal \
--alias phi4-production-v3
After registration, model escalation uses the custom model automatically for matching task types at Frugal tier. The cost drops. Quality on your specific workload is higher than the generic model it replaced because it was trained on your production traces, not on a general corpus.
Every export writes a WORM manifest recording sample hashes, filter criteria, Presidio scrub version, and Merkle root. If you’re processing sensitive documents, you can prove what entered the training dataset and that PII was removed before export.
There is no path in LiteLLM from “this call succeeded” to “the model handling similar calls next month is cheaper and more accurate.” The Flywheel is the mechanism that turns inference spend into a proprietary model asset. It compounds. LiteLLM doesn’t have it.
Payment Rails
LiteLLM has API key management and basic spend tracking.
clawql-inference ships four billing modes via clawql-payments:
Plan entitlements — pre-call monthly caps with structured 402 responses:
export CLAWQL_PAYMENTS_ENFORCE_INFERENCE=1
# HTTP 402 with: { "error": { "type": "insufficient_quota", "reset_at": "...", ... } }
Stripe metered billing — post-call usage reporting to Stripe Billing Meters:
export CLAWQL_PAYMENTS_REPORT_STRIPE_METER=1
export STRIPE_METER_EVENT_NAME=clawql_inference_calls
x402 per-request USDC micropayments — agents pay per call, no account required:
export CLAWQL_X402_ENFORCE=1
# HTTP 402 with x402 challenge
# Agent pays in USDC, presents proof, proceeds
# Sub-second settlement on Base L2
MPP session-based payments — pre-authorized spending limits for high-frequency workloads:
export CLAWQL_MPP_ENABLED=1
# Agent pre-authorizes a budget once
# Micropayments stream within the session
ClawQL’s own managed hosted tiers run on the same package. The EntitlementEnforcedGateway that enforces your plan limits is the same code your own deployment runs. Every payment event links to the specific inference call that triggered it via correlation_id in the WORM trail.
Three Usage Systems
clawql-inference tracks usage three ways, serving different purposes. Conflating them produces confusing dashboards.
Plan entitlements at $CLAWQL_HOME/Payments/usage.json — monthly inference call counts checked before each completion. The gate that enforces plan limits. Resets monthly.
Inference call store at calls.jsonl or a Postgres table — full InferenceRecord per completion with tokens, latency, model, tier, verdict, correlation_id. This is what clawql inference spend queries. The source for fine-tuning export. The audit record.
Virtual key USD budgets at $CLAWQL_HOME/Inference/virtual-keys.json — per-team spending caps in estimated USD. Separate from the plan limit.
All three can be active simultaneously. clawql inference spend output and your plan usage counter measure different things over different windows.
Observability
LiteLLM fires a callback after each completion. If the callback fails, the event is lost.
clawql-inference writes to the call store as part of the completion pipeline, before the response returns to the caller. The write is synchronous and mandatory, not optional and asynchronous.
On top of the call store:
- OTLP spans to Tempo (
CLAWQL_ENABLE_OTEL_TRACING=1) - Langfuse work-trace emission via OTLP (
CLAWQL_ENABLE_LANGFUSE=1, opt-out per ADR 0005) clawql inference logs— tail recent completions with model/tier/time filtersclawql inference trace --correlation-id <id>— full lifecycle of a single requestclawql inference spend --group-by model|tier|team|provider
Every escalation decision, cache hit, fallback attempt, entitlement check, and payment event carries the same correlation_id. A compliance query for a specific agent session is a single index lookup, not a join across several log sources.
Policy Manifest
LiteLLM configuration is environment variables and a YAML config file with no versioning or content-addressed anchoring.
clawql-inference reads a policy manifest at $CLAWQL_HOME/Inference/policy.yaml that merges with environment variables at runtime (env wins on conflicts):
policyVersion: '2026.07.01'
inference:
escalation:
enabled: true
tierMap:
frugal: ollama/phi4
standard: groq/llama-3.3-70b
frontier: anthropic/claude-sonnet-4
cache:
enabled: true
threshold: 0.92
ttl: 24h
fallback:
enabled: true
frugal: [ollama/phi4-backup, openai/gpt-4o-mini]
pipelineWorker:
enabled: true
schedule: '0 2 * * 0'
minSamples: 500
observability:
profile: external
resolveInferenceEffectiveEnv() is the same merge function used by both clawql inference policy show and clawql inference serve. What policy show displays is what the gateway runs.
Multi-Instance Deduplication
Running multiple inference serve replicas behind a load balancer creates a problem: the fine-tuning pipeline worker is a cron job that should run once per tick, not once per replica per tick.
When CLAWQL_INFERENCE_DATABASE_URL is configured, clawql-inference uses Postgres advisory locks keyed by pipeline schedule and UTC minute. One replica runs the tick. The others skip it without blocking. The lock releases when the tick completes. No external coordination service required.
Migration Path
Any client that respects OPENAI_BASE_URL works without code changes:
# Install
curl -fsSL https://clawql.com/install | bash
# Start the gateway
clawql inference serve --port 8080
# Point your existing client at it
export OPENAI_BASE_URL=http://127.0.0.1:8080/v1
export OPENAI_API_KEY=your-existing-key # passed through to the provider
Provider configuration:
export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=sk-ant-...
export OLLAMA_BASE_URL=http://127.0.0.1:11434
export GROQ_API_KEY=gsk_...
Use provider/model format for explicit routing (anthropic/claude-sonnet-4, ollama/phi4) or bare model IDs (gpt-4o, claude-sonnet-4) when a matching provider is configured.
Start with observation only and enable features incrementally:
# Week 1: call store recording only
clawql inference serve
# Week 2: semantic cache
export CLAWQL_INFERENCE_SEMANTIC_CACHE=1
# Week 3: model escalation
export CLAWQL_INFERENCE_ROUTING_ENABLED=1
# Week 4: fallback chains
export CLAWQL_INFERENCE_FALLBACK_ENABLED=1
export CLAWQL_INFERENCE_FALLBACK_STANDARD=groq/llama-3.3-70b,anthropic/claude-haiku-4
# Month 2: first fine-tuning export
clawql inference export --verdict passed --format openai-jsonl --output ./export.jsonl
After a week of production traffic, the call store has enough data to show whether semantic caching and model escalation make sense for your workload before you commit to either.
Where LiteLLM Still Leads
Provider breadth. 100+ providers, years of maintenance work. clawql-inference ships OpenAI, Anthropic, and Ollama with a plugin API for others. Teams that need Azure OpenAI, Bedrock, Vertex, Fireworks, and Replicate out of the box will hit gaps.
Python ecosystem. LangChain, LlamaIndex, DSPy, and most ML tooling are Python-native. clawql-inference is TypeScript. The REST surface is compatible with any HTTP client, but Python SDK integration is smoother with LiteLLM.
Managed hosting. LiteLLM Enterprise has a hosted option. ClawQL’s managed tiers include clawql-inference but are a broader platform commitment than a standalone hosted inference gateway.
Community size. More Stack Overflow answers, more blog posts, more example integrations. The clawql-inference documentation is thorough but the community is smaller.
Comparison Table
| Dimension | clawql-inference | LiteLLM |
|---|---|---|
| Language | TypeScript-native | Python |
| Supply chain | Cosign-signed, SBOM per build, Arweave manifest, startup hash verify | Standard PyPI; March 2026 compromise via dependency |
| Provider breadth | OpenAI, Anthropic, Ollama built-in; plugin API | 100+ providers |
| Routing | Outcome-driven model escalation | Load balancing and cost rules |
| Semantic cache | Embedding similarity, pgvector backend | Redis exact/near-match |
| Fine-tuning flywheel | Verdict-filtered export → PII scrub → job → tier-map registration | None |
| WORM audit | Every routing decision, cache hit, escalation, payment event with correlation_id | Callback-based logging |
| Payment rails | Stripe + x402 + MPP + ACP/AP2 planned | API key management |
| Policy manifest | Versioned YAML, governs live gateway, merged with env | Env + config file |
| Multi-instance dedup | Postgres advisory locks for pipeline worker | External coordination required |
| Startup integrity | clawql doctor --smoke verifies binary hash | None |
clawql-inference is available as a standalone npm package (npx clawql-inference) and as part of the ClawQL platform. Full reference: docs.clawql.com/inference/clawql-inference. Provider plugin docs and migration guide: docs.clawql.com.
