Architecture22 min read

Replacing LiteLLM After the March 2026 Supply Chain Compromise

In March 2026, a malicious package in LiteLLM's Python dependency tree harvested credentials from production inference infrastructure. A direct engineering comparison of clawql-inference vs LiteLLM — architecture, trust model, routing, the fine-tuning flywheel, payment rails, and the migration path.

In March 2026, a malicious package in LiteLLM’s Python dependency tree harvested credentials from production inference infrastructure. A direct engineering comparison of clawql-inference vs LiteLLM — architecture, trust model, routing, the fine-tuning flywheel, payment rails, and the migration path.

This pairs with Mini Shai-Hulud and supply chain controls, immutable releases, and the twelve layers of LLM cost for model escalation and the fine-tuning flywheel. Product reference: docs.clawql.com/inference/clawql-inference.


March 2026

A malicious package in LiteLLM’s Python dependency tree was used as an infection vector. Teams running LiteLLM in production CI/CD pipelines and inference infrastructure were exposed to credential harvesting. The LiteLLM team wasn’t responsible for the upstream package. Supply chain attacks against Python packages are a documented, growing, now open-sourced threat class — the Mini Shai-Hulud campaign that followed in May 2026 proved the tooling for this is freely available and actively copied. What happened in March is a consequence of how Python’s ecosystem works, not a failure specific to the LiteLLM codebase.

Teams evaluating their inference gateway after the incident were asking a question with no clean prior answer: what does it actually mean to trust a piece of infrastructure that sits between your application and every external model call?

This post runs through the comparison directly. Where LiteLLM is strong, that’s stated. Where clawql-inference is ahead, that’s demonstrated with specifics. Where clawql-inference has gaps, those are named.


What LiteLLM Is

LiteLLM is a Python proxy that provides a unified OpenAI-compatible API surface across 100+ LLM providers. You point your existing OpenAI SDK at LiteLLM, configure your providers, and your code reaches Anthropic, Groq, Together, Bedrock, Vertex, and everything else without changing a line of application code.

40,000 GitHub stars. YC-backed. Used in production at Netflix. The 100+ provider coverage represents years of maintenance work across dozens of provider API formats. That’s worth keeping in mind throughout the comparison — the question is whether LiteLLM covers what your team actually needs, not whether it’s a bad product.


Architecture

LiteLLM is a proxy. It sits between your application and LLM providers, normalizes the API surface, and routes calls. The proxy’s job ends when the provider responds. It has no knowledge of what your agent was trying to accomplish, whether the response was useful, what it cost relative to your plan limits, or whether the same response was requested 20 minutes ago.

clawql-inference closes the production loop. Every call passes through a decorator stack:

Traced → Observed → Entitlement → TokenEfficiency → Cache → Fallback → Configured → provider

Each layer is independently toggleable. ObservedInferenceGateway writes to the call store as part of the completion pipeline, before the response returns to the caller. The output of one call informs routing decisions on the next. Production traffic accumulates into training data. Training data produces a fine-tuned model. The model registers back into the routing tier.

LiteLLM logs to Langfuse or a configured callback. If the callback fails, the log is lost. If it’s misconfigured, nothing is recorded. The call store in clawql-inference is the primary record, written synchronously as part of the happy path.

Lesson: observation as a callback is observation when it works. Observation as a pipeline layer is observation always.


Trust Model

LiteLLM:

  • Python codebase with hundreds of dependency packages
  • The March 2026 compromise came through a dependency, not LiteLLM’s own code
  • Packages install from upstream PyPI with no content-addressed pinning
  • No SBOM per release
  • No startup integrity verification

clawql-inference:

  • TypeScript-native, same monorepo as the full ClawQL platform
  • Every release generates a CycloneDX SBOM recording every dependency and its hash
  • Container images Cosign-signed with SBOM attached
  • Layer 0 release manifest records SHA-256 of every artifact, anchored to Arweave
  • clawql doctor --smoke verifies the installed binary against the manifest hash on startup — a tampered binary fails before handling any requests
  • Integration catalog mirrors all provider adapters to ClawQL-controlled Harbor/R2 — upstream updates cannot reach users silently

The startup verification is what the March 2026 incident specifically called for. A compromised dependency that modifies the installed package after the fact keeps running. A modified clawql-inference binary is caught at next startup.

See immutable releases for the full Arweave anchoring and content-hash verification pattern.

Lesson: “installed from the official package” stopped being a sufficient trust claim in March 2026. Mirror, sign, pin, verify on boot.


Routing

LiteLLM distributes calls based on load balancing rules and cost configuration. It knows whether the provider returned HTTP 200. It doesn’t know whether the response was actually useful.

Model escalation in clawql-inference is outcome-driven. The router tracks whether tasks succeeded, whether confidence was high, and whether drift signals indicate a more capable tier is needed.

TierDefault modelRole
Frugalollama/phi4Local Ollama or custom fine-tuned model
Standardgroq/llama-3.3-70bPrimary cloud tier
Frontieranthropic/claude-sonnet-4Escalation target

Escalation rules: decomposed child tasks start at Frugal, top-level tasks start at Standard, failure escalates one notch and never skips tiers, combined_drift > 0.3 triggers agent coordination rather than further escalation. Every escalation is WORM-logged with the specific failure signal that caused it.

export CLAWQL_INFERENCE_ROUTING_ENABLED=1
export CLAWQL_INFERENCE_MODEL_FRUGAL=ollama/phi4
export CLAWQL_INFERENCE_MODEL_STANDARD=groq/llama-3.3-70b
export CLAWQL_INFERENCE_MODEL_FRONTIER=anthropic/claude-sonnet-4

# Pin to a single model and bypass escalation entirely
export CLAWQL_INFERENCE_MODEL_PIN=anthropic/claude-sonnet-4

LiteLLM’s retry logic handles transient provider errors by retrying the same model. Escalation to a more capable model on quality failure is a different mechanism entirely.


Semantic Cache

LiteLLM uses Redis for exact-match or near-match caching. The same prompt returns the cached response. Textually different prompts that ask the same question don’t match.

clawql-inference embeds every request and compares against cached embeddings using cosine similarity. “Summarize this document” and “give me a summary of this doc” hit the cache even though the text differs.

export CLAWQL_INFERENCE_SEMANTIC_CACHE=1
export CLAWQL_INFERENCE_CACHE_THRESHOLD=0.92
export CLAWQL_INFERENCE_CACHE_TTL=24h
export CLAWQL_EMBEDDING_MODEL=text-embedding-3-small

The cache is backed by pgvector with a Postgres inference database configured, falling back to in-memory for single-instance deployments.

Three design decisions worth naming:

Embedding failures fail open. If the embedding service is unavailable, the request proceeds to live inference. The cache is an optimization, never a gate.

Write operations are never cached. Only read-safe operations are eligible for cache hits. This is enforced in the gateway stack, not delegated to caller discipline.

Cache hits are flagged with cache_hit: true in the call store. The --exclude-cache-hits flag on export removes them from fine-tuning datasets, because training on cached responses instead of model outputs is a data quality problem.


The Fine-Tuning Flywheel

LiteLLM routes inference. The output of your agents today has no connection to what your agents cost or produce tomorrow. Every call is independent.

clawql-inference captures production traffic as training data, filters it by verified outcome, scrubs PII, and closes the loop back into the model tier that handles that task type.

# Export verdict-filtered training data
clawql inference export \
  --verdict passed \
  --tier frugal \
  --format openai-jsonl \
  --min-score 0.8 \
  --output ./training-data/$(date +%Y-%m).jsonl

# Submit fine-tuning job
clawql inference finetune \
  --dataset ./training-data/2026-07.jsonl \
  --base-model gpt-4o-mini \
  --provider openai

# Check status
clawql inference finetune status --job-id ftjob_abc123

# Register the resulting model as the new Frugal tier
clawql inference finetune register \
  --job-id ftjob_abc123 \
  --tier frugal \
  --alias phi4-production-v3

After registration, model escalation uses the custom model automatically for matching task types at Frugal tier. The cost drops. Quality on your specific workload is higher than the generic model it replaced because it was trained on your production traces, not on a general corpus.

Every export writes a WORM manifest recording sample hashes, filter criteria, Presidio scrub version, and Merkle root. If you’re processing sensitive documents, you can prove what entered the training dataset and that PII was removed before export.

There is no path in LiteLLM from “this call succeeded” to “the model handling similar calls next month is cheaper and more accurate.” The Flywheel is the mechanism that turns inference spend into a proprietary model asset. It compounds. LiteLLM doesn’t have it.


Payment Rails

LiteLLM has API key management and basic spend tracking.

clawql-inference ships four billing modes via clawql-payments:

Plan entitlements — pre-call monthly caps with structured 402 responses:

export CLAWQL_PAYMENTS_ENFORCE_INFERENCE=1
# HTTP 402 with: { "error": { "type": "insufficient_quota", "reset_at": "...", ... } }

Stripe metered billing — post-call usage reporting to Stripe Billing Meters:

export CLAWQL_PAYMENTS_REPORT_STRIPE_METER=1
export STRIPE_METER_EVENT_NAME=clawql_inference_calls

x402 per-request USDC micropayments — agents pay per call, no account required:

export CLAWQL_X402_ENFORCE=1
# HTTP 402 with x402 challenge
# Agent pays in USDC, presents proof, proceeds
# Sub-second settlement on Base L2

MPP session-based payments — pre-authorized spending limits for high-frequency workloads:

export CLAWQL_MPP_ENABLED=1
# Agent pre-authorizes a budget once
# Micropayments stream within the session

ClawQL’s own managed hosted tiers run on the same package. The EntitlementEnforcedGateway that enforces your plan limits is the same code your own deployment runs. Every payment event links to the specific inference call that triggered it via correlation_id in the WORM trail.


Three Usage Systems

clawql-inference tracks usage three ways, serving different purposes. Conflating them produces confusing dashboards.

Plan entitlements at $CLAWQL_HOME/Payments/usage.json — monthly inference call counts checked before each completion. The gate that enforces plan limits. Resets monthly.

Inference call store at calls.jsonl or a Postgres table — full InferenceRecord per completion with tokens, latency, model, tier, verdict, correlation_id. This is what clawql inference spend queries. The source for fine-tuning export. The audit record.

Virtual key USD budgets at $CLAWQL_HOME/Inference/virtual-keys.json — per-team spending caps in estimated USD. Separate from the plan limit.

All three can be active simultaneously. clawql inference spend output and your plan usage counter measure different things over different windows.


Observability

LiteLLM fires a callback after each completion. If the callback fails, the event is lost.

clawql-inference writes to the call store as part of the completion pipeline, before the response returns to the caller. The write is synchronous and mandatory, not optional and asynchronous.

On top of the call store:

  • OTLP spans to Tempo (CLAWQL_ENABLE_OTEL_TRACING=1)
  • Langfuse work-trace emission via OTLP (CLAWQL_ENABLE_LANGFUSE=1, opt-out per ADR 0005)
  • clawql inference logs — tail recent completions with model/tier/time filters
  • clawql inference trace --correlation-id <id> — full lifecycle of a single request
  • clawql inference spend --group-by model|tier|team|provider

Every escalation decision, cache hit, fallback attempt, entitlement check, and payment event carries the same correlation_id. A compliance query for a specific agent session is a single index lookup, not a join across several log sources.


Policy Manifest

LiteLLM configuration is environment variables and a YAML config file with no versioning or content-addressed anchoring.

clawql-inference reads a policy manifest at $CLAWQL_HOME/Inference/policy.yaml that merges with environment variables at runtime (env wins on conflicts):

policyVersion: '2026.07.01'
inference:
  escalation:
    enabled: true
    tierMap:
      frugal: ollama/phi4
      standard: groq/llama-3.3-70b
      frontier: anthropic/claude-sonnet-4
  cache:
    enabled: true
    threshold: 0.92
    ttl: 24h
  fallback:
    enabled: true
    frugal: [ollama/phi4-backup, openai/gpt-4o-mini]
  pipelineWorker:
    enabled: true
    schedule: '0 2 * * 0'
    minSamples: 500
  observability:
    profile: external

resolveInferenceEffectiveEnv() is the same merge function used by both clawql inference policy show and clawql inference serve. What policy show displays is what the gateway runs.


Multi-Instance Deduplication

Running multiple inference serve replicas behind a load balancer creates a problem: the fine-tuning pipeline worker is a cron job that should run once per tick, not once per replica per tick.

When CLAWQL_INFERENCE_DATABASE_URL is configured, clawql-inference uses Postgres advisory locks keyed by pipeline schedule and UTC minute. One replica runs the tick. The others skip it without blocking. The lock releases when the tick completes. No external coordination service required.


Migration Path

Any client that respects OPENAI_BASE_URL works without code changes:

# Install
curl -fsSL https://clawql.com/install | bash

# Start the gateway
clawql inference serve --port 8080

# Point your existing client at it
export OPENAI_BASE_URL=http://127.0.0.1:8080/v1
export OPENAI_API_KEY=your-existing-key  # passed through to the provider

Provider configuration:

export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=sk-ant-...
export OLLAMA_BASE_URL=http://127.0.0.1:11434
export GROQ_API_KEY=gsk_...

Use provider/model format for explicit routing (anthropic/claude-sonnet-4, ollama/phi4) or bare model IDs (gpt-4o, claude-sonnet-4) when a matching provider is configured.

Start with observation only and enable features incrementally:

# Week 1: call store recording only
clawql inference serve

# Week 2: semantic cache
export CLAWQL_INFERENCE_SEMANTIC_CACHE=1

# Week 3: model escalation
export CLAWQL_INFERENCE_ROUTING_ENABLED=1

# Week 4: fallback chains
export CLAWQL_INFERENCE_FALLBACK_ENABLED=1
export CLAWQL_INFERENCE_FALLBACK_STANDARD=groq/llama-3.3-70b,anthropic/claude-haiku-4

# Month 2: first fine-tuning export
clawql inference export --verdict passed --format openai-jsonl --output ./export.jsonl

After a week of production traffic, the call store has enough data to show whether semantic caching and model escalation make sense for your workload before you commit to either.


Where LiteLLM Still Leads

Provider breadth. 100+ providers, years of maintenance work. clawql-inference ships OpenAI, Anthropic, and Ollama with a plugin API for others. Teams that need Azure OpenAI, Bedrock, Vertex, Fireworks, and Replicate out of the box will hit gaps.

Python ecosystem. LangChain, LlamaIndex, DSPy, and most ML tooling are Python-native. clawql-inference is TypeScript. The REST surface is compatible with any HTTP client, but Python SDK integration is smoother with LiteLLM.

Managed hosting. LiteLLM Enterprise has a hosted option. ClawQL’s managed tiers include clawql-inference but are a broader platform commitment than a standalone hosted inference gateway.

Community size. More Stack Overflow answers, more blog posts, more example integrations. The clawql-inference documentation is thorough but the community is smaller.


Comparison Table

Dimensionclawql-inferenceLiteLLM
LanguageTypeScript-nativePython
Supply chainCosign-signed, SBOM per build, Arweave manifest, startup hash verifyStandard PyPI; March 2026 compromise via dependency
Provider breadthOpenAI, Anthropic, Ollama built-in; plugin API100+ providers
RoutingOutcome-driven model escalationLoad balancing and cost rules
Semantic cacheEmbedding similarity, pgvector backendRedis exact/near-match
Fine-tuning flywheelVerdict-filtered export → PII scrub → job → tier-map registrationNone
WORM auditEvery routing decision, cache hit, escalation, payment event with correlation_idCallback-based logging
Payment railsStripe + x402 + MPP + ACP/AP2 plannedAPI key management
Policy manifestVersioned YAML, governs live gateway, merged with envEnv + config file
Multi-instance dedupPostgres advisory locks for pipeline workerExternal coordination required
Startup integrityclawql doctor --smoke verifies binary hashNone

clawql-inference is available as a standalone npm package (npx clawql-inference) and as part of the ClawQL platform. Full reference: docs.clawql.com/inference/clawql-inference. Provider plugin docs and migration guide: docs.clawql.com.

About the author

Daniel Smith builds ClawQL, an agent operating system for token-efficient discovery and execution over APIs — with observability, hardened tool boundaries, and production routing for LLM workloads. He writes here about the systems problems behind shipping agents.