How typed entity schemas, permission-aware relationship graphs, and kinetic MCP writes transform AI agents from JSON-blob processors into typed, auditable business intelligence — with ClawQL’s .cqe format, fixture-backed reads, LOW/MEDIUM kinetic tools, and an honest map of what is shipped vs roadmap.
The Problem Nobody Names Correctly
When enterprise AI projects fail — and most of them do, eventually — the post-mortem usually blames the model. Wrong model choice. Insufficient context. Hallucinations. Prompt engineering failures.
These diagnoses are almost always wrong. The real problem is earlier in the stack.
Consider what happens when you give an AI agent access to a company’s CRM, ERP, and document management system. The agent receives tool call responses that look like this:
{
"id": "acc-8821",
"cust_id": "C-4421",
"status": 2,
"type": "ENT",
"owner": "jsmith",
"created": 1720000000,
"val": 48500.0,
"currency": "USD"
}
What is status: 2? What does type: "ENT" mean? Is owner: "jsmith" a user ID, an email, a display name? What does val represent — annual contract value, monthly, one-time? Is cust_id the same as id in the customer system or a different identifier?
The agent doesn’t know. Neither does the developer who wrote the integration. The answers are tribal knowledge, scattered across wiki pages, Slack threads, and the institutional memory of the person who built the CRM integration in 2019.
When the agent calls update_account with status: 3, it doesn’t know what status 3 means. When it queries accounts with type: "SMB", it doesn’t know if "SMB" is a valid value. When it joins account data with customer data using cust_id, it doesn’t know if that foreign key is always populated or frequently null.
This is the enterprise data problem that has existed since the first integration between two enterprise systems. Semantic fragmentation — the same real-world concept represented differently in every system, with no shared vocabulary to resolve the differences.
Traditional enterprise systems were built to store and retrieve records efficiently. They were not designed to preserve meaning across system boundaries. When an organization adds more applications, data pipelines, and AI systems, the cost of reconstructing meaning at every integration point compounds. Analysts reconcile terminology manually. Pipelines fail when field names change. AI models trained on one system’s data fail to generalize across others because the underlying concepts were never aligned.
AI makes this problem worse in two ways. First, AI agents operate at scale — one agent making thousands of tool calls compounds the semantic ambiguity at machine speed. Second, AI agents make decisions based on their interpretation of the data — and wrong interpretations at machine speed cause real-world damage.
The Enterprise Ontology is the solution to this problem. Not a new idea — semantic web researchers were working on this in 2001. But the implementation approach has changed fundamentally now that we have AI agents that can reason about typed schemas, generate typed queries, and respect typed constraints without constant human oversight.
What an Enterprise Ontology Actually Is
An Ontology, in the sense we’re using it, is a formally defined vocabulary of concepts in a domain — their types, their properties, their relationships, and the constraints that govern them. It is not a database schema. It is not an API contract. It is the semantic layer that sits above both, giving names and meaning to concepts that the database stores as integers and the API serializes as JSON.
The clearest mental model: OOP taken to its logical extreme.
Object-Oriented Programming gave us the insight that modeling the world as typed objects with properties, methods, and relationships produces more maintainable and comprehensible systems than modeling it as raw data and procedures. The Enterprise Ontology takes that insight and applies it to the entire enterprise information space — not just the software, but the real-world entities the software represents.
OOP: class Contract { parties: Party[], status: ContractStatus, sign(): void }
Database: CREATE TABLE contracts (id, status, effective_date, value...)
Ontology: Contract is a typed object with:
- Properties sourced from the database
- Relationships traversable to Organization and Individual
- Read methods generated as typed MCP tools
- Write methods generated as kinetic actions requiring authorization
- PII fields declared for automatic redaction before LLM exposure
- Audit level specified for every property change
Three properties make the Enterprise Ontology different from standard OOP in ways that matter for enterprise AI:
Provenance. In OOP, a Contract object is constructed at runtime from whatever data you hand to the constructor. In an Enterprise Ontology, a Contract object has a verifiable chain from the raw data source through the indexing pipeline to the typed object the agent sees. Every property value traces back to the specific database row or document page it came from. That chain is in the WORM audit log.
Permission-awareness. In OOP, method access control is typically application-level — you check permissions in the method body. In an Enterprise Ontology, permission-awareness is structural. The agent cannot see entities it doesn’t have authorization for. The graph traversal stops at authorization boundaries. The object doesn’t exist in the agent’s world if the agent isn’t authorized to see it — not an access denied error, structural invisibility.
Kinetic designation. In OOP, a method that writes to an external system looks the same at the call site as a method that reads. In an Enterprise Ontology, write methods are structurally different from read methods — they carry a kinetic designation, they require authorization before executing, they enter a Transaction Sandbox, and they generate a specific audit entry type that’s distinct from read operations. The type system encodes the safety boundary.
Why Now: Microsoft Fabric Just Validated the Entire Approach
Before diving into implementation, it’s worth noting that in mid-2026, Microsoft Fabric shipped their Ontology feature as part of Fabric IQ. The Fabric Ontology defines enterprise concepts as entity types (like Customer), properties (like a Customer’s name and email), and relationships (like Customer places Order), while clarifying the constraints of these terms. After defining the ontology, organizations bind the entity type definitions to real data so downstream tools can share the same language. Both humans and AI agents can use this language for cross-domain reasoning and decision-ready actions.
This is exactly the architecture described in this post — shipped by Microsoft as a first-class product feature. The validation matters not because Microsoft’s implementation is the right one (it isn’t — it has the same proprietary lock-in problem as Palantir’s Ontology), but because the largest enterprise software company in the world just confirmed that typed entity schemas for AI agents is the right architectural approach, not an academic exercise.
The difference between ClawQL’s Ontology and Microsoft Fabric’s: ClawQL’s is open format, version-controlled in Git, manifest-governed, independently auditable, and not tied to any cloud vendor’s data platform. Your Ontology schema is a .cqe file in your Git repository. If you want to move it, you git clone it.
The Three-Layer Architecture
The Enterprise Ontology has three distinct layers that build on each other. Each is independently valuable. Each feeds the next.
Layer 1: The Entity Schema
The Entity Schema defines what things are — their types, properties, data sources, PII declarations, and audit requirements. This is the .cqe file format.
# ontology/entities/Contract.cqe
---
type: ontology_entity
title: Contract
description: 'Legal agreement between two or more parties, establishing rights and obligations'
version: '1.2.0'
domain: legal
pii_classification: CONFIDENTIAL
worm_ref: sha256:a1b2c3d4... # links to WORM entry recording this schema version
---
properties:
contract_id:
type: string
required: true
indexed: true
immutable: true # cannot be changed after creation
description: 'Unique contract identifier'
title:
type: string
required: true
description: 'Human-readable contract name'
status:
type: enum
values: [draft, active, suspended, expired, terminated]
required: true
default: draft
description: 'Current lifecycle status'
audit_on_change: true # every status change generates WORM entry
effective_date:
type: date
required: true
expiry_date:
type: date
required: false
validation: 'expiry_date > effective_date'
value:
type: money
required: false
pii: false
kinetic_change_limit: '100000.00' # changes above this require AP2 mandate
parties:
type: array
items: { $ref: '#/entities/Organization' }
immutable: true # no mutation generated — party changes = new contract
primary_contact:
type: { $ref: '#/entities/Individual' }
pii: true # redacted before LLM exposure
required: false
document_refs:
type: array
items: { type: string } # Onyx document IDs
description: 'Linked documents in IDP pipeline'
sources:
- type: sql
connection: ${VAULT:contracts_db_ro}
table: contracts
id_column: contract_id
column_map:
cust_id: parties
status: status # int → enum mapping
val: value
- type: nextcloud
path: /company/legal/contracts/
classifier: classify_document
link_by: contract_id # links documents to entities by ID
audit:
level: STANDARD
pii_fields: [primary_contact, parties.contact_email]
change_log: true
Several things in this schema are doing real work beyond documentation:
The pii: true declarations drive Presidio redaction automatically — any LLM call that includes primary_contact data has those fields stripped before the prompt is assembled. The agent never sees raw PII unless its ATRClaims explicitly authorize it.
The immutable: true on parties means no mutation is generated for that field. An agent that tries to modify the parties on an existing Contract doesn’t get an access denied error — it gets a schema validation error explaining that party changes require a new Contract entity. The business rule is in the schema, not scattered across application code.
The sources block is what makes the Ontology live rather than theoretical. clawql ontology generate reads this block and configures Onyx to index Contract entities from both the SQL database and Nextcloud. The agent doesn’t know whether a Contract was sourced from the database or a document — it sees a typed Contract object either way.
The worm_ref in the frontmatter links this schema version to the WORM entry that recorded its publication. Every schema change is auditable — a regulator asking “what was your Contract entity definition on March 1, 2026?” gets a deterministic answer from the manifest history and the WORM log.
Layer 2: The Relationship Graph
Relationships between entities form a traversable graph. This is where the Enterprise Ontology becomes more powerful than a collection of database schemas — relationships create a navigable knowledge graph that agents can traverse with typed, permission-aware queries.
# ontology/relationships/contract_parties.cqe
---
type: ontology_relationship
title: ContractParties
description: 'Links a Contract to the Organizations that are party to it'
version: '1.0.0'
worm_ref: sha256:b2c3d4e5...
---
from: Contract
to: Organization
cardinality: many_to_many
via_table: contract_parties
via_columns:
from: contract_id
to: organization_id
traversal:
forward:
name: parties
description: 'Organizations that are party to this contract'
permission: READ_CONTRACT
reverse:
name: contracts
description: 'Contracts this organization is party to'
permission: READ_ORGANIZATION
constraints:
minimum: 2 # a contract requires at least 2 parties
parties_must_be_active: true # party Organizations must have status: active
The permission fields on forward and reverse traversal mean graph queries respect authorization boundaries automatically. An agent traversing from Contract → parties → Organization will only see Organizations that the agent’s ATRClaims authorize it to read. The graph is permission-aware at the traversal level, not just at the entity level.
clawql ontology generate reads relationship definitions and generates typed graph traversal MCP tools:
// Generated MCP tools from relationship definitions
search_contracts(query: string, status?: ContractStatus): Contract[]
get_contract(contract_id: string): Contract
get_contract_parties(contract_id: string): Organization[]
get_organization_contracts(organization_id: string, status?: ContractStatus): Contract[]
list_contracts_expiring(days: number, party_id?: string): Contract[]
The agent calls these typed tools instead of constructing raw SQL or API calls. The semantic ambiguity from the opening example — what does status: 2 mean? — is gone. The agent calls get_contract("acc-8821") and receives:
{
contract_id: "acc-8821",
title: "Enterprise Software License Agreement — Acme Corp",
status: "active", // typed enum, not integer 2
effective_date: "2026-01-15",
expiry_date: "2027-01-14",
value: { amount: 48500.00, currency: "USD" },
parties: [
{ organization_id: "org-4421", name: "Acme Corporation", type: "customer" }
]
// primary_contact: REDACTED — PII field, agent lacks PII_READ ATRClaim
}
Typed. Semantic. Clean. The agent knows what every field means because the schema defines it. The agent can’t accidentally use the wrong status value because the enum is enforced at the tool layer.
Layer 3: The Action Schema — Reads and Writes at the Right Level
This is where the Ontology closes the loop from knowledge representation to safe autonomous execution.
Read actions generate automatically from entity and relationship definitions — get_contract, search_contracts, get_contract_parties. No schema author needs to define them. They’re derived from the Entity Schema and the Relationship Graph.
Write actions — mutations — require explicit definition. And here is where the GraphQL mutation / @kinetic directive architecture delivers its value.
Why GraphQL mutations are the right transport layer:
GraphQL’s query/mutation distinction is a semantic contract at the schema level. Mutations are expected to have side effects. Queries are expected to be safe and idempotent. This distinction is structural, not just conventional — GraphQL clients and servers treat them differently. An agent can introspect the schema and discover which operations are mutations before calling them.
Where GraphQL mutations alone fall short:
The mutation/query distinction is binary. It tells you “this has side effects” but not what kind. updateContractStatus and initiatePayment are both mutations. They are not the same kind of mutation. One updates a field in a database. The other initiates a financial transaction with real-world consequences. GraphQL treats them identically at the schema level.
There’s also no native authorization semantics, no rollback protocol, no blast radius constraint, and no AP2 mandate binding in standard GraphQL. These require a governance layer on top.
The @kinetic directive:
The @kinetic directive adds graded governance semantics to GraphQL mutations without changing how the agent calls them. The call site looks identical for low-risk and high-risk mutations. The PEP intercepts based on the directive and applies the appropriate validation, staging, and audit.
type Mutation {
# Low-risk — field update, single record, immediate execution
updateContractStatus(contractId: ID!, status: ContractStatus!, reason: String): Contract
@kinetic(
riskLevel: LOW
requiresMandate: false
blastRadius: SINGLE_RECORD
rollbackProtocol: FIELD_RESTORE
auditLevel: STANDARD
)
# Medium-risk — value change, requires AP2 mandate above threshold
updateContractValue(contractId: ID!, newValue: MoneyInput!, approvalRef: String): Contract
@kinetic(
riskLevel: MEDIUM
requiresMandate: true
mandateType: AP2_FINANCIAL
changeLimit: "100000.00"
blastRadius: SINGLE_RECORD
rollbackProtocol: FIELD_RESTORE
auditLevel: DETAILED
)
# High-risk — bulk operation, canary deployment, human approval
bulkRenewContracts(contractIds: [ID!]!, extensionDays: Int!): BulkRenewalResult
@kinetic(
riskLevel: HIGH
requiresMandate: true
mandateType: AP2_BULK_OPERATION
blastRadius: MULTI_RECORD
canary: { initialBatch: 5, progression: [5, 25, 100, ALL], analysisInterval: "30s" }
rollbackProtocol: SNAPSHOT_RESTORE
auditLevel: FORENSIC
requiresHumanInLoop: true
)
# Critical — financial transaction, full AP2 mandate chain, reversal window
initiateContractPayment(contractId: ID!, amount: MoneyInput!): PaymentResult
@kinetic(
riskLevel: CRITICAL
requiresMandate: true
mandateType: AP2_PAYMENT
transactionLimit: "50000.00"
blastRadius: FINANCIAL
rollbackProtocol: REVERSAL
reversalWindow: "24h"
auditLevel: FORENSIC
requiresHumanInLoop: true
)
}
Each risk level maps to a different execution path in the Transaction Sandbox:
LOW: PEP validates ATRClaims → field snapshot for rollback → immediate execution → WORM entry (STANDARD).
MEDIUM: PEP validates ATRClaims → AP2 mandate verified → field snapshot → execution (or staged for human review if above changeLimit) → WORM entry (DETAILED).
HIGH: PEP validates ATRClaims → AP2 mandate verified → canary execution (5 records first, analysis gate, then 25, then 100, then ALL) → WORM entry per canary stage (FORENSIC).
CRITICAL: PEP validates ATRClaims → AP2 Payment Mandate verified → Transaction Sandbox stages action → mandatory human review in Command Deck → atomic execution with reversal hook → WORM entry (FORENSIC).
The Argo Rollouts / Pulumi Parallel
When reading the bulkRenewContracts canary configuration above, engineers familiar with Argo Rollouts will recognize the pattern immediately. Progressive delivery — 5% → 25% → 100% with automated analysis gates — is exactly what Argo Rollouts applies to deployments. The @kinetic canary configuration applies the same pattern to agent actions.
This is not a coincidence. Both problems are structurally identical: how do you make a stateful change to a system with real-world consequences, in a way that is atomic, recoverable, and observable?
Argo Rollouts solved this for deployments. Pulumi solved this for infrastructure changes — state file tracking exact resource state, plan-before-execute, partial rollback based on what actually succeeded rather than what was planned. ClawQL’s kinetic execution layer applies both solutions to agent actions.
The implementation is direct: executor: ARGO_WORKFLOW in the @kinetic directive routes the action through Argo Workflows for multi-step execution with DAG semantics, step-level retry, artifact passing between steps, and durable suspension for human-in-loop approval. executor: PULUMI routes infrastructure-changing actions through the Pulumi Automation API, getting state management and partial rollback for free.
For application-level mutations (database updates, CRM writes, payment initiations), the native Transaction Sandbox implements the same conceptual model — explicit state, plan-before-execute, surgical partial rollback, durable suspension.
The IDP pipeline is already a proof of concept: document processing is an Argo Workflow (Tika → Steganography Detection → Gotenberg → Stirling-PDF → Archive → Onyx). Every step has typed inputs and outputs. Step failures retry with backoff. The processed document passes between steps as an Argo artifact. The agent drives this via the Argo API wrapped by the Transaction Sandbox with WORM audit.
The File Format: .cqe, .cqm, .cqw, .cqk
ClawQL defines four open file formats for the Ontology layer. Each is published as an Apache 2.0 specification at docs.clawql.com. Third-party tools can implement support for any of them without a license.
.cqe — ClawQL Entity. Ontology entity definitions. OKF-compatible Markdown + YAML frontmatter with type: ontology_entity enforced and ClawQL-specific frontmatter fields (sources, pii_fields, audit level, worm_ref) required. clawql ontology generate processes .cqe files to generate typed MCP tools and configure Onyx indexing.
.cqm — ClawQL Manifest. The EnterpriseGovernance manifest, release manifest, and policy manifest. YAML-based internally with a ClawQL-specific schema that clawql-release lint validates. The extension signals to CI pipelines, editors, and the VS Code extension to apply ClawQL-specific validation.
.cqw — ClawQL Workflow. Argo Workflow templates enriched with ClawQL-specific annotations — @kinetic directive parameters, WORM correlation ID injection, blast radius declarations. Syntactically standard Argo YAML but semantically ClawQL-aware.
.cqk — ClawQL Knowledge. OKF-compatible knowledge bundle entries produced by ClawQL processes — memory entries, task results, runbooks with execution history. Distinguished from generic .md OKF files by the required worm_ref field that ties every knowledge entry to its WORM audit record.
The file extension is a signal to the toolchain, not a branding exercise. The VS Code extension validates .cqe files against the Entity Schema spec in real time. clawql-release lint validates .cqm files against the governance spec. clawql ontology generate processes .cqe files to generate tools. clawql doctor --smoke verifies .cqm manifest hashes on startup. Each extension earns its existence by being meaningfully different from existing formats in ways that enable different tooling behavior.
The OKF Integration: Memory and Ontology as One Vault
The memory vault and the Ontology schema are not separate systems. They are the same OKF-compatible directory structure with different type values in the frontmatter.
~/.ClawQL/
ontology/
index.md # OKF catalog of all entity types
log.md # changelog of schema changes
entities/
Contract.cqe
Patient.cqe
Organization.cqe
relationships/
contract_parties.cqe
actions/
initiate_payment.cqw
memory/
decisions/
arch-001.cqk # type: decision
context/
project-abc.cqk # type: context
task_results/
weekly-summary-2026-07-14.cqk # type: task_result
entities/ # populated entity instances
contracts/
contract-abc123.cqk # type: entity_instance, entity_type: Contract
clawql ontology generate reads ontology/entities/*.cqe and generates:
- Typed MCP tools registered in the gateway
- Onyx indexing configuration for each entity type’s sources
- The
ontology/index.mdOKF catalog
memory_ingest writes to memory/**/*.cqk with OKF-compatible frontmatter including worm_ref.
memory_recall reads ontology/index.md first to identify relevant entity types and memory categories, then loads specific .cqe and .cqk files. The index provides a structured catalog that lets the agent survey what’s available before loading full content — a meaningful token efficiency gain over loading everything or relying entirely on vector search.
The memory vault is stored in S3/R2, not Git. The schema (.cqe, .cqm, .cqw files) lives in Git — diffable, PR-reviewable, signed, Arweave-permanent on release. The data (.cqk files, populated entity instances, generated indexes) lives in R2 — grows without bound, synced to edge nodes via clawql sync on a hot/cold tiering model. Schema and data have different lifecycles and different homes. Git for definitions. R2 for instances.
Token Efficiency: How the Ontology Closes the Loop
The Ontology is not a separate concern from token efficiency. It is the architecture that makes the deeper token efficiency layers possible.
Layer 1 (Code Mode) works because of the Ontology. The search + execute two-tool pattern assumes there’s a meaningful search space — a catalog of typed operations with semantic descriptions that the agent can search against. The Ontology provides that catalog. Without typed entity definitions, search() returns raw API endpoints. With typed entity definitions, search() returns semantically meaningful operations like “renew a contract” and “flag a contract for legal review” that the agent can reason about.
Layer 2 (Response Trimming) is smarter with Ontology. GraphQL projection — trimming API responses to only the fields the agent actually depends on — becomes more precise when the agent is querying typed entity properties rather than arbitrary JSON paths. The agent asks for Contract.status and Contract.expiry_date explicitly. The trimmer knows exactly which fields to preserve.
Layer 4 (Cache Control) benefits from schema stability. The system prefix that Anthropic caches includes the tool definitions — the MCP tools generated from the Ontology schema. Because the Ontology schema is versioned and changes through explicit manifest version events rather than implicit runtime updates, the tool definitions are stable across sessions. The cache hit rate is higher because the stable prefix is longer.
Layer 5 (Semantic Cache) is more effective with typed queries. “Show me contracts expiring in 90 days” and “list contracts that expire within three months” are semantically identical queries over the Contract entity. The semantic cache recognizes this because both queries are typed — they query the same entity type with the same relationship traversal. The embedding similarity is higher for typed entity queries than for arbitrary natural language queries over JSON blobs.
Layer 8 (PAL Routing) escalates on schema violations. When the Model Escalation router sees the agent attempting to access an unauthorized entity type, or trying to invoke a kinetic action without the required mandate, it treats this as a failure signal — the current tier didn’t have sufficient context to operate within the schema constraints. It escalates to a more capable tier that can reason more carefully about the authorization requirements.
Layer 12 (Fine-Tuning Flywheel) produces better training data with Ontology. Fine-tuning examples that involve typed entity queries — get_contract("acc-8821") returning a structured Contract object — are more generalizable than examples involving arbitrary JSON manipulation. The fine-tuned Frugal tier learns the entity schema and the query patterns, not just the surface-level string manipulation. Each Flywheel cycle makes the Frugal tier more accurate specifically on the organization’s entity types.
The compound effect: a typed entity query that goes through all 12 layers — Code Mode search over typed entity catalog, response trimmed to declared properties, system prefix cached because schema is stable, semantically cached because query type is recognized, PAL-routed to Frugal tier because it’s a routine read, fine-tuned model used because domain is recognized — costs a small fraction of the naive equivalent and produces higher-quality output because the model is operating on typed, semantically clean data rather than ambiguous JSON blobs.
The Vertical Schema Packs
Not every organization can construct their Ontology from scratch. The time and expertise required is one of the primary complaints about Palantir’s implementation — their Ontology requires significant professional services engagement to build.
ClawQL ships vertical schema packs — pre-built .cqe bundles for specific industries that organizations adopt and customize rather than constructing from first principles.
Legal/M&A:
Contract, Matter, Party, Clause, Jurisdiction, Court, Filing
Relationships: matter_parties, contract_clauses, filing_court
Actions: execute_contract (kinetic: CRITICAL), file_motion (kinetic: HIGH)
PII declarations: party.personal_details, filing.personal_information
Healthcare:
Patient, Provider, Encounter, Prescription, Diagnosis, Facility
Relationships: patient_encounters, encounter_diagnoses, prescription_providers
Actions: update_medication (kinetic: HIGH, requiresMandate: AP2_CLINICAL)
PII declarations: patient.* (HIPAA covered — all fields marked pii: true by default)
Audit level: FORENSIC on all Patient entity modifications
Financial Services:
Account, Transaction, Customer, Instrument, Portfolio, Position
Relationships: account_transactions, portfolio_positions, customer_accounts
Actions: initiate_transfer (kinetic: CRITICAL, transactionLimit: configurable)
PII declarations: customer.personal_details, customer.identity_documents
Compliance: SOC2_TYPE2 manifest block required
Real Estate:
Property, Listing, Transaction, Party, Document, Inspection
Relationships: listing_property, transaction_parties, property_documents
Actions: accept_offer (kinetic: HIGH), close_transaction (kinetic: CRITICAL)
IDP integration: Document entity links to IDP pipeline for contract processing
Each pack is an Apache 2.0 OKF bundle. Organizations import it, review the entity definitions, customize property names and source mappings to match their specific systems, and have a working Ontology in a sprint rather than a quarter.
The customization workflow is intentionally self-service: clawql ontology import --pack legal adds the legal schema pack to the local .clawql/ontology/ directory. The schema builder in the Command Deck provides a visual editor for customization. clawql ontology lint validates the customized schema. clawql ontology generate produces the typed tools. The engineer reviews the generated tool definitions in a PR before they’re published.
The Palantir Comparison: Why Open Format Wins
Palantir’s Ontology is the most mature implementation of the typed enterprise entity concept in production today. It is genuinely powerful. Ontology provides a scaled, secure, and governed shared business model used across teams, agents, and workflows — Palantir’s version of this works and has worked for 20 years in some of the most demanding enterprise environments on earth.
The problem is structural. Palantir’s Ontology format is proprietary. It lives in Palantir’s systems. It requires Palantir’s tooling to read, write, or inspect. It cannot be versioned in Git alongside your code. It cannot be audited independently by a regulator who doesn’t have Palantir access. It cannot be exported to a different system without losing the semantic layer entirely.
Every organization that builds its Ontology on Palantir is building its organizational intelligence on a foundation it doesn’t own.
ClawQL’s approach: the Ontology schema is a directory of .cqe files in your Git repository. It is versioned with your code. It is reviewed in pull requests by your engineers. It is signed alongside your releases. It is Arweave-permanent — the schema that was active on any given date is retrievable forever from the Layer 0 release manifest. A regulator asking “what was your Contract entity definition on March 1?” gets a deterministic, independently verifiable answer from the WORM log and the Arweave manifest. No Palantir required. No ClawQL required. Just cryptographic proof.
Both humans and AI agents can use this language for cross-domain reasoning and decision-ready actions — this is the right outcome. ClawQL’s implementation gets there without the lock-in. Your schema, in your Git repository, in an open format, independently auditable.
The Ontology Builder
The schema format and CLI tooling come first. Before any UI, the stable YAML format for entity types, relationships, and action definitions needs to exist and be validated against real enterprise data models. clawql ontology lint validates schemas in CI. clawql ontology generate produces typed tools and Onyx configuration from schemas.
The visual schema builder — the Command Deck UI for non-engineers — comes second. Three panels:
Entity panel: Define entity types, their properties, and property metadata. Drag-and-drop property ordering. Inline documentation that becomes the OKF description. PII field marking with automatic redaction implications displayed. Source mapping to SQL tables, OpenAPI schemas, document classifiers.
Relationship panel: Draw relationships between entity types visually. The builder generates the relationship .cqe file and the graph traversal MCP tools automatically. Cardinality constraints and traversal permissions configurable through the UI.
Action panel: Define kinetic actions. Risk level selection — LOW, MEDIUM, HIGH, CRITICAL — with the governance implications of each displayed explicitly. Canary configuration for HIGH/CRITICAL actions. AP2 mandate type selection. Rollback protocol selection. The UI makes kinetic designation non-accidental — you cannot accidentally create a write action without a risk level, and you cannot accidentally omit the rollback protocol.
The builder generates .cqe files, commits them to a branch, and opens a PR for engineer review. The CISO authors the entity definitions. The engineer reviews the generated schema. Security is a PR review, not a post-deployment audit.
Getting Started in One Sprint
The path from “no Ontology” to “working typed entity schema with generated MCP tools” is a single sprint for an organization with one or two data sources to model:
# Day 1: Initialize
clawql ontology init
# Creates .clawql/ontology/ directory structure
# Day 1-2: Define your first entity
clawql ontology create-entity Contract
# Opens Contract.cqe in $EDITOR with template
# Fill in properties, sources, PII declarations
# Day 3: Validate and generate
clawql ontology lint
clawql ontology generate --dry-run
# Shows what MCP tools will be generated, what Onyx config will be created
# No changes made until you confirm
# Day 4: Test
clawql ontology generate
clawql inference serve --port 8080
# MCP tools now available: search_contracts, get_contract, list_contracts_expiring
# Point Cursor at localhost:8080/mcp and ask: "show me contracts expiring in 90 days"
# Day 5: Publish
clawql-release lint # validates manifest includes ontology schema version
clawql-release publish --tag v1.0.0
# Ontology schema version recorded in Arweave-permanent manifest
# Every schema change from now on is a manifest version event
By the end of the sprint: typed entity queries working in Cursor and Claude Code via MCP. Onyx indexing your first entity type from its data source. Schema versioned in Git. First schema version Arweave-permanent. The foundation for everything else — relationship graph, kinetic actions, vertical pack customization, visual schema builder — built on a stable format.
Conclusion: The Semantic Layer Your Agents Have Been Missing
The enterprise data problem is 30 years old. The same real-world entity is a “customer” in the CRM, a “counterparty” in the fraud system, and a “subject” in the compliance system. Every integration reconstructs meaning manually. Every AI agent ingests JSON blobs and infers semantics from context.
The Enterprise Ontology is the solution that the semantic web community proposed in 2001, that Palantir implemented in a proprietary form for defense and intelligence customers, that Microsoft just shipped in Fabric IQ, and that ClawQL implements in an open, version-controlled, independently auditable form that any organization can own.
The key insight — OOP taken to its logical extreme — is that modeling enterprise information as typed objects with provenance-tracked properties, permission-aware relationship traversal, and kinetic-designated write methods produces agents that are fundamentally more capable and fundamentally more trustworthy than agents operating on raw JSON. The semantic layer is not an academic exercise. It is the difference between an agent that infers meaning from context and an agent that knows meaning from schema.
Three properties make ClawQL’s implementation the right one for enterprise AI:
Open and portable. Your Ontology schema is .cqe files in your Git repository. If you want to move it, you git clone it. Apache 2.0 specification. No vendor lock-in.
Versioned and auditable. Every schema change is a manifest version event — signed, Arweave-permanent, WORM-logged. A regulator asking “what was your Patient entity definition on any given date?” gets a deterministic, independently verifiable answer.
Kinetically safe. Write operations are structurally different from read operations through the @kinetic directive on GraphQL mutations. Blast radius constraints, AP2 mandate verification, Argo-Rollouts-style canary execution, and Pulumi-style partial rollback are defined in the schema before any agent action executes. Autonomous agents that touch production systems operate within a typed safety envelope — not because you prompted them carefully, but because the schema enforces it.
The .cqe format specification is at docs.clawql.com/architecture/enterprise-ontology. The VS Code extension validates schemas in real time. The vertical packs give you a working Ontology in a sprint.
Your agents have been operating on JSON blobs. Give them types.
ClawQL’s Enterprise Ontology, including the .cqe/.cqm/.cqw/.cqk format specifications, vertical schema packs, and the Ontology Builder, is documented at docs.clawql.com/architecture/enterprise-ontology. The token efficiency architecture that the Ontology enables is at docs.clawql.com/architecture/token-efficiency.
