Data Contracts
A data contract is a written, machine-checkable agreement between the team that produces a dataset and the teams that use it, stating the data's structure, meaning, quality, freshness, and owner. It exists so that a change in a source system cannot silently break the reports, pipelines, and models that depend on its data.
itData engineering and analytics | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic: Data Contracts
A data contract is a promise, written in a file that software can read, from the team that makes some data to the teams that use it. The promise says what the data looks like, what it means, how good it has to be, how fresh it will be, and who to call when it is not. That is the whole idea. The rest of the course is about making the promise stick.
It exists because most data in a warehouse was never offered, only taken. A pipeline copies a service's database tables every night, other teams build dashboards and machine learning models on the copies, and the service team has no idea any of this is happening. Then they rename a column. The dashboards quietly start showing nonsense, and everyone learns about the dependency at the same moment, which is the worst possible moment. One practitioner called this kind of database schema a non-consensual API: everyone depends on it, and nobody agreed to anything.
Three ideas carry most of the weight. First, the producer owns the contract, meaning the team that emits the data signs up for the guarantees. Second, the contract has a version, so consumers can pin to one and a breaking change arrives as a new version instead of a surprise. Third, and this is the one people skip, the contract is enforced. A tool compares the old and new contract before a change merges, or refuses to build a model whose columns drifted, or tests the stored data every night. A contract that nothing checks is a well-formatted wish.
The file format to know is the Open Data Contract Standard, or ODCS, a YAML format run by the Bitol project under the Linux Foundation AI and Data Foundation. It has sections for the schema, quality rules, service levels such as latency and retention, the team, access roles, and where the data physically lives. The open-source datacontract command-line tool reads it, tests real data against it, and fails a build when someone tries to delete a column that consumers rely on.
Now the surprise. Declaring something in a contract does not make it true. In most cloud warehouses, only "this column cannot be null" is actually enforced by the database; a declared primary key on Snowflake or BigQuery is a polite note that nobody reads on insert. So the contract says one thing, the data does another, and only a real test notices. The second surprise is less technical: the YAML is the easy part. Getting two teams with different priorities to agree on what "correct" means is where the actual work happens, and no tool does that for anyone.
Where to go next: the Intro explains how contracts are built and enforced, from the four enforcement points to compatibility modes on event streams. The Cheatsheet holds every ODCS field, quality metric, and command in one place. The Practice Reference and the Exercise walk through writing a contract, catching bad data with it, and watching it block a breaking change, all on local files. Field Notes covers what goes wrong once real teams are involved.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://github.com/bitol-io/open-data-contract-standard
Supports
- JSON Schema odcs-json-schema-latest.json requires only apiVersion, kind, id, and version
- ODCS current version v3.2.0, Apache 2.0 license, Bitol project
- A data contract defines the agreement between a data producer and consumers
- Section list (fundamentals, schema, references, data quality, support, pricing, team, roles, SLA, servers, authoritative definitions, tags, custom properties)
- JSON Schema for editor validation, examples, and the media type application/odcs+yaml;version=3.2.0
- Citation naming the LF AI & Data Foundation as organization
- https://bitol-io.github.io/open-data-contract-standard/latest/
Supports
- Reader-friendly documentation of every ODCS section
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/fundamentals.md
Supports
- apiVersion, kind, id, and version are required; kind is DataContract
- status examples proposed, draft, active, deprecated, retired
- description purpose, usage, limitations; dataProduct deprecated since v3.1.0
- version is the version of the contract, apiVersion the version of the standard
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/schema.md
Supports
- Objects, properties, and elements terminology in ODCS v3
- logicalType values string, date, timestamp, time, number, integer, object, array, boolean, map, vector
- physicalType, physicalName, required (nulls not allowed, default false), unique, primaryKey, partitioned, classification, criticalDataElement
- enum, semanticType column, measure, dimension, deprecated flag, transformSourceObjects, transformLogic, dataGranularityDescription
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/data-quality.md
Supports
- Quality rule types text, library (default), sql, custom
- Library metrics nullValues, missingValues, invalidValues, duplicateValues, rowCount and their arguments
- Operators mustBe, mustNotBe, mustBeGreaterThan, mustBeGreaterOrEqualTo, mustBeLessThan, mustBeLessOrEqualTo, mustBeBetween, mustNotBeBetween
- unit rows or percent; dimension values accuracy, completeness, conformity, consistency, coverage, timeliness, uniqueness
- severity, businessImpact, scheduler cron and schedule; custom engines soda and greatExpectations; {object} and {property} placeholders
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/service-level-agreement.md
Supports
- slaProperties property and value pairs, element, unit, driver regulatory, analytics, operational
- SLA properties latency (preferred to freshness), frequency, retention, timeOfAvailability, availability, throughput, errorRate, generalAvailability, endOfSupport, endOfLife, timeToDetect, timeToNotify, timeToRepair
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/team.md
Supports
- team members with username, role, dateIn, dateOut; formerly stakeholders
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/roles.md
Supports
- roles list with access and first and second level approvers
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/support-communication-channels.md
Supports
- support channels with channel, tool, scope, url
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/pricing.md
Supports
- price with priceAmount, priceCurrency, priceUnit
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/infrastructure-servers.md
Supports
- servers entries describe one dataset per environment and technology with server, type, environment
- https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/references.md
Supports
- References section for relationships such as foreign keys, new in v3.1.0
- https://github.com/bitol-io/open-data-contract-standard/blob/main/CHANGELOG.md
Supports
- v3.2.0 adds enum, map, vector, context, semanticType, synonyms, deprecated flag, and ${VAR_NAME} variables
- v3.0.0 on 2024-10-21 renamed uuid to id and quantumName to dataProduct, dropped username, password, server and other connection fields, added servers and support
- v2.2.0 on 2023-07-27 renamed the template to Open Data Contract Standard and finalized the fork under the AIDA User Group
- v2.1.1 on 2023-04-26 open source version; v2.0.0 renamed tables to dataset
- https://github.com/bitol-io/open-data-contract-standard/releases
Supports
- Version 3.2.0 published 2026-09-08 and version 3.1.0 on 2025-12-08
- https://bitol.io/odcs-version-3-transforming-data-contracts-with-enhanced-flexibility-integration-and-global-collaboration/
Supports
- ODCS v3 announced 2024-10-21 with schema support for hierarchical, streaming, and unstructured data
- Technical Steering Committee of 16 people from 14 companies
- https://lfaidata.foundation/blog/2023/11/30/bitol-joins-lf-ai-data-as-new-sandbox-project/
Supports
- Bitol joined LF AI & Data as a sandbox project on 2023-11-30, formerly the data contract template used at PayPal
- https://jgp.ai/2023/05/01/paypal-open-sources-its-data-contract-template/
Supports
- PayPal released its data contract template under Apache 2 on 2023-05-01, version 2.1.1 with eight sections
- https://jgp.ai/2023/08/09/welcome-to-the-open-data-contract-standard/
Supports
- AIDA User Group released version 2.2 of the Open Data Contract Standard in August 2023
- https://github.com/datacontract/datacontract-specification
Supports
- Data Contract Specification deprecated with the release of ODCS v3.1.0 to focus on a single standard
- Supported in Data Contract CLI and Entropy Data until the end of 2026
- Data contract as structure, format, semantics, quality, and terms of use; "Think of an API, but for data"; follows OpenAPI and AsyncAPI conventions
- https://github.com/datacontract/datacontract-specification/blob/main/CHANGELOG.md
Supports
- Version 0.9.0 on 2023-09-12 was the first public release
- https://github.com/datacontract/datacontract-cli
Supports
- Open-source MIT-licensed Python CLI that natively supports ODCS to lint, test, and export
- Install with uv tool install datacontract-cli[all]; Docker image datacontract/cli; Python 3.10 to 3.12
- Commands init, lint, changelog, breaking, test, import, export, dbt sync, dbt test; credentials via DATACONTRACT_POSTGRES_USERNAME environment variables
- Entropy Data as the commercial tool to manage data contracts
- https://cli.datacontract.com/
Supports
- Data Contract CLI product homepage
- https://docs.datacontract.com/
Supports
- Testing guides for more than twenty data platforms including warehouses, object storage, local files, and Kafka
- Imports from SQL DDL, dbt, Avro, BigQuery, Glue, Excel, and CSV; exports including dbt-models, avro, jsonschema, html, sodacl, great-expectations
- https://docs.datacontract.com/compare-contract-versions
Supports
- datacontract breaking severities ERROR, WARNING, INFO and exit status 1 on ERROR
- INFO covers metadata changes and relaxed constraints
- https://docs.datacontract.com/testing/local
Supports
- Testing local CSV, Parquet, JSON, and Delta files with the duckdb extra and no credentials; import csv
- https://docs.datacontract.com/scheduling
Supports
- Test contracts in CI on every change and on a recurring schedule to detect data drift
- https://pypi.org/project/datacontract-cli/
Supports
- Data Contract CLI v1.2.2 run locally on 2026-09-26 to confirm lint, test failure output and exit status, breaking ERROR for a removed property and INFO for an added optional property, and export sql --dialect postgres
- https://docs.getdbt.com/docs/mesh/govern/model-contracts
Supports
- contract enforced true requires name and data_type for every column
- Preflight check that returned column names and types match, agnostic to order
- Supported materializations table, view, incremental with on_schema_change append_new_columns or fail; unsupported Python, materialized views, ephemeral
- Breaking changes list (remove column, change data_type, remove or modify constraint, delete, rename, or disable contracted model)
- Recommended for public models relied on downstream; pair with model versions
- https://docs.getdbt.com/reference/resource-properties/constraints
Supports
- Per-platform constraint enforcement table for not_null, primary_key, foreign_key, unique, check
- Most analytical platforms enforce only not_null; add a unique data test when primary_key is not enforced
- https://docs.getdbt.com/docs/mesh/govern/model-versions
Supports
- Downstream references can pin a specific model version during a migration window
- https://github.com/dbt-labs/dbt-core/releases/tag/v1.5.0
Supports
- dbt Core v1.5.0 released 2023-04-27 with model governance (access, contracts, versions)
- https://docs.confluent.io/platform/current/schema-registry/fundamentals/schema-evolution.html
Supports
- Compatibility types BACKWARD, BACKWARD_TRANSITIVE, FORWARD, FORWARD_TRANSITIVE, FULL, FULL_TRANSITIVE, NONE and their allowed changes
- Default BACKWARD checks only the latest version; transitive modes check all previous versions
- Upgrade consumers first for BACKWARD, producers first for FORWARD, independently for FULL
- https://docs.confluent.io/platform/current/schema-registry/fundamentals/data-contracts.html
Supports
- Data contract as schema plus integrity constraints, metadata, rules, and evolution
- CONDITION and TRANSFORM rule kinds, CEL and CEL_FIELD, onFailure ERROR, DLQ, NONE, JSONata migration rules
- Rules available only on Confluent Enterprise and Confluent Cloud Stream Governance Advanced
- https://docs.confluent.io/platform/current/schema-registry/develop/api.html
Supports
- PUT /config/{subject} with application/vnd.schemaregistry.v1+json and a compatibility body
- https://github.com/confluentinc/schema-registry
Supports
- Schema Registry licensed under the Confluent Community License except client and Avro libraries
- https://shopify.engineering/capturing-every-change-shopify-sharded-monolith
Supports
- CDC tightly couples external event consumers to the internal data model of the source, spreading breaking changes (2021-03-12)
- https://andrew-jones.com/blog/data-contracts/
Supports
- Data contracts proposed on 2021-04-08 as producer-designed events with a strongly defined, documented, versioned schema treated like an API
- https://dataproducts.substack.com/p/the-rise-of-data-contracts
Supports
- Databases treated as non-consensual APIs by ELT and CDC tools (2022-08-22)
- Data contracts as API-like agreements between service engineers and data consumers
- https://www.gable.ai/blog/why-you-cant-seem-to-adopt-data-contracts-no-matter-how-hard-you-try
Supports
- Contracts enforce existing agreement between teams but cannot produce it; they cover known, agreed critical boundaries
- Teams measuring percent covered instead of provenance; contracts reserved for agreed criticals
- https://martinfowler.com/articles/data-monolith-to-mesh.html
Supports
- Data mesh article of 2019-05-20; each data product defines integrity as a set of SLOs and publishes schemas
- https://benn.substack.com/p/data-contracts
Supports
- Skeptical reassessment of data contracts and the negotiation between application and analytics engineers
- https://www.datafold.com/blog/the-best-data-contract-is-the-pull-request
Supports
- Breaking changes should surface at development time, like compile errors, rather than in production
- https://github.com/sindresorhus/awesome
Supports
- Discovery index linking the Awesome Data Engineering list
- https://github.com/igorbarinov/awesome-data-engineering
Supports
- Lists Great Expectations, DQOps, Provero, daffy, Apache Avro, and Protocol Buffers
- https://github.com/AltimateAI/awesome-data-contracts
Supports
- Curated articles, books, videos, podcasts, and tools on data contracts, including Avo and Data Caterer
- https://greatexpectations.io/
Supports
- Open-source data validation framework with expectations about data
- https://dqops.com/
Supports
- Open-source data quality platform from profiling to monitoring
- https://github.com/provero-org/provero
Supports
- Declarative YAML data quality engine with data contracts, contract validate, and contract diff for breaking changes
- https://daffy.readthedocs.io/
Supports
- DataFrame validation decorators for pandas, Polars, Modin, and PyArrow checking columns, dtypes, nullability, and ranges
- https://data.catering/
Supports
- Test data generation and validation with metadata sources including Data Contract CLI and ODCS
- https://www.avo.app/
Supports
- Event data quality platform with schema management, implementation tools, observability, and a free tier
- https://avro.apache.org/
Supports
- Apache Avro data serialization system
- https://protobuf.dev/
Supports
- Protocol Buffers language-neutral data interchange format
- https://www.entropy-data.com/
Supports
- Manages data products and data contracts, ODCS support, free start tier
- https://www.gable.ai/
Supports
- Implements data contracts, blocks risky diffs, and provides lineage read from code; demo-based sales
- https://www.getdbt.com/
Supports
- dbt product homepage
- https://www.confluent.io/product/confluent-platform/data-compatibility/
Supports
- Confluent Schema Registry product page
- https://docs.aws.amazon.com/glue/latest/dg/schema-registry.html
Supports
- Schema as a data format contract between producers and consumers; compatibility modes BACKWARD, BACKWARD_ALL, FORWARD, FORWARD_ALL, FULL, FULL_ALL, NONE, DISABLED; free to use
- https://www.apicur.io/registry/
Supports
- Apache 2.0 schema registry with compatibility rules
- https://buf.build/docs/breaking/
Supports
- buf breaking detects breaking changes in Protobuf schemas
- https://github.com/sodadata/soda-core
Supports
- Soda Core described as a data contracts engine; soda contract verify; Elastic License 2.0
- https://www.soda.io/pricing
Supports
- Free tier available
- https://docs.datahub.com/docs/managed-datahub/observe/data-contract
Supports
- Data contract as verifiable assertions on physical assets, producer-owned, one per asset, failing when promoted assertions fail
- https://buf.build/
Supports
- Buf product homepage for Protobuf tooling and the Buf Schema Registry
- https://www.soda.io/
Supports
- Soda product homepage
- https://datahubproject.io/
Supports
- DataHub open-source metadata platform homepage
- https://aws.amazon.com/glue/
Supports
- AWS Glue product homepage, which includes the Glue Schema Registry
- https://docs.astral.sh/uv/
Supports
- uv Python package and tool installer used to install the Data Contract CLI
