openskills.info
Course Preview

Data Contracts

A data contract is a written, machine-checkable agreement between the team that produces a dataset and the teams that use it, stating the data's structure, meaning, quality, freshness, and owner. It exists so that a change in a source system cannot silently break the reports, pipelines, and models that depend on its data.

itData engineering and analytics

Don't Panic: Data Contracts

A data contract is a promise, written in a file that software can read, from the team that makes some data to the teams that use it. The promise says what the data looks like, what it means, how good it has to be, how fresh it will be, and who to call when it is not. That is the whole idea. The rest of the course is about making the promise stick.

It exists because most data in a warehouse was never offered, only taken. A pipeline copies a service's database tables every night, other teams build dashboards and machine learning models on the copies, and the service team has no idea any of this is happening. Then they rename a column. The dashboards quietly start showing nonsense, and everyone learns about the dependency at the same moment, which is the worst possible moment. One practitioner called this kind of database schema a non-consensual API: everyone depends on it, and nobody agreed to anything.

Three ideas carry most of the weight. First, the producer owns the contract, meaning the team that emits the data signs up for the guarantees. Second, the contract has a version, so consumers can pin to one and a breaking change arrives as a new version instead of a surprise. Third, and this is the one people skip, the contract is enforced. A tool compares the old and new contract before a change merges, or refuses to build a model whose columns drifted, or tests the stored data every night. A contract that nothing checks is a well-formatted wish.

The file format to know is the Open Data Contract Standard, or ODCS, a YAML format run by the Bitol project under the Linux Foundation AI and Data Foundation. It has sections for the schema, quality rules, service levels such as latency and retention, the team, access roles, and where the data physically lives. The open-source datacontract command-line tool reads it, tests real data against it, and fails a build when someone tries to delete a column that consumers rely on.

Now the surprise. Declaring something in a contract does not make it true. In most cloud warehouses, only "this column cannot be null" is actually enforced by the database; a declared primary key on Snowflake or BigQuery is a polite note that nobody reads on insert. So the contract says one thing, the data does another, and only a real test notices. The second surprise is less technical: the YAML is the easy part. Getting two teams with different priorities to agree on what "correct" means is where the actual work happens, and no tool does that for anyone.

Where to go next: the Intro explains how contracts are built and enforced, from the four enforcement points to compatibility modes on event streams. The Cheatsheet holds every ODCS field, quality metric, and command in one place. The Practice Reference and the Exercise walk through writing a contract, catching bad data with it, and watching it block a breaking change, all on local files. Field Notes covers what goes wrong once real teams are involved.

Where this skill leads

Relevant careers

See how this topic contributes to broader role-level skill maps.

Sources