Data Contracts engine for the modern data stack. https://www.soda.io
-
Updated
Oct 2, 2026 - Python
Data Contracts engine for the modern data stack. https://www.soda.io
Easy and flexible data contracts
A simple and easy to use Data Validation library for Python.
The DBT of ML, as Aligned describes data dependencies in ML systems, and reduce technical data debt
A kedro plugin to use pandera in your kedro projects
Decision provenance and data-flow tracking for coding agents: computes what changed in your code and data, records why it changed, and surfaces both before a silent failure ships.
Dex is the agent-native analytics engineering toolkit. Point it at your warehouse and your dbt project. It learns the landscape, authors your transformations, and tells you exactly what to fix when the schema drifts. Built for analytics engineers and data engineers who want more out of their coding agent.
A validation engine for Open Data Contract Standard (ODCS) data contracts. Define your validation rules once in a declarative YAML contract and get rich, actionable reports on your data's quality.
DCEE is a lightweight Python framework for validating data against contracts and enforcing SLA rules. Built on pandas and boto3, it provides simple, fast data validation without heavy dependencies.
Declarative data quality engine. Define checks in YAML, run anywhere.
Data Engineering infrastructure that controls inbound tabular data file feeds and validates data contracts
Your data, explored locally — and your AI agents kept on a leash. A federated data explorer with governed agentic access.
Open-source data contract enforcement — define, sync dbt, validate, block, report. Built on ODCS v3.1 + DuckDB.
Open-source, contract-driven data quality validation. Shift-left enforcement at the point of write — before data enters your pipeline.
YAML-first, domain-driven data governance for AI agents — teach agents your business domains, metrics, and rules before they write SQL
dqflow is a lightweight, contract-first data quality engine for modern data pipelines
LakeLogic Core — an open-source framework for executing OLC contracts across data platforms.
Data quality that just works. 3 lines of code, any data source, 10x faster. Snowflake, Databricks, Fabric, BigQuery, S3, Parquet & 16+ connectors.
Serializable, runtime-neutral contracts for market data, artifacts, strategy lifecycles, and execution across ML4T libraries.
CSV data-quality gates for GitHub Actions and ML pipelines. 46 contract checks, downloadable failure reports, a Python CLI and optional FastAPI.
To associate your repository with the data-contracts topic, visit your repo's landing page and select "manage topics."