A configuration-driven data engineering framework for Databricks. Onboarding a new dataset means writing a workflow YAML — never Python.
Most Databricks pipelines grow one notebook per dataset. This framework inverts that: a single typed engine reads a workflow YAML and carries every dataset through the same five layers — Start, Pipeline, Typing, Policies, Output. A new source is a new YAML file, not a new code path.
The project is alpha — see Status for what's implemented today.
Two tasks, two YAML blocks — CSV into Bronze, then Bronze into Silver with SCD Type 1
upsert on CLAIM_ID:
# Inbound → Bronze
source.origin: csv
source.path: /Volumes/dev/source_data/inbound/
source.directory: CLAIMS
output.verb: append
output.schema_name: claims
output.table: CLAIMS_BRONZE
# Bronze → Silver
source.origin: table
source.schema_name: claims
source.table: CLAIMS_BRONZE
output.verb: upsert
output.schema_name: claims
output.table: CLAIMS_SILVER
output.keys: CLAIM_ID
output.event_time.column: __EXPORT_DATENo code changes for either step — both are entries in a workflow YAML deployed through the bundle. See 03_write_verbs.md for the full verb matrix.
The source, increment strategy (source.increment_strategy) and verb to use at each step.
Every other combination, and the traps a valid config doesn't rule out, are in
08_usage.md.
| Step | Source | Strategy | Verb | Why |
|---|---|---|---|---|
| Inbound → Bronze | Files (csv, json) | checkpoint |
APPEND | Bronze keeps every export exactly as it arrived |
| External table → Bronze | A table we didn't create (e.g. a federated one) | full_read |
APPEND | One complete snapshot per run. Sees deletes |
delta_read |
APPEND | Only the rows changed since the last run, for a table too large to read whole. Deletes are invisible | ||
| Bronze → Silver | A table we stamped | checkpoint |
FULL | The newest export is the current state |
checkpoint |
UPSERT | Latest version per key (SCD1) | ||
checkpoint · watermark |
SCD2 | History of changes | ||
watermark |
COMPLETE_DELTA, snapshot_scope: delta |
Replays every export in order; a deletes feed retires keys |
- Config-driven onboarding — declare an origin, a write verb and a target; the framework validates the combination and runs it. No per-dataset Python.
- Data handling verbs —
APPEND,FULL,UPSERT,SCD2,COMPLETE_DELTA— layer-agnostic, so the same verb serves Inbound→Bronze or Bronze→Silver. - Gold is SQL (planned) — each Gold table will be a materialized view over Silver, in its
own
.sqlfile, so Silver's MERGEs never block it. See Gold. - A data quality gate, not a bolt-on — every batch is checked against a
Databricks DQX ruleset (you define this) before it's written;
errorrefuses the batch,warnlogs and lets it through. - Fail fast, fail loud — Invalid Configs or Failed Data QA stops the run with one aggregated error report. Nothing logs an error and reports success.
- Ships as a wheel — src-layout package deployed via Databricks Declarative Automation
Bundles and run through
python_wheel_taskentry points. No notebook logic, nosys.pathhacks.
- 00_overview.md — goals, principles, glossary, layer diagram
- 01_architecture.md — layers, protocols, execution flow, extension points
- 02_config_schema.md — the typed config schema, parameter by parameter
- 03_write_verbs.md — verb semantics with worked examples
- 04_policies.md — the data quality gate, driven by a DQX ruleset
- 05_testing.md — unit and platform testing
- 06_roadmap.md — phased implementation plan and current status
- 07_workspace_setup.md — workspace setup recipe: catalog naming, schemas, volumes, grants
- 08_usage.md — which source, strategy and verb to combine, and the traps
Targets Linux; on Windows, work inside WSL, since Spark doesn't run natively there (Required for running unit tests).
Needs a JDK (11 or 17, for PySpark), Python 3.11 (matches Databricks Runtime 15.4 LTS) and Poetry 2.x.
poetry env use python3.11 # once, to pin the interpreter
make install # runtime + dev dependencies
make test # unit tests with coverageReview the Makefile to lists the rest.
Platform tests are real Databricks jobs under platform_tests/ - with real life scenarios; see 05_testing.md.
The bundle in databricks.yml is configured to work with a Python wheel: it
builds the wheel, then deploys the bundle. Authenticate with a CLI profile or with
DATABRICKS_HOST / DATABRICKS_TOKEN. A new workspace first needs its catalogs, schemas and
volumes; see 07_workspace_setup.md.
databricks bundle deployEvery run that passes config validation writes an audit trail to your own catalog
(monitoring_{env}.audit.logs), not to vendor telemetry. Each row names the job, run and task
that produced it. The audit write comes first, so a run that can't log fails before touching
any data. Schema: 00_overview.md.
Start, Pipeline (CSV + JSON + table), Typing, Policies (DQX), and all five write verbs are implemented and tested. The SAS origin and the Gold example view are not implemented yet. Full phase-by-phase status: 06_roadmap.md.
Apache 2.0 — see LICENSE.