Skip to content
EscotoPublic

About

Configuration-driven data engineering framework for Databricks. Onboard datasets in YAML, not Python.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

dbx-flame

A configuration-driven data engineering framework for Databricks. Onboarding a new dataset means writing a workflow YAML — never Python.

PyPI Status CI python Databricks Runtime DQX License

Why

Most Databricks pipelines grow one notebook per dataset. This framework inverts that: a single typed engine reads a workflow YAML and carries every dataset through the same five layers — Start, Pipeline, Typing, Policies, Output. A new source is a new YAML file, not a new code path.

The project is alpha — see Status for what's implemented today.

Quick Look

Two tasks, two YAML blocks — CSV into Bronze, then Bronze into Silver with SCD Type 1 upsert on CLAIM_ID:

# Inbound → Bronze
source.origin: csv
source.path: /Volumes/dev/source_data/inbound/
source.directory: CLAIMS
output.verb: append
output.schema_name: claims
output.table: CLAIMS_BRONZE

# Bronze → Silver
source.origin: table
source.schema_name: claims
source.table: CLAIMS_BRONZE
output.verb: upsert
output.schema_name: claims
output.table: CLAIMS_SILVER
output.keys: CLAIM_ID
output.event_time.column: __EXPORT_DATE

No code changes for either step — both are entries in a workflow YAML deployed through the bundle. See 03_write_verbs.md for the full verb matrix.

Recommended Combinations

The source, increment strategy (source.increment_strategy) and verb to use at each step. Every other combination, and the traps a valid config doesn't rule out, are in 08_usage.md.

Step Source Strategy Verb Why
Inbound → Bronze Files (csv, json) checkpoint APPEND Bronze keeps every export exactly as it arrived
External table → Bronze A table we didn't create (e.g. a federated one) full_read APPEND One complete snapshot per run. Sees deletes
delta_read APPEND Only the rows changed since the last run, for a table too large to read whole. Deletes are invisible
Bronze → Silver A table we stamped checkpoint FULL The newest export is the current state
checkpoint UPSERT Latest version per key (SCD1)
checkpoint · watermark SCD2 History of changes
watermark COMPLETE_DELTA, snapshot_scope: delta Replays every export in order; a deletes feed retires keys

Key Capabilities

  • Config-driven onboarding — declare an origin, a write verb and a target; the framework validates the combination and runs it. No per-dataset Python.
  • Data handling verbs — APPEND, FULL, UPSERT, SCD2, COMPLETE_DELTA — layer-agnostic, so the same verb serves Inbound→Bronze or Bronze→Silver.
  • Gold is SQL (planned) — each Gold table will be a materialized view over Silver, in its own .sql file, so Silver's MERGEs never block it. See Gold.
  • A data quality gate, not a bolt-on — every batch is checked against a Databricks DQX ruleset (you define this) before it's written; error refuses the batch, warn logs and lets it through.
  • Fail fast, fail loud — Invalid Configs or Failed Data QA stops the run with one aggregated error report. Nothing logs an error and reports success.
  • Ships as a wheel — src-layout package deployed via Databricks Declarative Automation Bundles and run through python_wheel_task entry points. No notebook logic, no sys.path hacks.

Documentation

  1. 00_overview.md — goals, principles, glossary, layer diagram
  2. 01_architecture.md — layers, protocols, execution flow, extension points
  3. 02_config_schema.md — the typed config schema, parameter by parameter
  4. 03_write_verbs.md — verb semantics with worked examples
  5. 04_policies.md — the data quality gate, driven by a DQX ruleset
  6. 05_testing.md — unit and platform testing
  7. 06_roadmap.md — phased implementation plan and current status
  8. 07_workspace_setup.md — workspace setup recipe: catalog naming, schemas, volumes, grants
  9. 08_usage.md — which source, strategy and verb to combine, and the traps

Development

Targets Linux; on Windows, work inside WSL, since Spark doesn't run natively there (Required for running unit tests).

Needs a JDK (11 or 17, for PySpark), Python 3.11 (matches Databricks Runtime 15.4 LTS) and Poetry 2.x.

poetry env use python3.11   # once, to pin the interpreter
make install                # runtime + dev dependencies
make test                   # unit tests with coverage

Review the Makefile to lists the rest.

Platform tests are real Databricks jobs under platform_tests/ - with real life scenarios; see 05_testing.md.

Deploying to Databricks

The bundle in databricks.yml is configured to work with a Python wheel: it builds the wheel, then deploys the bundle. Authenticate with a CLI profile or with DATABRICKS_HOST / DATABRICKS_TOKEN. A new workspace first needs its catalogs, schemas and volumes; see 07_workspace_setup.md.

databricks bundle deploy

Traceability

Every run that passes config validation writes an audit trail to your own catalog (monitoring_{env}.audit.logs), not to vendor telemetry. Each row names the job, run and task that produced it. The audit write comes first, so a run that can't log fails before touching any data. Schema: 00_overview.md.

Status

Start, Pipeline (CSV + JSON + table), Typing, Policies (DQX), and all five write verbs are implemented and tested. The SAS origin and the Gold example view are not implemented yet. Full phase-by-phase status: 06_roadmap.md.

License

Apache 2.0 — see LICENSE.

About

Configuration-driven data engineering framework for Databricks. Onboard datasets in YAML, not Python.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages