Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 39 additions & 1 deletion CONTRIBUTING.rst
Original file line number Diff line number Diff line change
Expand Up @@ -322,6 +322,44 @@ If you prefer to use tox, these flags all work the same way.

tox tests/sql/parsing/queries/tpcds/test_tpcds.py::test_parsing_sparksql_tpcds_queries -- --tpcds

Property-Based Tests
--------------------

``datajunction-server/tests/property/`` holds `Hypothesis <https://hypothesis.readthedocs.io/>`_ tests. They generate
metric definitions, small fact and dimension tables, requests, and SQL expressions to check correctness rules. Most
execute DJ's SQL on DuckDB and compare it with a direct calculation or query:

- ``metric_decomposition_test.py``: a metric rolled up from its components equals the metric computed directly.
- ``sql_roundtrip_test.py``: parsing and printing SQL keeps its meaning.
- ``ast_printing_test.py``: AST-built expressions with unambiguous operator precedence print as equivalent SQL.
- ``random_graph_test.py``: SQL for a random graph and request matches a plain query over the raw tables.
- ``materialization_test.py``: a request served from a pre-aggregation matches the same request computed from raw
tables. Both paths use DJ-generated SQL; ``random_graph_test.py`` supplies the independent raw-table comparison.

``tests/property/scenario.py`` defines the immutable case, generated metric and filter specifications, and strategies.
Each metric's raw-table reference expression is written separately from its DJ definition. ``graph.py`` installs a
case through DJ's API and runs the direct DuckDB reference query. Add new case variants in ``scenario.py`` so the same
specifications can be reused across properties.

They run with the rest of the server suite. ``DJ_PBT_PROFILE`` selects ``dev`` (the default locally, random with
shrinking), ``ci`` (the default when ``CI`` is set, a fixed sequence without shrinking), or ``nightly`` (ten times the
examples, random with shrinking). CI runs only ``ci``, which replays the same examples every time, so it does not
search for new bugs. ``nightly`` is not scheduled anywhere; run it by hand for a deeper search. The API-backed properties need the same Postgres test setup as the server suite;
locally, the test fixtures start it through Docker.

.. code-block:: sh

cd datajunction-server
uv run pytest tests/property -n auto
DJ_PBT_PROFILE=nightly uv run pytest tests/property -n auto

In ``dev`` and ``nightly``, Hypothesis shrinks failures to a smaller example. The ``ci`` profile reports its failing
example without shrinking. The ``@reproduce_failure`` decorator in failure output can replay an example exactly.

Bugs found this way and not yet fixed are listed in ``tests/property/known_issues.py``. The generated tests exclude
affected inputs or metric families, and each listed bug has a small repro marked ``xfail(strict=True)``. Fixing a bug
makes its repro pass, which fails the run; then remove the marker and its corresponding exclusion.

Enabling ``pdb`` When Running Tests
-----------------------------------

Expand Down Expand Up @@ -432,4 +470,4 @@ The easiest way to fix it is to reset your database state using these commands (
root@...:/code# alembic upgrade head
...

After this, the `docker compose up` command should start the db_migration agent without problems.
After this, the `docker compose up` command should start the db_migration agent without problems.
2 changes: 2 additions & 0 deletions datajunction-server/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -161,6 +161,7 @@ testpaths = [
]
norecursedirs = [
"tests/helpers",
".hypothesis",
]

[tool.ruff.lint]
Expand Down Expand Up @@ -188,4 +189,5 @@ test = [
"sqlparse<1.0.0,>=0.4.3",
"asgi-lifespan>=2",
"mcp>=1.0.0",
"hypothesis>=6.168.1",
]
Empty file.
203 changes: 203 additions & 0 deletions datajunction-server/tests/property/ast_printing_test.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,203 @@
"""
Property: an expression tree built in code prints as SQL that means the same
thing.

DJ assembles SQL by constructing AST nodes (metric combiners, combined filters,
derived-metric substitution). This property covers expression trees whose
operators have unambiguous precedence while the known nested-parentheses bug
has a separate, fixed repro. Each tree is rendered by DJ and by a reference
renderer that parenthesizes every sub-expression, then evaluated on DuckDB.
"""

import duckdb
import pytest
from hypothesis import given, settings
from hypothesis import strategies as st

from datajunction_server.sql.parsing import ast
from tests.property.budget import examples
from tests.property.comparison import PropertyMismatch
from tests.property.known_issues import PRINTER_DROPS_PARENTHESES


class Expr:
"""A generated expression: its DJ AST and its fully parenthesized SQL."""

def __init__(self, build, reference: str):
self.build = build # fresh AST on each call; nodes carry parent links
self.reference = reference

def __repr__(self) -> str:
return self.reference


def column(name):
return Expr(lambda: ast.Column(ast.Name(name)), name)


def number(value):
return Expr(lambda: ast.Number(value), f"({value})" if value < 0 else str(value))


NULL = Expr(lambda: ast.Null(), "NULL")


def binary(op: ast.BinaryOpKind, left: Expr, right: Expr) -> Expr:
return Expr(
lambda: ast.BinaryOp(op=op, left=left.build(), right=right.build()),
f"({left.reference} {op.value} {right.reference})",
)


def negate(e: Expr) -> Expr:
return Expr(
lambda: ast.ArithmeticUnaryOp(
op=ast.ArithmeticUnaryOpKind.Minus,
expr=e.build(),
),
f"(-{e.reference})",
)


def not_(e: Expr) -> Expr:
return Expr(
lambda: ast.UnaryOp(op=ast.UnaryOpKind.Not, expr=e.build()),
f"(NOT {e.reference})",
)


def is_null(e: Expr, negated: bool) -> Expr:
return Expr(
lambda: ast.IsNull(expr=e.build(), negated=negated),
f"({e.reference} IS {'NOT ' if negated else ''}NULL)",
)


def between(e: Expr, low: Expr, high: Expr, negated: bool) -> Expr:
return Expr(
lambda: ast.Between(
expr=e.build(),
low=low.build(),
high=high.build(),
negated=negated,
),
f"({e.reference} {'NOT ' if negated else ''}BETWEEN "
f"{low.reference} AND {high.reference})",
)


def function(name: str, *args: Expr) -> Expr:
return Expr(
lambda: ast.Function(ast.Name(name), args=[a.build() for a in args]),
f"{name}({', '.join(a.reference for a in args)})",
)


def case(condition: Expr, result: Expr, otherwise: Expr) -> Expr:
return Expr(
lambda: ast.Case(
conditions=[condition.build()],
results=[result.build()],
else_result=otherwise.build(),
),
f"(CASE WHEN {condition.reference} THEN {result.reference} "
f"ELSE {otherwise.reference} END)",
)


ARITHMETIC = [ast.BinaryOpKind.Plus, ast.BinaryOpKind.Minus, ast.BinaryOpKind.Multiply]
COMPARISON = [
ast.BinaryOpKind.Eq,
ast.BinaryOpKind.NotEq,
ast.BinaryOpKind.Lt,
ast.BinaryOpKind.GtEq,
]
LOGICAL = [ast.BinaryOpKind.And, ast.BinaryOpKind.Or]

numeric_leaves = st.one_of(
st.sampled_from(["a", "b", "c"]).map(column),
st.integers(0, 3).map(number),
st.just(NULL),
)


base_comparison = st.builds(
binary,
st.sampled_from(COMPARISON),
numeric_leaves,
numeric_leaves,
)
numeric_atoms = st.one_of(
numeric_leaves,
st.builds(function, st.just("COALESCE"), numeric_leaves, numeric_leaves),
st.builds(case, base_comparison, numeric_leaves, numeric_leaves),
)
boolean_atoms = st.one_of(
st.builds(binary, st.sampled_from(COMPARISON), numeric_atoms, numeric_atoms),
st.builds(is_null, numeric_atoms, st.booleans()),
st.builds(between, numeric_atoms, numeric_leaves, numeric_leaves, st.booleans()),
)
expressions = st.one_of(
numeric_atoms,
st.builds(binary, st.sampled_from(ARITHMETIC), numeric_atoms, numeric_atoms),
numeric_atoms.map(negate),
boolean_atoms,
st.builds(binary, st.sampled_from(LOGICAL), boolean_atoms, boolean_atoms),
boolean_atoms.map(not_),
)


rows = st.lists(
st.tuples(*[st.one_of(st.none(), st.integers(-3, 3))] * 3),
min_size=1,
max_size=6,
)


def evaluate(conn, sql: str):
return conn.execute(f"SELECT {sql} FROM t ORDER BY rowid").fetchall()


@pytest.fixture(scope="module")
def conn():
with duckdb.connect(":memory:") as connection:
yield connection


@settings(max_examples=examples(500))
@given(expression=expressions, data=rows)
def test_printed_tree_means_the_same_as_the_tree(conn, expression, data):
conn.execute("CREATE OR REPLACE TABLE t (a INTEGER, b INTEGER, c INTEGER)")
conn.executemany("INSERT INTO t VALUES (?, ?, ?)", data)
expected = evaluate(conn, expression.reference)

printed = str(expression.build())
try:
actual = evaluate(conn, printed)
except duckdb.Error as exc:
raise AssertionError(
f"DJ's SQL no longer runs.\n tree: {expression.reference}\n"
f" DJ: {printed}\n error: {exc}",
) from exc
assert actual == expected, (
f"\n tree: {expression.reference}\n DJ: {printed}\n"
f" expected: {expected}\n got: {actual}"
)


@PRINTER_DROPS_PARENTHESES.xfail()
def test_known_issue_nested_negation_prints_as_comment(conn):
conn.execute("CREATE OR REPLACE TABLE t (a INTEGER, b INTEGER, c INTEGER)")
conn.execute("INSERT INTO t VALUES (1, 2, 3)")
expression = negate(negate(column("a")))
expected = evaluate(conn, expression.reference)
printed = str(expression.build())
try:
actual = evaluate(conn, printed)
except duckdb.Error as exc:
if printed.startswith("--"):
raise PropertyMismatch(
f"DJ printed {printed!r} for {expression!r}",
) from exc
raise
assert actual == expected
19 changes: 19 additions & 0 deletions datajunction-server/tests/property/budget.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
"""
How many examples each property runs.

Properties declare a budget for an ordinary PR run; `DJ_PBT_PROFILE=nightly`
multiplies it for longer searches.
"""

import os

SCALE = {"dev": 1, "ci": 1, "nightly": 10}


def profile() -> str:
default = "ci" if os.environ.get("CI") else "dev"
return os.environ.get("DJ_PBT_PROFILE", default)


def examples(n: int) -> int:
return n * SCALE[profile()]
18 changes: 18 additions & 0 deletions datajunction-server/tests/property/comparison.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
"""Value comparison and failure type for property-test correctness oracles."""

import math


class PropertyMismatch(AssertionError):
"""The result of a DJ operation differs from the independent answer."""


def same_value(actual, expected) -> bool:
"""Compare numbers approximately, but keep SQL NULL and non-finite values distinct."""
if actual is None or expected is None:
return actual is None and expected is None
if isinstance(actual, float) and math.isnan(actual):
return isinstance(expected, float) and math.isnan(expected)
if isinstance(expected, float) and math.isnan(expected):
return False
return math.isclose(actual, expected, rel_tol=1e-6, abs_tol=1e-6)
25 changes: 25 additions & 0 deletions datajunction-server/tests/property/comparison_test.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
"""The oracle must not conflate distinct SQL result values."""

import math

import pytest

from tests.property.comparison import same_value


@pytest.mark.parametrize(
("actual", "expected", "matches"),
[
(None, None, True),
(None, math.nan, False),
(None, math.inf, False),
(math.nan, math.nan, True),
(math.nan, math.inf, False),
(math.inf, math.inf, True),
(math.inf, -math.inf, False),
(1.0, 1.0 + 1e-7, True),
(1.0, 1.1, False),
],
)
def test_same_value(actual, expected, matches):
assert same_value(actual, expected) is matches
34 changes: 34 additions & 0 deletions datajunction-server/tests/property/conftest.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
"""
Hypothesis profiles for the property tests, chosen with `DJ_PBT_PROFILE`:

- dev (default locally): random examples; failures are saved and replayed.
- ci (default when `CI` is set): a fixed sequence of examples and no shrinking,
so runs are repeatable and a failure is reported quickly.
- nightly: ten times the examples, random, with shrinking.
"""

from hypothesis import HealthCheck, Phase, settings

from tests.property.budget import profile

COMMON = {
# Examples that create nodes through the API take well over the default
# 200ms, and vary with load.
"deadline": None,
"suppress_health_check": [
HealthCheck.function_scoped_fixture,
HealthCheck.too_slow,
],
"print_blob": True,
}

settings.register_profile("dev", **COMMON)
settings.register_profile(
"ci",
derandomize=True,
database=None,
phases=[Phase.explicit, Phase.reuse, Phase.generate],
**COMMON,
)
settings.register_profile("nightly", **COMMON)
settings.load_profile(profile())
Loading
Loading