Setup: the demo’s Orders schema, which every example on this page uses
import dagster as dg
import dataframely as dy
import polars as pl
import dagster_dataframely as dd
from dagster_dataframely_demo.schema import Orders
This package resolves each setting from three sources, each overriding the one before: the package default, then an environment variable, then the argument to dd.asset. An environment variable sets a default for the code location, and an argument overrides it for one asset. Each variable is DAGSTER_DATAFRAMELY_ plus the setting’s name in upper case.
check_granularity |
how many checks the package collapses the schema’s rules into: rule, column or schema |
rule |
schema_rules |
how the package reports the schema-level rules at column granularity: collapsed or per_rule |
collapsed |
statistics |
whether each materialization has statistics for the rows it wrote |
true |
max_failure_samples |
how many of the rows that failed a rule appear in that rule’s check |
5 |
row_sample |
how many valid rows and how many invalid rows appear in the materialization |
5 |
quarantine_dir |
the directory the package writes invalid rows to when you call the asset instead of running it |
unset: a call with invalid rows raises QuarantineDirError |
The package resolves the settings once, when you declare the asset, and the check specs and the validation both use those values.
quarantine_dir has two sources, not three, because dd.asset takes no argument for it (ADR-0006). The package reads its value when the writer writes the rows, not when you declare the asset. An empty value still raises InvalidSettingError when you declare the asset, not on the first call with invalid rows. For example, DAGSTER_DATAFRAMELY_QUARANTINE_DIR=${SCRATCH} is empty when SCRATCH is unset.
The package validates every source when it resolves the setting, including the package default and the argument. An invalid value raises InvalidSettingError, which names the value and its source. So statistics="false" raises instead of turning statistics on, and max_failure_samples=True raises instead of meaning one row.
A flag’s environment variable accepts true or false, and nothing else. Case does not matter, so TRUE is the same as true. 1, yes and on raise InvalidSettingError.
There is no fourth source and no set_default_*() function.
Changing check_granularity starts a new check history
rule gives each rule its own check and its own history. column gives one check per column with rules, dy_col__<column>, so a 40-column schema’s check list stays readable. A column’s check reports for every rule on that column. For a Struct column, Dataframely generates one inner_<field>_nullability rule per field, so a ten-field struct has ten rules and one check. schema gives one check, dy_schema__rules, for every rule.
schema_rules sets how the package reports the schema-level rules at column granularity, because they belong to no single column. collapsed puts them all in one check, dy_schema__rules, and per_rule gives each one its own check. The setting has no effect at rule or schema granularity.
A schema with no rules has only the column-schema check, at every granularity. That check is always present and always blocking.
Changing check_granularity on an asset that has already run starts a new check history. The asset no longer reports the old checks, so their history ends with the last run before the change. The new checks start with no history. Nothing moves the old history to the new checks, so choose the granularity before you deploy the asset.
Statistics and both samples are on by default
Each materialization has summary statistics for the rows it wrote. The statistics tables lists their columns.
Two of the settings write real rows of your data to the Dagster event log, and both are on by default. A deployment shares one event log, and you can export it. This package redacts nothing. If a column contains an email address, a name or an account number, the package writes that value to the log, where it stays.
max_failure_samples |
up to this many of the rows that failed each rule |
that rule’s asset check, under dy_failed_sample |
row_sample |
up to this many valid rows and this many invalid rows |
the materialization, under dataframely/valid_sample and dataframely/invalid_sample |
A sample is the first rows of the frame, not a random selection. A sample is absent, never empty: with no rows to show, the package writes no metadata key for it.
Dataframely’s comparable setting defaults to 0. This package defaults to 5, because a count shows how many rows failed a rule but not what they contain.
Set either setting to 0 to turn it off. Per asset:
@dd.asset(Orders, max_failure_samples=0, row_sample=0)
def orders(raw_orders: pl.DataFrame) -> pl.DataFrame:
return raw_orders
Or set it for a whole code location, in the deployment’s environment:
DAGSTER_DATAFRAMELY_MAX_FAILURE_SAMPLES=0
DAGSTER_DATAFRAMELY_ROW_SAMPLE=0
Turning the samples off does not change statistics, which is a separate setting. A sample shows a Decimal as a string, with every digit. The statistics tables convert it to a float for display.