# Settings


Setup: the demo's `Orders` schema, which every example on this page uses

``` python
import dagster as dg
import dataframely as dy
import polars as pl

import dagster_dataframely as dd
from dagster_dataframely_demo.schema import Orders
```


This package resolves each setting from three sources, each overriding the one before: the package default, then an environment variable, then the argument to `dd.asset`. An environment variable sets a default for the code location, and an argument overrides it for one asset. Each variable is `DAGSTER_DATAFRAMELY_` plus the setting's name in upper case.

| setting | what it sets | default |
|----|----|----|
| `check_granularity` | how many checks the package collapses the schema's rules into: `rule`, `column` or `schema` | `rule` |
| `schema_rules` | how the package reports the schema-level rules at `column` granularity: `collapsed` or `per_rule` | `collapsed` |
| `statistics` | whether each materialization has statistics for the rows it wrote | `true` |
| `max_failure_samples` | how many of the rows that failed a rule appear in that rule's check | `5` |
| `row_sample` | how many valid rows and how many invalid rows appear in the materialization | `5` |
| `quarantine_dir` | the directory the package writes invalid rows to when you call the asset instead of running it | unset: a call with invalid rows raises [QuarantineDirError](../reference/errors.QuarantineDirError.md#dagster_dataframely.errors.QuarantineDirError) |

The package resolves the settings once, when you declare the asset, and the check specs and the validation both use those values.

`quarantine_dir` has two sources, not three, because `dd.asset` takes no argument for it ([ADR-0006](https://github.com/ozanozbeker/dagster-dataframely/blob/main/docs/pre-1.0.md#adr-0006-the-assets-own-io-manager-writes-the-quarantine)). The package reads its value when the writer writes the rows, not when you declare the asset. An empty value still raises [InvalidSettingError](../reference/errors.InvalidSettingError.md#dagster_dataframely.errors.InvalidSettingError) when you declare the asset, not on the first call with invalid rows. For example, `DAGSTER_DATAFRAMELY_QUARANTINE_DIR=${SCRATCH}` is empty when `SCRATCH` is unset.

The package validates every source when it resolves the setting, including the package default and the argument. An invalid value raises [InvalidSettingError](../reference/errors.InvalidSettingError.md#dagster_dataframely.errors.InvalidSettingError), which names the value and its source. So `statistics="false"` raises instead of turning statistics on, and `max_failure_samples=True` raises instead of meaning one row.

**A flag's environment variable accepts `true` or `false`, and nothing else.** Case does not matter, so `TRUE` is the same as `true`. `1`, `yes` and `on` raise [InvalidSettingError](../reference/errors.InvalidSettingError.md#dagster_dataframely.errors.InvalidSettingError).

There is no fourth source and no `set_default_*()` function.


# Changing `check_granularity` starts a new check history

`rule` gives each rule its own check and its own history. `column` gives one check per column with rules, `dy_col__<column>`, so a 40-column schema's check list stays readable. A column's check reports for every rule on that column. For a `Struct` column, Dataframely generates one `inner_<field>_nullability` rule per field, so a ten-field struct has ten rules and one check. `schema` gives one check, `dy_schema__rules`, for every rule.

`schema_rules` sets how the package reports the schema-level rules at `column` granularity, because they belong to no single column. `collapsed` puts them all in one check, `dy_schema__rules`, and `per_rule` gives each one its own check. The setting has no effect at `rule` or `schema` granularity.

A schema with no rules has only the column-schema check, at every granularity. That check is always present and always blocking.

**Changing `check_granularity` on an asset that has already run starts a new check history.** The asset no longer reports the old checks, so their history ends with the last run before the change. The new checks start with no history. Nothing moves the old history to the new checks, so choose the granularity before you deploy the asset.


# Statistics and both samples are on by default

Each materialization has summary statistics for the rows it wrote. [The statistics tables](what-a-run-produces.md#the-statistics-tables) lists their columns.

> **Important: Important**
>
> Two of the settings write **real rows of your data to the Dagster event log**, and both are on by default. A deployment shares one event log, and you can export it. This package redacts nothing. If a column contains an email address, a name or an account number, the package writes that value to the log, where it stays.

| setting | what it writes | where |
|----|----|----|
| `max_failure_samples` | up to this many of the rows that failed each rule | that rule's asset check, under `dy_failed_sample` |
| `row_sample` | up to this many valid rows and this many invalid rows | the materialization, under `dataframely/valid_sample` and `dataframely/invalid_sample` |

A sample is the first rows of the frame, not a random selection. A sample is absent, never empty: with no rows to show, the package writes no metadata key for it.

Dataframely's comparable setting defaults to `0`. This package defaults to `5`, because a count shows how many rows failed a rule but not what they contain.

Set either setting to `0` to turn it off. Per asset:


``` python
@dd.asset(Orders, max_failure_samples=0, row_sample=0)
def orders(raw_orders: pl.DataFrame) -> pl.DataFrame:
    return raw_orders
```


Or set it for a whole code location, in the deployment's environment:

``` bash
DAGSTER_DATAFRAMELY_MAX_FAILURE_SAMPLES=0
DAGSTER_DATAFRAMELY_ROW_SAMPLE=0
```

Turning the samples off does not change `statistics`, which is a separate setting. A sample shows a `Decimal` as a string, with every digit. The statistics tables convert it to a float for display.
