# What a run produces


Setup: the demo's `Orders` schema, which every example on this page uses

``` python
import dagster as dg
import dataframely as dy
import polars as pl

import dagster_dataframely as dd
from dagster_dataframely_demo.schema import Orders
```


You launch a run as you would for any Dagster asset: from the UI, from a schedule or a sensor, or with `dg.materialize([orders], resources={...})`. This package reports a run's results in two places: the table's materialization and the asset checks.


# The table's materialization

The row count is under Dagster's own key, `dagster/row_count`. Every other key this package writes is in the `dataframely/` namespace, so these keys sort together, apart from Dagster's keys and your IO manager's.

| key | what it holds |
|----|----|
| `dagster/row_count` | how many rows passed validation, not how many the function returned |
| `dataframely/valid_sample` | the first few of those rows |
| `dataframely/valid_statistics/<group>` | one table per dtype group present |
| `dataframely/quarantine_address` | where the writer wrote the invalid rows, on the runs that wrote any |
| `dataframely/invalid_count` | how many rows failed at least one rule |
| `dataframely/invalid_sample` | the first few of those, rule columns included |
| `dataframely/invalid_by_rules` | which sets of rules the rows failed together, largest group first |

The last four are absent when no rows failed. `dataframely/invalid_by_rules` counts each invalid row once, under the set of rules it failed. So one bad upstream field that fails three rules appears as one group, not as three unrelated counts. It names rules as the quarantine's rule columns do, so its names match the check names only at `rule` granularity.


<figure class="figure">
<p><img src="../assets/images/invalid-by-rules.png" class="img-fluid figure-img" /></p>
<figcaption>The materialization of a run where 8 rows failed, and the writer wrote them to the quarantine. The third row of <code>dataframely/invalid_by_rules</code> is one row that failed three rules, counted once, not three times.</figcaption>
</figure>


Every value that contains rows is a Dagster table value, not markdown, so the UI shows a sortable table. The two counts are integers, and the quarantine address is a string.


<figure class="figure">
<p><img src="../assets/images/materialization-metadata.png" class="img-fluid figure-img" /></p>
<figcaption>The materialization of a run where no rows failed: the row count, one statistics table per dtype group present, the row sample, and the IO manager's own keys below them.</figcaption>
</figure>


# The statistics tables

There is one table for each dtype group present in the written rows. Each table has one row per column, in the frame's column order. A column whose dtype is in no group, such as a `List`, `Struct` or `Array`, appears in no table, because only a count and a null count would apply to it.

| group | dtypes | columns |
|----|----|----|
| `numeric` | `Int*`, `UInt*`, `Float*`, `Decimal` | `count`, `null_count`, `mean`, `std`, `min`, `p50`, `max` |
| `temporal` | `Date`, `Datetime`, `Time`, `Duration` | `count`, `null_count`, `min`, `max`, `span` |
| `string` | `String`, `Categorical`, `Enum`, `Binary` | `count`, `null_count`, `n_unique`, `min_len`, `max_len`, `n_empty` |
| `boolean` | `Boolean` | `count`, `null_count`, `n_true`, `n_false`, `true_rate` |

The package computes `mean`, `std`, `p50` and `true_rate`, so it rounds them to four decimal places. `min` and `max` are values from the data, so the table shows them exactly. The one exception is a `Decimal`: the table shows it as a float, so it rounds any value with more digits than a float holds.

Lengths are in bytes, the unit of Dataframely's `max_length` on a `String`, so this table matches the column constraints. The table shows a `span` in Polars' own duration form, such as `8d` or `1m 30s`, not ISO-8601.

**The string group has no statistic that shows values from the data**, at any setting. A `min` or `max` on an email column would write real addresses into the event log, which a deployment shares, anyone can export, and nothing deletes. Lengths and cardinality detect the same problems.

The package computes no statistics for the invalid rows.


# The checks

The column-schema check reports first and is blocking. Each rule set then has one check.

| check metadata | on which check |
|----|----|
| `dy_rule`, `dy_rule__expr` | a check reporting for one rule |
| `dy_failed_count` | a check reporting for one rule, when any row failed |
| `dy_failed_sample` | any check, when any row failed |
| `dy_rules` | a collapsed check: a row per rule with its failure count and expression |
| `dy_schema__errors` | the column-schema check, when it fails: a row per mismatched column |

A bound appears in `dy_rule__expr`, not in the check name, so changing a `min` neither renames the check nor starts a new check history. In a collapsed check, each row of `dy_failed_sample` has a `dy_rule` column with the rule that row failed. A collapsed check has no total failure count. Counts are per rule, and one row can fail several rules, so their sum is not a row count.

`max_failure_samples` applies per rule, not per check. So in a collapsed check, a rule that a thousand rows failed cannot fill the sample and leave out a rule that one row failed.

A run that raises [NoValidRowsError](../reference/errors.NoValidRowsError.md#dagster_dataframely.errors.NoValidRowsError) yields no materialization, so the package copies `dataframely/quarantine_address` onto every check result instead.

The Columns tab is not one of these two places. It comes from the asset definition, so the catalog shows it before the first run and still shows it after a failed run.


# Attaching your own metadata

Use the context to read the run. Use the return value to write the materialization.

**Return a `dg.MaterializeResult` with the frame as its `value`.** `@dg.asset` accepts the same return value. It is the only supported way to set a materialization's tags and data version. It also works under direct invocation:


``` python
@dd.asset(Orders)
def orders(raw_orders: pl.DataFrame) -> dg.MaterializeResult[pl.DataFrame]:
    return dg.MaterializeResult(
        value=raw_orders,
        metadata={"source": "stripe", "extracted_at": "2026-08-13"},
        data_version=dg.DataVersion("2026-08-13"),
        tags={"run/flavour": "backfill"},
    )
```


`value` is the frame to validate, and you have to set it. The package copies `metadata`, `data_version` and `tags` onto the table's materialization. The quarantine has no materialization event, so none of these fields apply to it.

Setting the result's `asset_key` or [check_results](../reference/wiring.check_results.md#dagster_dataframely.wiring.check_results) raises [MaterializeResultFieldError](../reference/errors.MaterializeResultFieldError.md#dagster_dataframely.errors.MaterializeResultFieldError), which names the field. The decorator sets the asset key from its declaration and the check results from the schema's rules.

**This package's own metadata keys take precedence.** If you return `dagster/row_count` or a key in the `dataframely/` namespace, the package's value replaces yours. The package copies every other key you return onto the materialization unchanged.


<figure class="figure">
<p><img src="../assets/images/returned-result-metadata.png" class="img-fluid figure-img" /></p>
<figcaption>An asset whose returned result set <code>source</code>, <code>extract/rows</code> and <code>extract/window</code>, and also set <code>dagster/row_count</code> to 999. The materialization shows the three keys unchanged, and <code>dagster/row_count</code> shows 12, the number of rows written.</figcaption>
</figure>


## Using the context

`context.add_asset_metadata` is the other way to attach metadata, and older Dagster examples use it. The package does not block it. The decorator builds one asset with one key, so the call works without an `asset_key`:


``` python
@dd.asset(Orders)
def orders(context: dg.AssetExecutionContext, raw_orders: pl.DataFrame) -> pl.DataFrame:
    context.add_asset_metadata({"source": "stripe"})
    return raw_orders
```


**It overrides this package's own keys, unlike a returned result.** Dagster merges the context's metadata last, after the package yields its results. So `context.add_asset_metadata({"dagster/row_count": 999})` shows 999 in the catalog for a table with two rows. The package cannot prevent this.

**This package does not support `context.set_data_version`.** It is not marked `@public` in Dagster, and it does not work under direct invocation. Return `data_version=` on a `dg.MaterializeResult` instead, which sets the same event tags.

Avoid `context.add_output_metadata`. Every asset check is an output, and this package always declares the column-schema check. So an asset built with `dd.asset` always has several outputs, and the call raises:

``` text
DagsterInvariantViolationError: Attempted to add metadata without providing output_name, but multiple outputs exist. Please provide an output_name to the invocation of `context.add_output_metadata`.
```

Naming the output works: Dagster names the asset's output `result`, whatever you call the asset and under every `key_prefix`. That name is internal to Dagster, and the catalog never shows it. Use `add_asset_metadata` instead.

> **Note: Note**
>
> This package has tested only `add_asset_metadata`, `set_data_version` and `add_output_metadata` on Dagster's context, and it will test no others. It guarantees nothing about the rest of the context, in this or any future Dagster version. If you find another method that works and is worth documenting, [open an issue](https://github.com/ozanozbeker/dagster-dataframely/issues).
