Changelog

This changelog is generated automatically from GitHub Releases.

v0.9.0

2026-09-22 · GitHub

Everything user-visible in 0.9 is prose. The package used to explain itself with metaphors that no upstream library uses, and every page, docstring, comment and error message now says what the code does in Dagster’s, Dataframely’s and Polars’ own words. Nothing behaves differently, and no check history moves.

Breaking

  • Two error classes are renamed. Each is raised in the same cases as before.
  • Most error messages are reworded, so code that matches on message text needs updating.
  • The floors move to dagster>=1.13.24, dataframely>=3.1.2 and polars>=1.44.2. The tests and every measurement in this repo run against those versions.
0.8 0.9
dd.errors.NothingSurvivedError dd.errors.NoValidRowsError
dd.errors.UnnameableColumnError dd.errors.InvalidColumnNameError

Documentation

  • The documentation is a site now: https://ozanozbeker.com/dagster-dataframely. The guide is one page per topic, and every code cell runs when the site builds, so an example that stops being true fails the build. USER_GUIDE.md is gone.
  • A demo code location backs the examples. A prek hook keeps the README’s code blocks in sync with the demo’s modules.
  • The decisions and the measurements the package was built on are in one file, docs/pre-1.0.md. It replaces docs/adr/ and docs/research/. Every claim in it was re-checked against the code, and a claim the code no longer supports was deleted rather than corrected. ARCHITECTURE.md is gone, and its diagrams are in the guide.

The full upgrade log is in CHANGELOG.md.

v0.8.0

2026-09-14 · GitHub

A LazyFrame return is now filtered in the engine, so disk staging is gone. The rest of the release settles names on the words Dagster, Dataframely and Polars already own: nothing behaves differently under the renames, and no check history moves.

Breaking

  • temp_dir and DAGSTER_DATAFRAMELY_TEMP_DIR are gone. Schema.filter runs in one collect_all on the streaming engine, so there is no staging file to place: drop the argument and the variable, and nothing else changes. The measurements are in docs/pre-1.0.md.
  • The decorator is dd.asset. dy_asset put a user-typed name inside dy_, which otherwise names only what this package generates into Dagster. Dagster owns the word asset, and Dataframely settles the same question the same way, shadowing 28 Polars names under its own alias. Import the module, not the name: a bare from dagster_dataframely import asset reads as Dagster’s own.
  • Two functions are renamed for what they return, and nothing about either changes but the name.
0.7 0.8
[@dd](https://github.com/dd).dy_asset(Orders) [@dd](https://github.com/dd).asset(Orders)
dd.build_quarantine_spec dd.quarantine_spec
dd.wiring.process dd.wiring.validation_results
  • A terminology pass renamed what this package had invented where an upstream word already existed.
0.7 0.8 the word it borrows
multi_column_rules= schema_rules= Dataframely calls these schema-level rules, and a single-column primary_key is one, so “multi-column” was false
dd.MultiColumnRules dd.SchemaRules
DAGSTER_DATAFRAMELY_MULTI_COLUMN_RULES DAGSTER_DATAFRAMELY_SCHEMA_RULES
schema_rules="schema" schema_rules="collapsed" the old value repeated the setting’s own name
dataframely/valid_stats/<family> dataframely/valid_statistics/<group> Polars’ DataTypeGroup

Added

  • dd.wiring.check_results answers every check check_specs declares, for an asset that writes its own storage. It is validation_results without the write, so it never raises ValidationAbortError or NothingSurvivedError, and it takes severity rather than deriving one.
  • A run of a quarantined asset fails with QuarantineKeyCollisionError before its body when another asset in the code location already materializes <name>_quarantine (ADR-0007). Declaring quarantine=True beside such an asset used to lose the invalid rows in silence.
  • ReservedColumnError and CheckNameCollisionError now come from every public function that takes a schema, not only from check_specs (ADR-0008).
  • UnnameableColumnError joins them, for a column spelled in anything but A-Za-z0-9_. A column comes by such a name through dy.Column(alias=...); the fix is to rename the alias. An alias holding Dataframely’s own | delimiter used to surface as KeyError on a column nobody declared.

Fixed

  • DAGSTER_DATAFRAMELY_QUARANTINE_DIR is read where the invalid rows are written rather than where the asset is declared. A deployment that sets it after the module imported is now read, and a call holding nothing back needs no directory at all.
  • InvalidSettingError no longer sends you to a quarantine_dir= argument that does not exist. A setting with no argument now names the two sources it does resolve through, and says outright that there is no third.

Documentation

The README is a landing page and a quick start; USER_GUIDE.md is everything else. ARCHITECTURE.md is new: the two paths, the six outcomes, and a code reference index. dataframely/invalid_by_rules is now documented as naming each rule the way the quarantine’s own columns name it, at every granularity.

The full upgrade log is in CHANGELOG.md.

v0.7.0

2026-09-02 · GitHub

The quarantine leaves the graph and lands wherever your IO manager puts things. This package ships no storage of its own any more, so a DuckDB or Snowflake shop keeps its invalid rows in the warehouse instead of losing them to a base_dir that has no meaning there.

Removed

  • Both IO managers, both factories, the CSV codecs and the schema carrier. 1988 lines out. dagster-polars writes Polars frames to a filesystem or object store and dagster-duckdb-polars writes them to a warehouse. Every storage-backed test moved onto dagster-polars before the delete, so nothing here was covering ground those don’t (#109).
  • UnwritableDtypeError, QuarantineSettingError and wiring.quarantine_table_schema. The first belonged to the CSV writer. The second has no setting left to guard, because nothing about the quarantine is configurable now. The third described a Columns tab that a quarantine spec carries itself (#109, #113).

Changed

  • dataframely_asset is dy_asset, and the schema is its only positional argument (#113).
  • quarantine= is a bool, and the quarantine is no longer an out. True means “wherever this asset’s manager puts things”, which needs no configuration, so quarantine=dg.AssetOut() has nothing left to say. It also has no node in the graph: build_quarantine_spec(Orders, orders) is what puts one there, and only if you want it drawn (#109, #113).
  • The asset’s own IO manager writes the invalid rows, addressed by asset key. The writer borrows the manager off the step’s output handle rather than declaring a resource key, because Dagster validates required_resource_keys at bind time and every direct invocation would then have to supply a manager it has no use for. ADR-0006 records the decision and pins the seven private Dagster APIs it rests on (#109, #113).
  • Invalid rows are written on every exit that has them, aborts included. As a second out they were skipped on an abort, because the out was never yielded (#113).
  • Metadata moves under dataframely/. dagster-polars writes table and stats onto the same materialization this package was writing sample and stats/* onto, so a reader saw two samples and two statistics blocks, with stats sorting against stats/numeric. The namespace removes a collision that was already live rather than adding a convention. Check-result metadata stays dy_, because a check name cannot hold a slash: dg.AssetCheckSpec(name="dataframely/x") constructs fine and then fails when the spec becomes an op output (#116).
  • SchemaShapeError is ColumnSchemaError and dy_schema__dtypes is dy_schema__columns. “Shape” is Polars’ word for something else, and the check has always compared columns and dtypes both (#116).
  • The quarantine location override is deferred rather than shipped. ADR-0006 promised one; the three cases it would cover are two spellings of where and one of who, and a root means nothing to a warehouse. Pre-1.0 a rename costs more than the wait, so the surface waits for a user (#110).

Added

  • build_quarantine_spec(Orders, orders) returns the spec that draws the quarantine as a child of the valid asset, with its own Columns tab (#109).
  • wiring.quarantine_path, wiring.file_writer, wiring.delegating_writer and the QuarantineWriter protocol. A called asset reaches no IO manager, so DAGSTER_DATAFRAMELY_QUARANTINE_DIR is the fallback address for direct invocation, and QuarantineDirError says so when it is unset (#109, #113).
  • The README has an ## Upgrading from 0.6 section, and the whole of it is rewritten on the settled vocabulary. The glossary lost twelve terms (#116).

Upgrading

Every asset key and every rule column is unchanged, so check history survives all of this except the column-schema rename. The full table is in Upgrading from 0.6.

0.6 0.7 note
dd.dataframely_asset(schema=Orders) dd.dy_asset(Orders) the schema is positional
quarantine=dg.AssetOut() quarantine=True it’s a bool, and True needs no configuration
dd.DataframelyParquetIOManager dagster_polars.PolarsParquetIOManager this package ships no IO manager
dd.DataframelyCSVIOManager none the CSV codecs went with it; use parquet or a warehouse
the quarantine was a second out dd.build_quarantine_spec(Orders, orders) and only if you want it in the graph
dd.errors.SchemaShapeError dd.errors.ColumnSchemaError “shape” was Polars’ word for something else
dd.errors.QuarantineSettingError none nothing on the quarantine is configurable now
dd.errors.UnwritableDtypeError none it belonged to the CSV writer
dd.wiring.quarantine_table_schema none a quarantine spec carries its own Columns tab
dd.wiring.process(..., quarantine_key=...) dd.wiring.process(..., quarantine_writer=...)
dy_schema__dtypes dy_schema__columns this orphans that check’s history
sample dataframely/valid_sample
stats/<family> dataframely/valid_stats/<family>
dy_quarantine_path dataframely/quarantine_address it can be an asset key now, not just a file path
dy_rejected_count dataframely/invalid_count
dy_rejected_sample dataframely/invalid_sample
dy_rejected_rules dataframely/invalid_by_rules

dy_failed_count and dy_failed_sample are unchanged. So are check_granularity, multi_column_rules, statistics, max_failure_samples, row_sample and temp_dir.

quarantine=True adds a context parameter to the asset whether or not your function declared one, because the writer is built from the execution context and nothing else can reach it. Calling one directly now takes a dg.build_asset_context() first. An asset without a quarantine is unchanged.

v0.6.0

2026-09-02 · GitHub

A partition that has no source data, and never will, can now say so. Return None and the asset skips.

Added

  • Returning None from a decorated function skips the asset. Nothing is validated, neither out materializes, and the run stays green, so the partition stays unmaterialized instead of going green with zero rows or red with an error. None joins the four things a decorated function could already return (#95).
  • The existence test stays yours. The decorator never catches FileNotFoundError, or anything else, to decide this for you. It cannot tell a file that is legitimately absent from a path that is misconfigured, and guessing wrong would turn a broken pipeline into a silently missing partition.
  • Every check still reports on a skipped run, and passes. A check spec is a non-optional op output whatever the asset declares, so a step that answers none of them fails outright. The rules run over Schema.create_empty() and report what that says: every rule evaluated, over zero rows, none violated. Dagster attaches those evaluations to no materialization, so a passing check on a skipped partition does not claim to have checked the last one that had data.
  • The README has an ## A partition with no data section, with the monthly x distributor grid the feature was written for.

Changed

  • Every docstring and comment is rewritten in plain language. No behaviour moves. The diff is large and touches nearly every module.
  • dg.MaterializeResult(value=None) is refused rather than read as the skip. A returned result exists to put metadata, tags or a data version on a materialization, and a skipped run has none.
  • Dependency floors now name the versions the test suite runs against: dagster>=1.13.20, polars>=1.44.1 (#96). pydantic>=2 and universal-pathlib>=0.2.0 stay loose, since they guard a direct import rather than track a version.

Upgrading

Nothing you wrote has to change unless a decorated function of yours already returned None.

A forgotten return no longer raises. _require_frame used to catch it, because None was not a value any correct function produced. It is now the skip, so the mistake and the intent are the same object and the guard cannot tell them apart. A function that falls off the end materializes nothing and reports passing checks. The missing partition is what makes it visible.

v0.5.0

2026-08-13 · GitHub

The quarantine is now drawn as a child of the table it came from. One edge in the graph; the rest of this release is what it took to be sure of it.

Changed

  • The quarantine’s only parent is the valid asset. A [@dg](https://github.com/dg).multi_asset gives every out every input, so the quarantine used to render as a second child of every upstream, which said it was reachable and materializable on its own. It is not: the two are outs of one step. ADR-0003 records the decision (#92).
  • The README has an ## Automation section, which it had none of. automation_condition and freshness_policy never reach the quarantine, nothing needs wiring between the two tables because one step writes both, and automating on invalid rows is your own asset pointed at the quarantine’s key (#92).
  • The pin advice in the quick start reads >=0.5,<0.6. It had named <0.2 since 0.2.0 shipped, so it excluded every published version of the package.

Upgrading

Nothing you wrote has to change. The decorator’s parameters, the asset keys, the check names and the quarantine’s rule columns are unchanged, and a run does what it did before. What moves is lineage.

An upstream selection over the quarantine now includes the valid table. AssetSelection.upstream(orders_quarantine) reaches orders, where it used to reach only the raw tables. Anything defining a job or a backfill that way selects one more asset than it did.

A quarantine on an asset with no upstream of its own now reads stale after a clean run. A clean run skips the quarantine rather than writing an empty table, so its last write is older than the valid table’s. An asset with any upstream at all already read stale on those runs, so this is new only for a root asset.

The edge is asset-grained rather than row-grained. No row in the quarantine came from the valid table, since Schema.filter splits one frame into two, so column-lineage tooling will trace through it and be wrong about where those rows came from. Stated in the README and in the ADR rather than left to be discovered.

v0.4.0

2026-08-13 · GitHub

Two things a user reaches for constantly: a decorated asset you can call in a unit test, and a dg.MaterializeResult you can return. The root namespace went from 26 names to 7 on the way, so this release is mostly a one-line import change and then more surface than before.

Added

  • A decorated asset is directly invocable. Calling it hands back the validated frame, the quarantine frame and every check result as ordinary Python objects, with no run, no IO manager and no instance. dg.build_asset_context() supplies a context when the decorated function declares one, partition_key included. This is Dagster’s documented unit-testing path, and it cost every user of this package a whole run before (#72).
  • A transform may return a dg.MaterializeResult carrying the frame. value is the frame to validate; metadata, data_version and tags fold onto the valid table’s materialization. It is the only supported route to a materialization’s tags and data version. The package’s own metadata keys are applied last, so a returned dagster/row_count loses to the one that was counted, and asset_key and check_results are refused by name. Returning a bare frame is byte-for-byte what it was (#77).
  • The asset’s description falls back to the schema’s docstring before Dagster’s fallback to the transform’s, so description=MySchema.__doc__ at the call site is no longer a thing anyone writes (#75).

Changed

  • Errors moved to dd.errors, hand-wiring moved to dd.wiring. Ten of the root’s twenty-four names were things you reach for only once something has gone wrong, and seven more were for an arrangement most users never build (#76).
  • DataFramePartitions and LazyFramePartitions are gone. An alias that hides dict[str, pl.DataFrame] subtracts information at the call site, which is fatal for a name exported to teach that shape (#76).
  • process takes asset keys rather than the execution context. valid_key and quarantine_key replace the context parameter, which the function only ever used to call asset_key_for_output twice. ADR-0001 records the decision (#66).
  • The op is named after the whole asset key. [@dataframely](https://github.com/dataframely)_asset(key_prefix="sales", name="orders") builds an op called sales__orders, which is how [@dg](https://github.com/dg).asset names its own. An asset name alone is not unique across a code location; two assets sharing a name under different prefixes built two ops called the same thing (#70).
  • The frame guard’s message names three routes out, not one. Sending the reader to a plain [@dg](https://github.com/dg).asset was the whole of the old advice, which was wrong for anyone who wanted metadata on a validated table (#77).
  • The README is a quick start and a user guide rather than a tour of mechanisms, and the hand-wiring section ends in eighty lines of raw Dagster and Dataframely that do what the decorator does (#78).

Upgrading

Imports move; nothing else you wrote has to change:

# 0.3.0
from dagster_dataframely import SchemaShapeError, check_specs, process

# 0.4.0
from dagster_dataframely.errors import SchemaShapeError
from dagster_dataframely.wiring import check_specs, process

Hand-wired assets pass keys instead of the context:

# 0.3.0
yield from process(Orders, frame, context)

# 0.4.0
yield from process(Orders, frame, valid_key=context.asset_key_for_output("orders"))

Annotate a fan-in with the shape itself, dict[str, pl.DataFrame] or dict[str, pl.LazyFrame], in place of the two removed aliases.

An asset that declares a key_prefix has a new op name. Run config, re-execution from a step key and step-level concurrency addressed orders and now address sales__orders. Asset keys, check names and quarantine rule columns are unchanged, so no check history is orphaned. An asset with no key_prefix is unaffected.

An asset carrying both a schema docstring and a transform docstring, passing no description=, now renders the schema’s. Pass the transform’s as description= at the call site if that was the one you wanted.

v0.3.0

2026-08-12 · GitHub

Every term this package invented, audited against what Dagster and dataframely say themselves. Where an upstream word already existed, that word wins. Behaviour is unchanged: this release is naming.

Changed

  • SchemaGateError is now SchemaShapeError. It is raised when a frame’s columns or dtypes disagree with the schema, and the package already called that comparison the frame’s shape. “Gate” was a word neither upstream uses (#60).
  • process(good_out=) is now process(valid_out=). Schema.filter returns “the validated rows” and FailureInfo.invalid() returns “the invalid rows”, so the package had renamed dataframely’s own axis to good and rejected. Only hand-wiring passes this argument; dataframely_asset passes it for you (#60).
  • Four more coinages were retired internally, none of them on the public surface: constraint for pill, decorator for door, rule set for bucket, and staging for the local temp file a lazy plan streams to (#60).

Added

  • CONTEXT.md at the repo root: the project’s glossary, one entry per term, each listing the words it deliberately avoids (#60).

Upgrading

Two renames, both mechanical:

SchemaGateError     ->  SchemaShapeError
process(good_out=)  ->  process(valid_out=)

The asset-check name dy_schema__dtypes is unchanged, so upgrading does not orphan any check history. Nothing else on the public surface moved: dataframely_asset and both IO managers take the arguments they did in 0.2.0.

If you only use [@dataframely](https://github.com/dataframely)_asset and never catch SchemaGateError by name, there is nothing to change.

v0.2.0

2026-08-12 · GitHub

Laziness, as far as each path can carry it. A pl.LazyFrame now reads back as an unexecuted scan, streams to a local parquet before validation, and on a plain asset streams all the way to storage.

Added

  • Lazy reads. Annotate an input pl.LazyFrame and the IO manager hands back an unexecuted scan, so a downstream filter or select prunes rows and columns before anything is decoded. The scan rides the same fsspec handle the eager read uses, so a cloud scheme needs no second set of credentials (#52).
  • Lazy writes on a plain asset. Return a pl.LazyFrame from a [@dg](https://github.com/dg).asset and the plan is sunk through the streaming engine to a local temp file, then promoted to base_dir once it succeeded. Peak memory is the engine’s buffers rather than the whole frame. A failing plan writes nothing to the destination and leaves a file already there untouched (#54).
  • Temp landing for a validated transform. A dataframely_asset whose transform returns a pl.LazyFrame streams that plan to a local parquet, then validates it eagerly as before. Peak memory becomes the size of the landed frame rather than the plan’s high-water mark, which is the saving for a transform with a large intermediate (#53).
  • temp_dir, on dataframely_asset and as DAGSTER_DATAFRAMELY_TEMP_DIR, deciding which disk a lazy plan lands on. Unset, that is the system temp directory, which in a container is its ephemeral disk (#53).
  • DataFramePartitions and LazyFramePartitions, the shapes a fan-in over every partition has to annotate. The obvious spelling, pl.DataFrame, fails Dagster’s type check after every partition has already been read (#35, #52).

Changed

  • A pl.LazyFrame output used to collect into memory, log a warning, and write. It now sinks and promotes. A run therefore needs writable temp space it did not need before, and parquet row groups can differ from an eager write of the same frame. CSV is byte-identical either way.
  • A pl.LazyFrame input annotation used to fail Dagster’s type check.

Fixed

  • The quarantine’s co-occurrence table is now ordered, biggest group first. It came off an unordered group_by, so two runs of the same data diffed as though something had changed.

Upgrading

Nothing to change in your code. If your executor has no writable temp space, or a small one, set DAGSTER_DATAFRAMELY_TEMP_DIR at a volume before returning a pl.LazyFrame from anything.

The reasoning behind where laziness stops, and the measurements it rests on, are in docs/pre-1.0.md.

v0.1.0

2026-08-11 · GitHub

The first release with the whole of v1 in it.

0.0.1 was a seven-module snapshot pushed mid-build to hold the name on PyPI. This is the package the README describes.

One declaration

[@dd](https://github.com/dd).dataframely_asset(schema=Orders)
def orders(raw_orders: pl.DataFrame) -> pl.DataFrame:
    return raw_orders.select("order_id", "amount")

From that alone: the catalog’s Columns tab fills in before the asset has ever run, with dtypes, descriptions, nullability, uniqueness, the primary key and one pill per remaining constraint; every dataframely rule reports through an asset check with its own pass/fail history, behind a blocking gate check that compares columns and dtypes before a single row is filtered. Add quarantine=dg.AssetOut() and rejected rows land in a sibling asset carrying the original columns plus one outcome column per rule.

In the box

  • dataframely_asset, with [@dg](https://github.com/dg).asset’s vocabulary over a multi_asset mechanism.
  • DataframelyParquetIOManager and DataframelyCSVIOManager, writing to a local directory or to s3://, gs:// and az://. The CSV manager encodes the five dtypes a cell cannot hold, each with a declared inverse, so the frame read back equals the frame written.
  • Summary statistics on every materialization, as one table per dtype family.
  • Row samples, on by default: five of the rows each rule rejected, and five of the rows the asset wrote. Those are real rows in the Dagster event log, which is shared, exportable and not redacted. Read Statistics and both samples are on by default before running this against data you care about.
  • A three-tier settings chain: the package default, then a DAGSTER_DATAFRAMELY_* environment variable, then the argument on the asset.
  • The kit, for a wiring the decorator does not offer.
  • py.typed.

What it deliberately does not do

  • It never casts. A frame whose shape is not the schema’s aborts before anything is written.
  • Without a quarantine, every row has to be good. There is no lenient mode and no strict flag. Dropping rows is a line you write in your own asset body.
  • Validation materializes. The habitat is post-landing transformation, not ingestion scale. Sinking lazily is tracked in #27.
  • dy.Collection is not supported. Declare one asset per member.

Compatibility

Python 3.12, 3.13 and 3.14, against dagster>=1.13.16, dataframely>=3.0.0 and polars>=1.43.2.

Pre-1.0: the public surface is pinned by a test rather than held by convention, but a 0.x minor is where a breaking change lands. Pin dagster-dataframely>=0.1,<0.2 if that matters to you.

What’s Changed

  • feat: route rejected rows to a quarantine sibling asset by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/28
  • feat: render schema constraints as pills, a table key, and check descriptions by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/29
  • docs: record what partitioned assets actually do by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/30
  • feat: collapse the check list through a three-tier settings chain by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/36
  • feat: store polars frames as CSV through lossless codecs by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/37
  • feat: profile every materialization as four statistics tables by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/38
  • feat: sample rejected rows into checks and written rows onto materializations by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/39
  • feat: close the public surface and rewrite the README by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/40
  • fix: validate a flag’s argument tier by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/41
  • chore: release 0.1.0 by @ozanozbeker in https://github.com/ozanozbeker/dagster-dataframely/pull/42

New Contributors

  • @ozanozbeker made their first contribution in https://github.com/ozanozbeker/dagster-dataframely/pull/28

Full Changelog: https://github.com/ozanozbeker/dagster-dataframely/compare/v0.0.1…v0.1.0

v0.0.1

2026-08-07 · GitHub

Name-reservation release. The public surface is not stable yet; see #15 for the spec.

Working today

  • [@dataframely](https://github.com/dataframely)_asset attaches a dataframely Schema to a Dagster asset and derives one asset check per validation rule, statically from the schema.
  • DataframelyParquetIOManager writes Parquet to a local directory or to s3://, gs:// and az://.

Not yet

Quarantine (#19), the rule-text renderer (#20), the CSV IO manager (#22), and the README rewrite (#26) — the current README understates what ships.

Requires Python 3.12+.