Guide¶
tusk turns a set of related tables into a single wide feature matrix. The workflow is always the same three steps:
- Describe your data. Build a database: register each
table, say which column is its primary key and which column records when a
row became knowable, then link the tables with relationships. Those are
declarations tusk takes on trust — call
validate()when you want them checked against the data. - Synthesize. Call
deep_feature_synthesis(). It walks the relationship graph, stacking primitives up tomax_depth, and returns both the feature matrix and the feature definitions that produced it. - Re-apply. Feed those definitions back to
apply_features()to compute the same columns on new data.
Everything tusk builds is a narwhals expression, so the whole pipeline is one query plan on the backend you already use. Nothing is materialized until you ask for it.