关于Featuretools/DFS生成特征向量类型及疏密性的技术咨询
Hey there, great questions about Featuretools' DFS output—let's unpack them clearly:
1. What type of feature vectors does Featuretools/DFS generate?
Featuretools' Deep Feature Synthesis (DFS) spits out Pandas DataFrames by default (or Dask DataFrames if you're working with distributed, large-scale datasets). The actual feature columns fall into standard data types you'd expect in tabular data:
- Numeric types (
int64,float64): Come from aggregations likemean,sum, orcount, or date transformations like extracting themonthfrom an order date. - Boolean types (
bool): Generated from logical checks—thinkis_weekend(order_date)orhas_any(returns)(flagging if a customer ever made a return). - Categorical/object types: Either raw categorical features passed through DFS, or derived features like the
mode(most frequent value) of a categorical column, or combined category attributes.
For example, running DFS on an e-commerce dataset might give you columns like mean(order_items.price) (float64), count(customer.orders) (int64), and is_holiday(order_date) (bool) in your final feature set.
2. Are the feature vectors dense, sparse, or dependent on factors?
By default, DFS outputs dense feature vectors because Pandas DataFrames use dense storage out of the box. But sparsity (or the ability to work with sparse formats) depends on a few key factors:
- Raw data sparsity: If your input data has lots of missing values or zero-inflated columns, the generated features will inherit that sparsity. For example, if 90% of customers never made a return, a feature like
sum(returns.amount)will be 0 or NaN for almost all rows. - Feature cardinality and type: High-cardinality categorical features (especially if you one-hot encode them later using
featuretools.encode_features) can create extremely sparse matrices. Aggregation features that only have non-zero values for a tiny subset of rows also act like sparse features. - Manual conversion: Featuretools has a handy
pandas_to_sparseutility that turns a dense DataFrame into a sparse Pandas DataFrame (usingSparseArrayfor memory efficiency). You can also export features to ascipy.sparsematrix if you need to work with sparse formats for machine learning models that handle them well.
So to sum up: Default is dense, but sparsity can come from your data or feature choices, and you can explicitly convert to sparse formats when it makes sense for memory or performance.
内容的提问来源于stack exchange,提问作者Henry Thornton

