You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何确定DFS及FeatureTools中Deep Feature Synthesis使用的原语列表?

How to Choose Primitives for Deep Feature Synthesis (DFS) & Implement This in FeatureTools

Great question! Picking the right primitives for Deep Feature Synthesis (DFS) is crucial—they directly shape the quality of the features you generate, and choosing wisely can save you from sifting through noise while capturing the signals your model needs. Let’s break this down into two clear parts:


1. How to Determine Which Primitives to Use

There’s no universal list, but these guidelines will help you narrow it down based on your data and problem:

  • Start with data types and entity relationships

    • For numeric fields: Use aggregation primitives like sum, mean, median, max, min, or transform primitives like rolling_mean (for time-series data).
    • For categorical fields: Go with count_unique, mode, or percent_true (if binary).
    • For time fields: Use time_since_last, days_in_month, or year to extract temporal signals.
    • For entity relationships (e.g., one-to-many between users and transactions): Use aggregation primitives to roll up child entity data to the parent (e.g., total spend per user).
  • Align with business logic
    Think about what features would matter for your use case. For example:

    • In e-commerce: "Total spend in the last 30 days" is more actionable than "lifetime total spend" for predicting near-term purchases.
    • In fraud detection: "Number of failed login attempts in the last hour" is a stronger signal than total failed attempts ever.
  • Start simple, then iterate
    Don’t overload your DFS with every primitive at once. Begin with core primitives (sum, count, mean, mode) and evaluate model performance. Then add more specific primitives (like skew, kurtosis, or time_since_first) based on feature importance scores.

  • Avoid redundant primitives
    If you already have sum and count for a field, mean is just sum/count—you might not need both unless your model explicitly benefits from redundant signals. Prioritize features that add unique information.

  • Balance performance and interpretability
    Complex primitives (like custom aggregations) might boost model accuracy, but if you need explainable results, stick to simpler, easier-to-understand primitives.


2. How to Select Primitives in FeatureTools

FeatureTools gives you full control over which primitives to use. Here’s how to implement your selection:

List available primitives first

Before choosing, check what’s built-in with this command:

import featuretools as ft

# List all aggregation primitives
print(ft.list_primitives("aggregation"))

# List all transform primitives
print(ft.list_primitives("transform"))

Manually specify primitives in dfs()

Use the agg_primitives and trans_primitives parameters to pass your chosen list:

# Define your target primitives
agg_primitives = ["sum", "mean", "count_unique", "max"]
trans_primitives = ["year", "month", "time_since_previous"]

# Run DFS with your selected primitives
features, feature_defs = ft.dfs(
    dataframes=your_dataframes_dict,
    target_dataframe_name="target_entity",
    agg_primitives=agg_primitives,
    trans_primitives=trans_primitives,
    verbose=True
)

Exclude unwanted primitives

If you want to use most built-in primitives except a few, use exclude_agg_primitives or exclude_trans_primitives:

features, feature_defs = ft.dfs(
    dataframes=your_dataframes_dict,
    target_dataframe_name="target_entity",
    exclude_agg_primitives=["mode"],  # Skip mode if rare categories are common
    verbose=True
)

Add custom primitives

If built-in primitives don’t cover your use case, define your own and include them in the list:

from featuretools.primitives import AggregationPrimitive
from featuretools.variable_types import Numeric

# Custom aggregation to calculate average transaction amount
class AvgTransaction(AggregationPrimitive):
    name = "avg_transaction"
    input_types = [Numeric]
    return_type = Numeric

    def get_function(self):
        def avg(x):
            return x.mean()
        return avg

# Use the custom primitive alongside built-ins
agg_primitives = ["sum", AvgTransaction]
features, feature_defs = ft.dfs(
    dataframes=your_dataframes_dict,
    target_dataframe_name="customers",
    agg_primitives=agg_primitives,
    verbose=True
)

Automate feature filtering post-DFS

After generating features, use featuretools.select_features to keep only impactful ones (based on model importance or correlation), which can inform your primitive choices for future runs:

from sklearn.ensemble import RandomForestClassifier
from featuretools.selection import select_features

# Train a quick model to get feature importances
model = RandomForestClassifier()
model.fit(features, target_labels)

# Select top 20 features
selected_features = select_features(features, target_labels, model=model, top_n=20)

Remember, the best primitive list comes from experimentation. Test small sets, iterate based on model performance, and keep your business context front and center.

内容的提问来源于stack exchange,提问作者Harish Rajula

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:43:22