You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用featuretools的get_pandas_data_slice?求示例与参数解析

Using featuretools.EntitySet.get_pandas_data_slice(): Examples & Parameter Deep Dive

Got it, let's break down how to use this method with concrete examples and clear explanations for those tricky parameters you mentioned. I'll start with a full working example, then dive into each parameter in detail.

First, Let's Build a Sample EntitySet

Let's create a simple e-commerce EntitySet with three entities: users, orders, and products. This will give us a concrete foundation to demonstrate the method.

import featuretools as ft
import pandas as pd

# Create sample data
users_df = pd.DataFrame({
    "user_id": [1, 2, 3],
    "name": ["Alice", "Bob", "Charlie"],
    "signup_date": pd.date_range("2023-01-01", periods=3)
})

orders_df = pd.DataFrame({
    "order_id": [101, 102, 103, 104],
    "user_id": [1, 1, 2, 3],
    "product_id": [201, 202, 201, 203],
    "order_date": pd.date_range("2023-01-05", periods=4)
})

products_df = pd.DataFrame({
    "product_id": [201, 202, 203],
    "product_name": ["Laptop", "Phone", "Tablet"],
    "price": [999, 699, 499]
})

# Create EntitySet
es = ft.EntitySet(id="ecommerce")
es = es.entity_from_dataframe(
    entity_id="users",
    dataframe=users_df,
    index="user_id",
    time_index="signup_date"
)
es = es.entity_from_dataframe(
    entity_id="orders",
    dataframe=orders_df,
    index="order_id",
    time_index="order_date"
)
es = es.entity_from_dataframe(
    entity_id="products",
    dataframe=products_df,
    index="product_id"
)

# Add relationships
es = es.add_relationship(ft.Relationship(es["users"]["user_id"], es["orders"]["user_id"]))
es = es.add_relationship(ft.Relationship(es["products"]["product_id"], es["orders"]["product_id"]))

Full Usage Example: Filter the EntitySet to a Subset

Let's say we want to extract all data related to user 1 (Alice). Here's how to use get_pandas_data_slice() to do that:

# Get the subset of data for user 1
subset_data = es.get_pandas_data_slice(
    filter_entity_ids={"users": [1]},
    index_eid="users",
    instances=[1]
)

# Print the result
print("Filtered Users Data:")
print(subset_data["users"])
print("\nFiltered Orders Data:")
print(subset_data["orders"])
print("\nFiltered Products Data:")
print(subset_data["products"])

Expected Output:

Filtered Users Data:
       name signup_date
user_id                 
1      Alice  2023-01-01

Filtered Orders Data:
         user_id  product_id order_date
order_id                               
101            1         201 2023-01-05
102            1         202 2023-01-06

Filtered Products Data:
            product_name  price
product_id                     
201               Laptop    999
202                Phone    699

Now let's break down the key parameters you asked about.

Parameter Deep Dive

1. filter_entity_ids

  • What it does: A dictionary that explicitly defines which rows to keep for specific entities. The keys are entity IDs (like "users"), and the values are lists of index values (like [1]) to retain for that entity.
  • When to use: Use this if you want to directly filter one or more entities, regardless of relationships. For example, if you only wanted to keep products with IDs 201 and 203, you could add {"products": [201, 203]} to this dict.
  • Note: If you specify this alongside index_eid and instances, the method will combine the filters—only rows that meet both the relationship-based filter (from index_eid/instances) and the explicit filter_entity_ids will be kept.

2. index_eid

  • What it does: The ID of the "root" entity you want to use as the starting point for filtering. All related entities will be filtered to only include rows connected to the instances specified in instances.
  • When to use: This is essential when you want to extract a subset of data based on relationships (e.g., "all orders and products related to user 1"). The method will traverse the EntitySet's relationships to pull in all connected data.
  • Example: In our sample, setting index_eid="users" tells Featuretools to start from the users entity and pull in all linked orders and products.

3. instances

  • What it does: A list of index values from the index_eid entity that you want to retain. These are the specific "root" instances that the filtered subset will revolve around.
  • When to use: Pair this with index_eid to define exactly which root entities to include. For example, if we wanted data for both user 1 and user 2, we'd set instances=[1, 2].
  • Note: This parameter only works when index_eid is specified—if you don't set index_eid, instances will be ignored.

Additional Tip: Combining Filters

You can mix filter_entity_ids with index_eid/instances for more precise control. For example, if we want user 1's data but only keep orders from January 6th onwards:

# First, get the order IDs from user 1 that are on or after 2023-01-06
filtered_order_ids = orders_df[(orders_df["user_id"] == 1) & (orders_df["order_date"] >= "2023-01-06")]["order_id"].tolist()

# Use combined filters
subset_data = es.get_pandas_data_slice(
    filter_entity_ids={"orders": filtered_order_ids},
    index_eid="users",
    instances=[1]
)

print(subset_data["orders"])

Output:

user_id  product_id order_date
order_id                               
102            1         202 2023-01-06

内容的提问来源于stack exchange,提问作者Nick Bernini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:12:56