如何使用featuretools的get_pandas_data_slice?求示例与参数解析
featuretools.EntitySet.get_pandas_data_slice(): Examples & Parameter Deep Dive Got it, let's break down how to use this method with concrete examples and clear explanations for those tricky parameters you mentioned. I'll start with a full working example, then dive into each parameter in detail.
First, Let's Build a Sample EntitySet
Let's create a simple e-commerce EntitySet with three entities: users, orders, and products. This will give us a concrete foundation to demonstrate the method.
import featuretools as ft import pandas as pd # Create sample data users_df = pd.DataFrame({ "user_id": [1, 2, 3], "name": ["Alice", "Bob", "Charlie"], "signup_date": pd.date_range("2023-01-01", periods=3) }) orders_df = pd.DataFrame({ "order_id": [101, 102, 103, 104], "user_id": [1, 1, 2, 3], "product_id": [201, 202, 201, 203], "order_date": pd.date_range("2023-01-05", periods=4) }) products_df = pd.DataFrame({ "product_id": [201, 202, 203], "product_name": ["Laptop", "Phone", "Tablet"], "price": [999, 699, 499] }) # Create EntitySet es = ft.EntitySet(id="ecommerce") es = es.entity_from_dataframe( entity_id="users", dataframe=users_df, index="user_id", time_index="signup_date" ) es = es.entity_from_dataframe( entity_id="orders", dataframe=orders_df, index="order_id", time_index="order_date" ) es = es.entity_from_dataframe( entity_id="products", dataframe=products_df, index="product_id" ) # Add relationships es = es.add_relationship(ft.Relationship(es["users"]["user_id"], es["orders"]["user_id"])) es = es.add_relationship(ft.Relationship(es["products"]["product_id"], es["orders"]["product_id"]))
Full Usage Example: Filter the EntitySet to a Subset
Let's say we want to extract all data related to user 1 (Alice). Here's how to use get_pandas_data_slice() to do that:
# Get the subset of data for user 1 subset_data = es.get_pandas_data_slice( filter_entity_ids={"users": [1]}, index_eid="users", instances=[1] ) # Print the result print("Filtered Users Data:") print(subset_data["users"]) print("\nFiltered Orders Data:") print(subset_data["orders"]) print("\nFiltered Products Data:") print(subset_data["products"])
Expected Output:
Filtered Users Data: name signup_date user_id 1 Alice 2023-01-01 Filtered Orders Data: user_id product_id order_date order_id 101 1 201 2023-01-05 102 1 202 2023-01-06 Filtered Products Data: product_name price product_id 201 Laptop 999 202 Phone 699
Now let's break down the key parameters you asked about.
Parameter Deep Dive
1. filter_entity_ids
- What it does: A dictionary that explicitly defines which rows to keep for specific entities. The keys are entity IDs (like
"users"), and the values are lists of index values (like[1]) to retain for that entity. - When to use: Use this if you want to directly filter one or more entities, regardless of relationships. For example, if you only wanted to keep products with IDs 201 and 203, you could add
{"products": [201, 203]}to this dict. - Note: If you specify this alongside
index_eidandinstances, the method will combine the filters—only rows that meet both the relationship-based filter (fromindex_eid/instances) and the explicitfilter_entity_idswill be kept.
2. index_eid
- What it does: The ID of the "root" entity you want to use as the starting point for filtering. All related entities will be filtered to only include rows connected to the instances specified in
instances. - When to use: This is essential when you want to extract a subset of data based on relationships (e.g., "all orders and products related to user 1"). The method will traverse the EntitySet's relationships to pull in all connected data.
- Example: In our sample, setting
index_eid="users"tells Featuretools to start from theusersentity and pull in all linked orders and products.
3. instances
- What it does: A list of index values from the
index_eidentity that you want to retain. These are the specific "root" instances that the filtered subset will revolve around. - When to use: Pair this with
index_eidto define exactly which root entities to include. For example, if we wanted data for both user 1 and user 2, we'd setinstances=[1, 2]. - Note: This parameter only works when
index_eidis specified—if you don't setindex_eid,instanceswill be ignored.
Additional Tip: Combining Filters
You can mix filter_entity_ids with index_eid/instances for more precise control. For example, if we want user 1's data but only keep orders from January 6th onwards:
# First, get the order IDs from user 1 that are on or after 2023-01-06 filtered_order_ids = orders_df[(orders_df["user_id"] == 1) & (orders_df["order_date"] >= "2023-01-06")]["order_id"].tolist() # Use combined filters subset_data = es.get_pandas_data_slice( filter_entity_ids={"orders": filtered_order_ids}, index_eid="users", instances=[1] ) print(subset_data["orders"])
Output:
user_id product_id order_date order_id 102 1 202 2023-01-06
内容的提问来源于stack exchange,提问作者Nick Bernini

