如何使用Pandas筛选包含列表中任一子串的DataFrame行?
Hey there! The problem with your current code is that pandas.Series.str.contains() expects a regular expression pattern string as input, not a list directly. When you pass a list to it, pandas doesn’t know how to check for any of the substrings—so let’s fix that with a quick adjustment.
Step 1: Convert Your List to a Regex Pattern
First, we’ll join your product list into a single regex pattern where each substring is separated by | (this acts as an OR operator in regex). This tells pandas to match any row where the column contains one or more of the substrings in your list.
Step 2: Run the Corrected Filter
Here’s the fixed code—note I adjusted the column name to match your sample data (Product instead of 'Goods Shipped'; double-check this matches your actual DataFrame column name!):
import pandas as pd # Sample data matching your example data = { 'Name': ['David', 'Meghan', 'Melanie', 'Aaron', 'Venus', 'Abigail', 'Sophia'], 'Product': ['PLASTIC BOTTLE', 'PLASTIC COVER', 'PLASTIC CUP', 'PLASTIC BOWL', 'PLASTIC KNIFE', 'PLASTIC CONTAINER', 'PLASTIC LID'] } df_plastic = pd.DataFrame(data) product = ['LID', 'TABLEWARE', 'CUP', 'COVER', 'CONTAINER', 'PACKAGING'] # Create regex pattern: "LID|TABLEWARE|CUP|COVER|CONTAINER|PACKAGING" pattern = '|'.join(product) # Filter rows where Product contains any substring from the pattern df_plastic_prod = df_plastic[df_plastic['Product'].str.contains(pattern)] # Verify the result print(df_plastic_prod) df_plastic_prod.info()
Expected Output
Running this code will give you the filtered DataFrame you’re looking for:
Name Product 1 Meghan PLASTIC COVER 2 Melanie PLASTIC CUP 5 Abigail PLASTIC CONTAINER 6 Sophia PLASTIC LID
Bonus: Case-Insensitive Matching
If you want to match substrings regardless of uppercase/lowercase (e.g., catch plastic lid or Plastic Cover), add the case=False parameter to str.contains():
df_plastic_prod = df_plastic[df_plastic['Product'].str.contains(pattern, case=False)]
内容的提问来源于stack exchange,提问作者Matthias Gallagher

