如何判断DataFrame列是否包含特定列表(忽略顺序)
Got it, here's a straightforward way to solve this problem—since order doesn't matter, sets are the ideal tool here because they ignore both order and duplicate values.
Step 1: Set Up Your DataFrame
First, let's confirm we're working with the same data you provided:
import pandas as pd df = pd.DataFrame() df['Col1'] = [['B'],['A','D','B'],['D','C']] df['Col2'] = [1,2,4]
Step 2: Define Your Target Elements as a Set
Convert your target list [B,A,D] into a set to enable order-agnostic checks:
target = {'B', 'A', 'D'}
Step 3: Check for Matching Rows
We'll iterate over each row in Col1, convert the row's list to a set, and use issubset() to see if our target set is fully contained within the row's set. The any() function will tell us if at least one row meets this condition:
has_all_elements = any(target.issubset(set(row)) for row in df['Col1']) print(has_all_elements) # Output: True
This works because the second row's list ['A','D','B'] converts to the same set as our target—so target.issubset(...) returns True for that row, and any() correctly identifies that there's at least one matching row.
Bonus: Add a Column for Row-Level Checks
If you want to see which individual rows contain all target elements, you can add a new column to the DataFrame using apply():
df['Has_Target_Elements'] = df['Col1'].apply(lambda x: target.issubset(set(x)))
This will give you:
Col1 Col2 Has_Target_Elements 0 [B] 1 False 1 [A, D, B] 2 True 2 [D, C] 4 False
You can still get the overall result with df['Has_Target_Elements'].any().
内容的提问来源于stack exchange,提问作者Ewdlam

