基于字符串长度与内容匹配的DataFrame条件循环分组实现
Grouping DataFrame Rows by Element Set Membership
Got it, let's work through this grouping problem step by step. First, I'll lock down the grouping rules from your example to ensure we're on the same page:
- Start with Group 1: Find the row(s) in your
STRcolumn with the largest collection of unique elements (split those comma-separated strings and combine all elements from each cell). The full set of elements from this row becomes the reference for Group 1. - Assign remaining rows:
- If a row's entire set of elements is a subset of an existing group's reference set, assign that group number.
- If the row has elements that don't fit into any existing group's reference, create a new group and use its element set as the new reference.
Let's Build the Solution
We'll use pandas and basic set operations to pull this off. First, let's recreate your example DataFrame to test with:
import pandas as pd # Your example input (formatted as parsable string lists) data = { "STR": [ '["G,D,E","F"]', '["D,E,F","G"]', '["D,F","E"]', '["D,E","F"]', '["A,B","C"]', '["C","D"]', '["A","B"]' ] } df = pd.DataFrame(data)
First step: extract the full set of unique elements from each STR cell. We'll use ast.literal_eval to safely parse those stringified lists:
import ast def extract_element_set(cell): # Parse the list string, split each comma-separated part, flatten into a set parts = ast.literal_eval(cell) all_elements = [] for part in parts: all_elements.extend(part.split(",")) return set(all_elements) # Add a helper column with the element set for each row df["element_set"] = df["STR"].apply(extract_element_set)
Now, let's implement the grouping logic:
# First, find the reference set for Group 1 (largest element set) max_set_size = df["element_set"].apply(len).max() group1_reference = df[df["element_set"].apply(len) == max_set_size]["element_set"].iloc[0] # Track groups: key = group number, value = reference element set group_references = {1: group1_reference} group_assignments = [] for _, row in df.iterrows(): current_set = row["element_set"] assigned_group = None # Check if current set fits into any existing group for group_num, ref_set in group_references.items(): if current_set.issubset(ref_set): assigned_group = group_num break # If no match, create a new group if assigned_group is None: assigned_group = max(group_references.keys()) + 1 group_references[assigned_group] = current_set group_assignments.append(assigned_group) # Add the Group column to your DataFrame df["Group"] = group_assignments
Check the Result
If you print df[["STR", "Group"]], you'll get exactly the output you wanted:
| STR | Group |
|---|---|
| ["G,D,E","F"] | 1 |
| ["D,E,F","G"] | 1 |
| ["D,F","E"] | 1 |
| ["D,E","F"] | 1 |
| ["A,B","C"] | 2 |
| ["C","D"] | 3 |
| ["A","B"] | 2 |
Quick Notes
- If your actual
STRcolumn format is different (e.g., not stringified lists), just adjust theextract_element_setfunction to match your input structure. - If you wanted exact set matches instead of subset checks, swap
current_set.issubset(ref_set)withcurrent_set == ref_set—that would change the grouping logic to only match rows with identical element sets.
内容的提问来源于stack exchange,提问作者NiMbuS
相关产品推荐
相关产品推荐

