You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字符串长度与内容匹配的DataFrame条件循环分组实现

Grouping DataFrame Rows by Element Set Membership

Got it, let's work through this grouping problem step by step. First, I'll lock down the grouping rules from your example to ensure we're on the same page:

  1. Start with Group 1: Find the row(s) in your STR column with the largest collection of unique elements (split those comma-separated strings and combine all elements from each cell). The full set of elements from this row becomes the reference for Group 1.
  2. Assign remaining rows:
    • If a row's entire set of elements is a subset of an existing group's reference set, assign that group number.
    • If the row has elements that don't fit into any existing group's reference, create a new group and use its element set as the new reference.

Let's Build the Solution

We'll use pandas and basic set operations to pull this off. First, let's recreate your example DataFrame to test with:

import pandas as pd

# Your example input (formatted as parsable string lists)
data = {
    "STR": [
        '["G,D,E","F"]',
        '["D,E,F","G"]',
        '["D,F","E"]',
        '["D,E","F"]',
        '["A,B","C"]',
        '["C","D"]',
        '["A","B"]'
    ]
}
df = pd.DataFrame(data)

First step: extract the full set of unique elements from each STR cell. We'll use ast.literal_eval to safely parse those stringified lists:

import ast

def extract_element_set(cell):
    # Parse the list string, split each comma-separated part, flatten into a set
    parts = ast.literal_eval(cell)
    all_elements = []
    for part in parts:
        all_elements.extend(part.split(","))
    return set(all_elements)

# Add a helper column with the element set for each row
df["element_set"] = df["STR"].apply(extract_element_set)

Now, let's implement the grouping logic:

# First, find the reference set for Group 1 (largest element set)
max_set_size = df["element_set"].apply(len).max()
group1_reference = df[df["element_set"].apply(len) == max_set_size]["element_set"].iloc[0]

# Track groups: key = group number, value = reference element set
group_references = {1: group1_reference}
group_assignments = []

for _, row in df.iterrows():
    current_set = row["element_set"]
    assigned_group = None
    
    # Check if current set fits into any existing group
    for group_num, ref_set in group_references.items():
        if current_set.issubset(ref_set):
            assigned_group = group_num
            break
    
    # If no match, create a new group
    if assigned_group is None:
        assigned_group = max(group_references.keys()) + 1
        group_references[assigned_group] = current_set
    
    group_assignments.append(assigned_group)

# Add the Group column to your DataFrame
df["Group"] = group_assignments

Check the Result

If you print df[["STR", "Group"]], you'll get exactly the output you wanted:

STRGroup
["G,D,E","F"]1
["D,E,F","G"]1
["D,F","E"]1
["D,E","F"]1
["A,B","C"]2
["C","D"]3
["A","B"]2

Quick Notes

  • If your actual STR column format is different (e.g., not stringified lists), just adjust the extract_element_set function to match your input structure.
  • If you wanted exact set matches instead of subset checks, swap current_set.issubset(ref_set) with current_set == ref_set—that would change the grouping logic to only match rows with identical element sets.

内容的提问来源于stack exchange,提问作者NiMbuS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:35:59