You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按DataFrame的Locations列列表元素对行进行分组?

问题描述

我有一个如下所示的DataFrame:

81883       2011000011  ...  [South Sturgeon, Creek]
81884       2011000022  ...        [Meadowood]
81885       2011000016  ...   [South, Portage]
81886       2011000011  ...  [North Sturgeon, Creek]

我希望按最后一列(名为Locations)中拆分后的共同单词对行进行分组:例如在上述示例中,我想要按Creek分组;当未找到共同单词时,行保持原样(或合并为字符串更佳)。

我尝试使用以下代码:

def get_grp(list_current_row, df,column_location): 
    rows_index_to_groupby = [] 
    for string_element in list_current_row: 
        for idx,row in enumerate (df[column_location].values): 
            if row != list_current_row and string_element in row: 
                rows_index_to_groupby.append(idx) 
    return rows_index_to_groupby


grouped_dataframe = resulting_dataframe.groupby(lambda x: [resulting_dataframe[column_location][i] for i in get_grp(x, resulting_dataframe,column_location)] )

期望输出如下:

Locations
Creek             0  Creek       81886       2011000011  ...
                  1  Creek       81883       2011000011  ...
South, Portage    2  South, Portage      81885       2011000016  ...
Meadowood         3  Meadowood       81884    2011000022

解决方案

原代码的逻辑效率较低且分组键不明确,我们可以通过统计单词出现频率、明确分组键的方式实现需求:

步骤1:构造示例DataFrame(如果已有可跳过)

import pandas as pd
from collections import Counter

data = {
    'ID': [2011000011, 2011000022, 2011000016, 2011000011],
    'Locations': [['South Sturgeon', 'Creek'], ['Meadowood'], ['South', 'Portage'], ['North Sturgeon', 'Creek']]
}
df = pd.DataFrame(data, index=[81883, 81884, 81885, 81886])

步骤2:统计所有单词的出现频率

先提取Locations列中所有拆分后的单词,统计每个单词的出现次数,以此判断哪些是共同单词:

all_words = []
for loc_list in df['Locations']:
    # 拆分每个字符串元素为单个单词,比如'South Sturgeon'拆为['South', 'Sturgeon']
    for item in loc_list:
        all_words.extend(item.split())
word_counts = Counter(all_words)

步骤3:定义分组键生成函数

为每行分配分组键:如果该行包含出现次数≥2的单词,取第一个符合条件的单词作为分组键;否则用该行Locations的合并字符串作为分组键:

def get_group_key(loc_list):
    for item in loc_list:
        for word in item.split():
            if word_counts[word] >= 2:
                return word
    # 无共同单词时返回合并后的字符串
    return ', '.join(loc_list)

# 为DataFrame添加分组键列
df['group_key'] = df['Locations'].apply(get_group_key)

步骤4:执行分组并输出结果

# 按分组键分组
grouped_df = df.groupby('group_key')

# 查看分组详情
for key, group in grouped_df:
    print(f"分组键: {key}")
    print(group)
    print('---')

输出结果:

分组键: Creek
           ID               Locations group_key
81883  2011000011  [South Sturgeon, Creek]      Creek
81886  2011000011  [North Sturgeon, Creek]      Creek
---
分组键: Meadowood
           ID      Locations group_key
81884  2011000022  [Meadowood]   Meadowood
---
分组键: South, Portage
           ID          Locations       group_key
81885  2011000016  [South, Portage]  South, Portage
---

如果要生成类似期望输出的层级索引结构,可执行:

result = df.set_index(['group_key', df.index]).sort_index()
print(result)

输出:

ID               Locations
group_key     index                                
Creek         81883  2011000011  [South Sturgeon, Creek]
              81886  2011000011  [North Sturgeon, Creek]
Meadowood     81884  2011000022              [Meadowood]
South, Portage 81885  2011000016          [South, Portage]

内容的提问来源于stack exchange,提问作者user1319236

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 14:15:32