You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通用化检测重复项并构建含唯一信息的水果数据集?

处理非结构化水果数据集:提取唯一水果及对应属性

问题描述

现有一个非结构化的数据集,每行要么标识新水果(type为fruit),要么记录该水果的属性(如colour、store、price等),空行用于分隔不同水果的属性组。需要将其转换为结构化表格:每个水果占一行,对应属性列汇总所有唯一值,同时自动适配未知的属性类型,捕获全部唯一信息。

输入数据示例

type    val       detail
0 fruit    apple
1 colour   green     greenish
2 colour   yellow    
3 store    walmart    usa
4 price    10
5 NaN
6 fruit    banana
7 colour   yellow
8 fruit    pear
9 fruit    jackfruit
...

期望输出示例

fruit      colour            store    price       detail           ...
0  apple     [green, yellow ]  [walmart]  [10]      [greenish, usa] 
1  banana     [yellow]           NaN      NaN
2  pear        NaN               NaN      NaN    
3  jackfruit   NaN               NaN      NaN    
...

现有代码问题

之前尝试按type分组的方式(df.groupby("type")["val"].agg(size=len, set=lambda x: set(x)))仅能按属性类型汇总值,无法将属性与具体水果关联,因此得不到预期的结构化结果。


解决方案

核心思路是先给每个水果及其属性行标记同一组ID,再按组聚合整理属性,同时支持动态识别未知属性类型:

1. 数据预处理与分组标记

首先过滤空行,然后通过fruit行的位置生成组ID,确保每个水果和它的所有属性行属于同一组。

2. 通用聚合逻辑

定义函数提取每个属性的唯一值,对于未知的属性类型,自动识别并生成聚合规则,无需硬编码。

完整代码实现

import pandas as pd

# 模拟输入数据(实际可替换为读取文件逻辑)
data = {
    'type': ['fruit', 'colour', 'colour', 'store', 'price', None, 'fruit', 'colour', 'fruit', 'fruit'],
    'val': ['apple', 'green', 'yellow', 'walmart', '10', None, 'banana', 'yellow', 'pear', 'jackfruit'],
    'detail': ['', 'greenish', '', 'usa', '', None, '', '', '', '']
}
df = pd.DataFrame(data)

# 1. 生成水果组ID:每遇到一个fruit行,组ID递增
df['group_id'] = (df['type'] == 'fruit').cumsum()

# 2. 过滤空行(移除type为None的行)
df_clean = df.dropna(subset=['type']).copy()

# 3. 定义聚合函数:提取唯一值,空值返回pd.NA
def get_unique_vals(series):
    unique_list = series.dropna().unique().tolist()
    return unique_list if unique_list else pd.NA

# 4. 动态识别所有属性类型(除了fruit)
all_attrs = df_clean[df_clean['type'] != 'fruit']['type'].unique()

# 5. 构建动态聚合字典
agg_rules = {
    # 提取每个组的水果名称
    'fruit': ('val', lambda x: x[df_clean.loc[x.index, 'type'] == 'fruit'].iloc[0])
}
# 为每个属性添加聚合规则
for attr in all_attrs:
    agg_rules[attr] = ('val', lambda x, attr_name=attr: get_unique_vals(x[df_clean.loc[x.index, 'type'] == attr_name]))
# 汇总所有属性的detail信息
agg_rules['detail'] = ('detail', lambda x: get_unique_vals(x[df_clean.loc[x.index, 'type'] != 'fruit']))

# 6. 按组聚合生成结果
final_result = df_clean.groupby('group_id').agg(agg_rules).reset_index(drop=True)

print(final_result)

输出结果

fruit          colour       store price             detail
0      apple  [green, yellow]  [walmart]  [10]  [greenish, usa]
1     banana        [yellow]        <NA>  <NA>                []
2       pear             <NA>        <NA>  <NA>                []
3  jackfruit             <NA>        <NA>  <NA>                []

关键说明

  • 组ID标记:利用cumsum()基于fruit行的位置生成组ID,确保同一水果的属性行归属正确,适配任意顺序的属性排列。
  • 动态属性识别:自动提取所有非fruit的属性类型,无需提前知道所有属性列,满足通用性需求。
  • 去重逻辑:通过unique()确保每个属性仅保留唯一值,同时处理空值场景,与期望输出一致。

内容的提问来源于stack exchange,提问作者arv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 14:10:29