You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按组生成不区分大小写的行内容包含标记列并灵活删除标记行?

解决方案

数据准备

先将输入数据转换为pandas DataFrame:

import pandas as pd

data = {
    'Type': ['Fruit', 'Fruit', 'Fruit', 'Fruit', 'Cutlery'],
    'value': ['apple', 'App le', 'Apple yes', 'Apple', 'Spoon']
}
df = pd.DataFrame(data)

生成dup和dup_index列

核心思路是先统一转换为小写实现不区分大小写的判断,再按Type分组,逐行检查同组内其他行的value是否包含当前行内容:

# 生成小写版本的value,用于不区分大小写匹配
df['lower_value'] = df['value'].str.lower()

# 初始化结果列
df['dup'] = ''
df['dup_index'] = ''

# 按Type分组处理每组数据
for _, group in df.groupby('Type'):
    group_indices = group.index.tolist()
    group_lower_vals = group['lower_value'].tolist()
    
    for idx, current_val in zip(group_indices, group_lower_vals):
        # 筛选同组中其他行里包含当前值的索引
        matched = [i for i, val in zip(group_indices, group_lower_vals) if i != idx and current_val in val]
        
        if matched:
            df.loc[idx, 'dup'] = 'yes'
            df.loc[idx, 'dup_index'] = ','.join(map(str, matched))

# 清理临时列
df = df.drop('lower_value', axis=1)

执行后得到的结果如下:

Type       value  dup dup_index
0     Fruit       apple  yes       2,3
1     Fruit      App le              
2     Fruit    Apple yes  yes         
3     Fruit      Apple  yes       0,2
4  Cutlery       Spoon              

删除dup列为"yes"的行

只需简单过滤即可实现:

# 过滤掉dup为"yes"的行
filtered_df = df[df['dup'] != 'yes'].reset_index(drop=True)

过滤后的结果:

Type    value dup dup_index
0     Fruit   App le              
1  Cutlery    Spoon              

内容的提问来源于stack exchange,提问作者asd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 04:49:59