You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法排查Pandas字符串过滤代码错误:移除子串型短字符串问题

问题描述

我有一个Pandas DataFrame的dtc_mined列,列中值以|分隔,示例值如下:

P18A253|P18A0|P18A2|P18A043|P2B61

其中包含长度为5的字符串(如P18A2)和长度为7的字符串(如P18A043)。我的需求是:若某一5长度字符串是任意7长度字符串的子串,则移除该5长度字符串,期望输出为:

P18A253|P18A043|P2B61

我尝试了以下两段代码,但无法找出错误所在:

第一段尝试代码

import pandas as pd

# Sample DataFrame
data = {'dtc_mined': ['P18A253|P18A0|P18A2|P18A043|P2B61']}
df = pd.DataFrame(data)

# Split the values and create sets of 5 and 7 character words
df['split_values'] = df['dtc_mined'].str.split('|')
df['words_5'] = df['split_values'].apply(lambda lst: set(word for word in lst if len(word) == 5))
df['words_7'] = df['split_values'].apply(lambda lst: set(word for word in lst if len(word) == 7))

# Remove 5-character words that have a corresponding 7-character word
df['filtered_values'] = df.apply(lambda row: '|'.join(word for word in row['split_values'] if len(word) == 7 or word not in row['words_7']), axis=1)

# Drop intermediate columns and display the result
result = df.drop(['split_values', 'words_5', 'words_7'], axis=1)
print(result)

第二段尝试代码

# Remove 5-character words that have a corresponding 7-character word

def Check1(row):
    for word in row['words_5']:
        if word not in row['words_7']:
            row['words_7'].add(word)
    return row['words_7']

df['filtered_values'] = df.apply(Check1, axis=1)
错误分析

第一段代码的问题

核心逻辑错误:代码中用word not in row['words_7']判断是否保留5长度字符串,这是在检查5长度字符串是否等于某个7长度字符串,而非是否是子串。比如P18A2是P18A253的子串,但它不等于任何7长度字符串,所以这段代码会错误保留它,导致结果不符合预期。

第二段代码的问题

逻辑完全偏离需求:这段代码遍历5长度字符串,把不在7长度集合里的5长度字符串添加到7长度集合中,最后返回这个混合集合。它根本没有实现“移除是7长度子串的5长度字符串”的逻辑,反而把不需要的5长度字符串也混入结果。

正确实现方案

我们需要先识别出所有是任意7长度字符串子串的5长度字符串,再过滤掉这些字符串。代码如下:

import pandas as pd

# 示例DataFrame
data = {'dtc_mined': ['P18A253|P18A0|P18A2|P18A043|P2B61']}
df = pd.DataFrame(data)

def filter_dtc(row):
    split_vals = row['dtc_mined'].split('|')
    # 提取所有7长度的字符串
    words_7 = [w for w in split_vals if len(w) == 7]
    # 找出所有需要移除的5长度字符串
    to_remove = set()
    for w5 in [w for w in split_vals if len(w) == 5]:
        for w7 in words_7:
            if w5 in w7:
                to_remove.add(w5)
                break  # 匹配到一个就停止检查,提升效率
    # 过滤并拼接结果
    filtered = [w for w in split_vals if len(w) ==7 or w not in to_remove]
    return '|'.join(filtered)

df['filtered_values'] = df.apply(filter_dtc, axis=1)
print(df[['dtc_mined', 'filtered_values']])

运行后,filtered_values列会输出预期结果:P18A253|P18A043|P2B61

内容的提问来源于stack exchange,提问作者Jawed Sheikh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 06:35:13