You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件的Pandas-Numpy Select关联 多语言字符串拼写清洗问题

问题排查与修正方案

核心问题原因

  • 数据类型不匹配:你示例中new_df['Country']字段为整数类型,而条件判断中使用了字符串格式的数值(比如'1'/'2'),整数与字符串的相等判断永远为False,所有条件均未命中,最终返回默认值None导致生成空列。
  • 执行逻辑冗余:你预先对全列执行了6次apply计算,无论行对应的Country是什么,都重复计算了全量数据的6种语种匹配,数据量稍大时会严重浪费性能。

修正方案

方案1:优化写法(推荐)

通过词表映射逐行处理,逻辑更清晰、执行效率更高:

import pandas as pd
from nltk.corpus import words

# 提前加载好各语种词表,此处仅为示例占位
# list_fr = 法语词表列表
# list_de = 德语词表列表
# list_it = 意大利语词表列表
# list_es = 西班牙语词表列表
# list_se = 瑞典语词表列表

# 建立Country数值与对应语种词表的映射
country_word_map = {
    1: words.words(),
    2: list_fr,
    3: list_de,
    4: list_it,
    5: list_es,
    6: list_se,
    12: words.words()
}

# 定义逐行清洗函数
def clean_text(row):
    country = row['Country']
    text = row['OneCol'].strip()
    if country not in country_word_map:
        return None
    # 提前将词表转为集合可大幅提升匹配速度
    valid_word_set = set(country_word_map[country])
    input_words = text.split()
    # 如需忽略大小写匹配,可改为 w.lower() in valid_word_set
    matched_words = [w for w in input_words if w in valid_word_set]
    return ' '.join(matched_words)

# 应用到DataFrame生成结果列
new_df['ResultG'] = new_df.apply(clean_text, axis=1)

方案2:原有np.select写法修复

仅需将条件中的字符串数值改为整数,即可正常运行:

# 去掉条件中数值的引号,改为整数匹配
conditions = [new_df["Country"].isin([1,2,12]), 
              new_df["Country"]==3, 
              new_df["Country"]==4,  
              new_df["Country"]==5, 
              new_df["Country"]==6, 
              new_df["Country"]==10]

applied_actions =  [new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(words.words())))), 
                    new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_fr)))), 
                    new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_de)))),
                    new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_it)))),
                    new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_es)))),
                    new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_se))))]

new_df['ResultG'] = np.select(conditions, applied_actions, default=None)

注意事项

如果词表和输入文本的大小写不统一,会导致匹配失败,建议提前把词表全部转为小写,匹配时把输入单词也转为小写判断,避免漏匹配。

内容的提问来源于stack exchange,提问作者Qthry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 09:48:03