基于条件的Pandas-Numpy Select关联 多语言字符串拼写清洗问题
问题排查与修正方案
核心问题原因
- 数据类型不匹配:你示例中
new_df['Country']字段为整数类型,而条件判断中使用了字符串格式的数值(比如'1'/'2'),整数与字符串的相等判断永远为False,所有条件均未命中,最终返回默认值None导致生成空列。 - 执行逻辑冗余:你预先对全列执行了6次
apply计算,无论行对应的Country是什么,都重复计算了全量数据的6种语种匹配,数据量稍大时会严重浪费性能。
修正方案
方案1:优化写法(推荐)
通过词表映射逐行处理,逻辑更清晰、执行效率更高:
import pandas as pd from nltk.corpus import words # 提前加载好各语种词表,此处仅为示例占位 # list_fr = 法语词表列表 # list_de = 德语词表列表 # list_it = 意大利语词表列表 # list_es = 西班牙语词表列表 # list_se = 瑞典语词表列表 # 建立Country数值与对应语种词表的映射 country_word_map = { 1: words.words(), 2: list_fr, 3: list_de, 4: list_it, 5: list_es, 6: list_se, 12: words.words() } # 定义逐行清洗函数 def clean_text(row): country = row['Country'] text = row['OneCol'].strip() if country not in country_word_map: return None # 提前将词表转为集合可大幅提升匹配速度 valid_word_set = set(country_word_map[country]) input_words = text.split() # 如需忽略大小写匹配,可改为 w.lower() in valid_word_set matched_words = [w for w in input_words if w in valid_word_set] return ' '.join(matched_words) # 应用到DataFrame生成结果列 new_df['ResultG'] = new_df.apply(clean_text, axis=1)
方案2:原有np.select写法修复
仅需将条件中的字符串数值改为整数,即可正常运行:
# 去掉条件中数值的引号,改为整数匹配 conditions = [new_df["Country"].isin([1,2,12]), new_df["Country"]==3, new_df["Country"]==4, new_df["Country"]==5, new_df["Country"]==6, new_df["Country"]==10] applied_actions = [new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(words.words())))), new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_fr)))), new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_de)))), new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_it)))), new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_es)))), new_df['OneCol'].apply(lambda row: ' '.join(list(set(row.split()) & set(list_se))))] new_df['ResultG'] = np.select(conditions, applied_actions, default=None)
注意事项
如果词表和输入文本的大小写不统一,会导致匹配失败,建议提前把词表全部转为小写,匹配时把输入单词也转为小写判断,避免漏匹配。
内容的提问来源于stack exchange,提问作者Qthry
相关产品推荐
相关产品推荐

