Pandas中如何将替换后的复合词拆分为列表独立元素
解决DataFrame中错误词替换并拆分单词的问题
示例数据构造
首先根据你提供的结构创建两个DataFrame:
import pandas as pd # 创建colloquial DataFrame colloquial = pd.DataFrame({ 'wrong': ['sheis', 'taht', 'diedwhen'], 'correct': ['she is', 'that', 'died when'] }) # 创建dataset DataFrame dataset = pd.DataFrame({ 'review': [ ['shewas', 'wrong', 'but', 'her', 'intentions', 'are', 'good'], ['is', 'taht', 'you'], ['he', 'diedwhen', 'he', 'was', 'young'] ] })
核心处理步骤
- 构建错误词-正确词映射字典
把colloquial的两列转换成字典,方便快速查找替换:
correction_map = colloquial.set_index('wrong')['correct'].to_dict()
- 定义处理函数
这个函数会遍历每个评论的单词,替换错误词后,将带空格的正确词拆分为独立单词,最后合并成新的单词列表:
def fix_and_split_review(review_words): fixed_words = [] for word in review_words: # 替换错误词,无匹配则保留原词 corrected_str = correction_map.get(word, word) # 拆分带空格的字符串为独立单词,加入结果列表 fixed_words.extend(corrected_str.split()) return fixed_words
- 应用函数到dataset的review列
用apply方法批量处理所有评论:
dataset['fixed_review'] = dataset['review'].apply(fix_and_split_review)
结果展示
处理后的dataset:
| review | fixed_review |
|---|---|
| ['shewas', 'wrong', 'but', 'her', 'intentions', 'are', 'good'] | ['shewas', 'wrong', 'but', 'her', 'intentions', 'are', 'good'] |
| ['is', 'taht', 'you'] | ['is', 'that', 'you'] |
| ['he', 'diedwhen', 'he', 'was', 'young'] | ['he', 'died', 'when', 'he', 'was', 'young'] |
说明
- 对于无匹配的错误词(比如示例中的
shewas),会直接保留原词 - 替换后带空格的字符串(比如
diedwhen替换成died when)会自动拆分成两个独立单词,合并到结果列表中
内容的提问来源于stack exchange,提问作者andryan86
相关产品推荐
相关产品推荐

