You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中如何将替换后的复合词拆分为列表独立元素

解决DataFrame中错误词替换并拆分单词的问题

示例数据构造

首先根据你提供的结构创建两个DataFrame:

import pandas as pd

# 创建colloquial DataFrame
colloquial = pd.DataFrame({
    'wrong': ['sheis', 'taht', 'diedwhen'],
    'correct': ['she is', 'that', 'died when']
})

# 创建dataset DataFrame
dataset = pd.DataFrame({
    'review': [
        ['shewas', 'wrong', 'but', 'her', 'intentions', 'are', 'good'],
        ['is', 'taht', 'you'],
        ['he', 'diedwhen', 'he', 'was', 'young']
    ]
})

核心处理步骤

  1. 构建错误词-正确词映射字典
    把colloquial的两列转换成字典,方便快速查找替换:
correction_map = colloquial.set_index('wrong')['correct'].to_dict()
  1. 定义处理函数
    这个函数会遍历每个评论的单词,替换错误词后,将带空格的正确词拆分为独立单词,最后合并成新的单词列表:
def fix_and_split_review(review_words):
    fixed_words = []
    for word in review_words:
        # 替换错误词,无匹配则保留原词
        corrected_str = correction_map.get(word, word)
        # 拆分带空格的字符串为独立单词,加入结果列表
        fixed_words.extend(corrected_str.split())
    return fixed_words
  1. 应用函数到dataset的review列
    用apply方法批量处理所有评论:
dataset['fixed_review'] = dataset['review'].apply(fix_and_split_review)

结果展示

处理后的dataset:

reviewfixed_review
['shewas', 'wrong', 'but', 'her', 'intentions', 'are', 'good']['shewas', 'wrong', 'but', 'her', 'intentions', 'are', 'good']
['is', 'taht', 'you']['is', 'that', 'you']
['he', 'diedwhen', 'he', 'was', 'young']['he', 'died', 'when', 'he', 'was', 'young']

说明

  • 对于无匹配的错误词(比如示例中的shewas),会直接保留原词
  • 替换后带空格的字符串(比如diedwhen替换成died when)会自动拆分成两个独立单词,合并到结果列表中

内容的提问来源于stack exchange,提问作者andryan86

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 22:46:02