如何移除Pandas数据框列中每个单词末尾的‘br’字母组合?
移除Pandas数据框列中单词末尾的"br"字符
问题背景
需要处理Pandas数据框中的某一列(每行都是独立句子),移除每个单词末尾的"br"字母——这些是之前清理HTML内容时未妥善处理<br>标签导致的无效后缀,比如出现somebr、andbr这类错误词汇。
错误示例
Sentence = here are somebr examples of poorly written paragraphs andbr well-written paragraphsbr on the same topicbr how do they compare?
期望结果
Sentence: here are some examples of poorly written and well-written paragraphs on the same topic how do they compare?
核心要求
- 仅移除单词末尾的"br",不影响句子结构和其他正常单词
- 保留
brutish、breathtaking、ember这类中间含"br"的正常词汇 - 无需要保留的末尾为"br"的单词
解决方案
使用Pandas的str.replace()方法配合正则表达式,精准匹配单词末尾的"br"并替换为空,避免误改正常词汇。
代码实现
import pandas as pd # 假设数据框为df,目标列名为'Sentence' df['Sentence'] = df['Sentence'].str.replace(r'(\w+)br\b', r'\1', regex=True)
正则逻辑说明
(\w+):捕获单词的主体部分(一个或多个字母/数字/下划线),通过括号分组便于后续引用br:精准匹配单词末尾的"br"后缀\b:单词边界标记,确保"br"处于单词结尾位置,不会匹配到brutish这类中间含"br"的单词\1:替换为之前捕获的单词主体,即去掉末尾"br"后的内容
测试验证
用示例句子测试:
test_str = "here are somebr examples of poorly written paragraphs andbr well-written paragraphsbr on the same topicbr how do they compare? Also, brutish and breathtaking should stay unchanged." cleaned_str = pd.Series([test_str]).str.replace(r'(\w+)br\b', r'\1', regex=True)[0] print(cleaned_str)
输出结果:
here are some examples of poorly written paragraphs and well-written paragraphs on the same topic how do they compare? Also, brutish and breathtaking should stay unchanged.
可见错误后缀被移除,正常含"br"的单词未受影响,完全符合需求。
内容的提问来源于stack exchange,提问作者user15516822
相关产品推荐
相关产品推荐

