如何在Pandas DataFrame中选择性移除类似'1)'的字符串组合?
移除Pandas DataFrame中"数字+右括号"类序号标记
要精准识别并移除"1)"、"2)"这类组合,核心是匹配一个或多个数字紧跟右括号的模式,用正则表达式就能搞定,不会误删其他数字内容。
核心实现代码
假设你的DataFrame目标列名为content,直接用str.replace结合正则替换:
import pandas as pd # 示例DataFrame df = pd.DataFrame({ 'content': ['"1) some text WH-1162" some words: 1011,4; 2) some other text: 1 pc; 3) CBHU8512454, number:2; 8) Code:000;'] }) # 替换掉所有"数字+右括号"的组合 df['content'] = df['content'].str.replace(r'\d+\)', '', regex=True)
效果说明
处理前的内容:
"1) some text WH-1162" some words: 1011,4; 2) some other text: 1 pc; 3) CBHU8512454, number:2; 8) Code:000;
处理后的内容:
" some text WH-1162" some words: 1011,4; some other text: 1 pc; CBHU8512454, number:2; Code:000;
扩展适配
如果序号前有空格(比如" 8)"),把正则改成r'\s*\d+\)',可以连前面的空格一起移除:
df['content'] = df['content'].str.replace(r'\s*\d+\)', '', regex=True)
内容的提问来源于stack exchange,提问作者Madina
相关产品推荐
相关产品推荐

