使用正则表达式删除Pandas DataFrame中的指定索引结构子串
在Pandas DataFrame中批量删除包含endIndex/startIndex/value的结构化子串
嘿,我来帮你搞定这个问题!你需要在Pandas的DataFrame里批量删掉那些带endIndex、startIndex、value的结构化子串,不管是单独的{endIndex:...}块还是嵌套在其他结构里的,对吧?我整理了一个实用的方案,一步步来:
1. 构造匹配目标结构的正则表达式
首先得写一个能精准匹配这类结构的正则。考虑到可能有嵌套(比如你提到的date-and-time:{city:{endIndex:...}}),我用了支持递归的正则来处理嵌套的JSON对象,确保能把包含三个目标字段的整个结构都匹配到:
import re # 匹配同时包含endIndex、startIndex、value的JSON对象(支持嵌套) pattern = r'\{(?:[^{}]|(?R))*\bendIndex\b(?:[^{}]|(?R))*\bstartIndex\b(?:[^{}]|(?R))*\bvalue\b(?:[^{}]|(?R))*\}'
如果你的场景里这些字段都是按endIndex→startIndex→value的顺序出现,且嵌套不复杂,也可以用更简单的非递归版本,运行更快:
# 匹配顺序固定的非嵌套结构 simple_pattern = r'\{endIndex:\s*-?\d+,\s*startIndex:\s*-?\d+,\s*value:[^}]*\}'
2. 在Pandas中批量替换
接下来就是把这个正则应用到DataFrame的字符串列上。你可以选择处理所有字符串列,或者只针对特定列(比如你的Json列):
处理所有字符串列
import pandas as pd # 假设你的DataFrame叫df # 获取所有字符串类型的列 str_columns = df.select_dtypes(include=['object']).columns # 对每个字符串列执行替换 df[str_columns] = df[str_columns].apply(lambda col: col.str.replace(pattern, '', regex=True))
只处理特定列(比如Json列)
如果你只需要清理Json列,直接针对它操作就行:
df['Json'] = df['Json'].str.replace(pattern, '', regex=True)
3. 清理残留的无效格式
替换后可能会留下一些多余的逗号或者空的键值对(比如number-of-shots-fired:,),可以再加一步清理:
# 替换掉像`,}`这样的无效格式 cleanup_pattern = r',\s*\}' df[str_columns] = df[str_columns].apply(lambda col: col.str.replace(cleanup_pattern, '}', regex=True)) # 如果想删掉空的键值对(比如`number-of-shots-fired:`),可以用这个 empty_key_pattern = r'\w+:\s*(?=,|\})' df[str_columns] = df[str_columns].apply(lambda col: col.str.replace(empty_key_pattern, '', regex=True))
举个实际例子
用你提供的示例数据来测试:
# 构造示例DataFrame sample_data = { "_id": ["ObjectId('5ad13551df70d8025b7811db')"], "Article url": ["http://wreg.com/2015/08/02/mpd-officer-killed-in-the-line-of-duty-identified/"], "Full text": ["Blah blah"], "Json": ["{circumstances:{number-of-shots-fired:{endIndex:-1,startIndex:-1,value:},type-of-gun:{endIndex:-1,startIndex:-1,value:}}}]"] } df = pd.DataFrame(sample_data) # 执行替换和清理 df['Json'] = df['Json'].str.replace(pattern, '', regex=True) df['Json'] = df['Json'].str.replace(cleanup_pattern, '}', regex=True) df['Json'] = df['Json'].str.replace(empty_key_pattern, '', regex=True)
处理后Json列的结果会变成:{circumstances:{}}}],完美去掉了所有带目标字段的结构!
内容的提问来源于stack exchange,提问作者Deik
相关产品推荐
相关产品推荐

