You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式删除Pandas DataFrame中的指定索引结构子串

在Pandas DataFrame中批量删除包含endIndex/startIndex/value的结构化子串

嘿,我来帮你搞定这个问题!你需要在Pandas的DataFrame里批量删掉那些带endIndex、startIndex、value的结构化子串,不管是单独的{endIndex:...}块还是嵌套在其他结构里的,对吧?我整理了一个实用的方案,一步步来:

1. 构造匹配目标结构的正则表达式

首先得写一个能精准匹配这类结构的正则。考虑到可能有嵌套(比如你提到的date-and-time:{city:{endIndex:...}}),我用了支持递归的正则来处理嵌套的JSON对象,确保能把包含三个目标字段的整个结构都匹配到:

import re

# 匹配同时包含endIndex、startIndex、value的JSON对象(支持嵌套)
pattern = r'\{(?:[^{}]|(?R))*\bendIndex\b(?:[^{}]|(?R))*\bstartIndex\b(?:[^{}]|(?R))*\bvalue\b(?:[^{}]|(?R))*\}'

如果你的场景里这些字段都是按endIndex→startIndex→value的顺序出现,且嵌套不复杂,也可以用更简单的非递归版本,运行更快:

# 匹配顺序固定的非嵌套结构
simple_pattern = r'\{endIndex:\s*-?\d+,\s*startIndex:\s*-?\d+,\s*value:[^}]*\}'

2. 在Pandas中批量替换

接下来就是把这个正则应用到DataFrame的字符串列上。你可以选择处理所有字符串列,或者只针对特定列(比如你的Json列):

处理所有字符串列

import pandas as pd

# 假设你的DataFrame叫df
# 获取所有字符串类型的列
str_columns = df.select_dtypes(include=['object']).columns

# 对每个字符串列执行替换
df[str_columns] = df[str_columns].apply(lambda col: col.str.replace(pattern, '', regex=True))

只处理特定列(比如Json列)

如果你只需要清理Json列,直接针对它操作就行:

df['Json'] = df['Json'].str.replace(pattern, '', regex=True)

3. 清理残留的无效格式

替换后可能会留下一些多余的逗号或者空的键值对(比如number-of-shots-fired:,),可以再加一步清理:

# 替换掉像`,}`这样的无效格式
cleanup_pattern = r',\s*\}'
df[str_columns] = df[str_columns].apply(lambda col: col.str.replace(cleanup_pattern, '}', regex=True))

# 如果想删掉空的键值对(比如`number-of-shots-fired:`),可以用这个
empty_key_pattern = r'\w+:\s*(?=,|\})'
df[str_columns] = df[str_columns].apply(lambda col: col.str.replace(empty_key_pattern, '', regex=True))

举个实际例子

用你提供的示例数据来测试:

# 构造示例DataFrame
sample_data = {
    "_id": ["ObjectId('5ad13551df70d8025b7811db')"],
    "Article url": ["http://wreg.com/2015/08/02/mpd-officer-killed-in-the-line-of-duty-identified/"],
    "Full text": ["Blah blah"],
    "Json": ["{circumstances:{number-of-shots-fired:{endIndex:-1,startIndex:-1,value:},type-of-gun:{endIndex:-1,startIndex:-1,value:}}}]"]
}
df = pd.DataFrame(sample_data)

# 执行替换和清理
df['Json'] = df['Json'].str.replace(pattern, '', regex=True)
df['Json'] = df['Json'].str.replace(cleanup_pattern, '}', regex=True)
df['Json'] = df['Json'].str.replace(empty_key_pattern, '', regex=True)

处理后Json列的结果会变成:{circumstances:{}}}],完美去掉了所有带目标字段的结构!

内容的提问来源于stack exchange,提问作者Deik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:46:32