You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则提取多行文本双大括号间内容并转为Pandas DataFrame?

解决方法

你的代码存在几个关键问题:

  • re.match 仅匹配字符串开头,而你的文本开头是换行和[,根本匹配不到目标内容
  • 即便匹配成功,贪婪模式的.*会把所有内容合并成一个字符串,无法直接转换为DataFrame
  • 原文本里的字典键(比如specialty)没有引号,不属于合法JSON格式,直接解析会报错

下面是正确的实现方案:

步骤1:修复文本格式,转为合法JSON

先给无引号的键添加双引号,这是后续解析的必要前提:

import re
import json
import pandas as pd

s = """
[
          {
            specialty: "Anatomic/Clinical Pathology",
            one: " 12,643 ",
            two: " 8,711 ",
            three: " 385 ",
            four: " 520 ",
            five: " 3,027 ",
          },
          {
            specialty: "Nephrology",
            one: " 11,407 ",
            two: " 9,964 ",
            three: " 140 ",
            four: " 316 ",
            five: " 987 ",
          },
          {
            specialty: "Vascular Surgery",
            one: " 3,943 ",
            two: " 3,586 ",
            three: " 48 ",
            four: " 13 ",
            five: " 296 ",
          },
        ]
"""

# 给无引号的键加上双引号
fixed_s = re.sub(r'(\w+):', r'"\1":', s)

步骤2:解析JSON并转换为DataFrame

用json.loads解析修复后的文本,再直接转换为DataFrame:

# 解析JSON数据
data = json.loads(fixed_s)
# 转换为Pandas DataFrame
df = pd.DataFrame(data)

# 可选操作:清理数值列的空格和逗号,转为整数类型
df[['one', 'two', 'three', 'four', 'five']] = df[['one', 'two', 'three', 'four', 'five']].apply(
    lambda x: x.str.strip().str.replace(',', '').astype(int)
)

print(df)

最终输出

specialty    one   two  three  four  five
0  Anatomic/Clinical Pathology  12643  8711    385   520  3027
1                   Nephrology  11407  9964    140   316   987
2             Vascular Surgery   3943  3586     48    13   296

内容的提问来源于stack exchange,提问作者dallascow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 06:10:41