如何用Python正则提取多行文本双大括号间内容并转为Pandas DataFrame?
解决方法
你的代码存在几个关键问题:
re.match仅匹配字符串开头,而你的文本开头是换行和[,根本匹配不到目标内容- 即便匹配成功,贪婪模式的
.*会把所有内容合并成一个字符串,无法直接转换为DataFrame - 原文本里的字典键(比如
specialty)没有引号,不属于合法JSON格式,直接解析会报错
下面是正确的实现方案:
步骤1:修复文本格式,转为合法JSON
先给无引号的键添加双引号,这是后续解析的必要前提:
import re import json import pandas as pd s = """ [ { specialty: "Anatomic/Clinical Pathology", one: " 12,643 ", two: " 8,711 ", three: " 385 ", four: " 520 ", five: " 3,027 ", }, { specialty: "Nephrology", one: " 11,407 ", two: " 9,964 ", three: " 140 ", four: " 316 ", five: " 987 ", }, { specialty: "Vascular Surgery", one: " 3,943 ", two: " 3,586 ", three: " 48 ", four: " 13 ", five: " 296 ", }, ] """ # 给无引号的键加上双引号 fixed_s = re.sub(r'(\w+):', r'"\1":', s)
步骤2:解析JSON并转换为DataFrame
用json.loads解析修复后的文本,再直接转换为DataFrame:
# 解析JSON数据 data = json.loads(fixed_s) # 转换为Pandas DataFrame df = pd.DataFrame(data) # 可选操作:清理数值列的空格和逗号,转为整数类型 df[['one', 'two', 'three', 'four', 'five']] = df[['one', 'two', 'three', 'four', 'five']].apply( lambda x: x.str.strip().str.replace(',', '').astype(int) ) print(df)
最终输出
specialty one two three four five 0 Anatomic/Clinical Pathology 12643 8711 385 520 3027 1 Nephrology 11407 9964 140 316 987 2 Vascular Surgery 3943 3586 48 13 296
内容的提问来源于stack exchange,提问作者dallascow
相关产品推荐
相关产品推荐

