You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Regex从字符串中精准提取指定键值对(如firstName)

问题

抓取到一段包含多组JSON键值对的网页字符串,想要提取其中firstName、lastName等字段生成列表。尝试用正则表达式re.findall('"firstName":"(.*)\S$",st)时,匹配结果附带了多余内容,如何让正则在姓名的结束引号位置停止匹配?

示例字符串代码:

st = '{"accountId":405266,"firstName":"Quaran","lastName":"McPherson","accountIdentifier":"StudentAthlete","profilePicUrl":"https://pbs.twimg.com/profile_images/1331475329014181888/4z19KrCf.jpg","networkProfileCode":"quaran-mcpherson","hasDeals":true,"activityMin":11,"sports":["Men\'s Basketball","Basketball"],"currentTeams":["Nebraska Cornhuskers"],"previousTeams":[],"facebookReach":null,"twitterReach":619,"instagramReach":0,"linkedInReach":null},{"accountId":375964,"firstName":"Micole","lastName":"Cayton","accountIdentifier":"StudentAthlete","profilePicUrl":"https://opendorsepr.blob.core.windows.net/media/375964/20220622223838_46dbe3fd-a683-436b-84d4-90c84a5af35f.jpg","networkProfileCode":"micole-cayton","hasDeals":true,"activityMin":16,"sports":["Basketball","Women\'s Basketball"],"currentTeams":["Minnesota Golden Gophers"],"previousTeams":["Cal Berkeley Golden Bears"],"facebookReach":0,"twitterReach":1273,"instagramReach":5700,"linkedInReach":null}'

解决方法

一、优化正则表达式

之前的正则用(.*)贪婪匹配,会尽可能抓取到最后一个符合条件的位置,导致多余内容。可以通过两种方式修正:

方式1:非贪婪匹配

将.*改为.*?,让正则匹配到第一个结束引号"就停止:

import re

# 提取firstName
first_names = re.findall(r'"firstName":"(.*?)"', st)
# 提取lastName
last_names = re.findall(r'"lastName":"(.*?)"', st)

print(first_names)  # 输出: ['Quaran', 'Micole']
print(last_names)   # 输出: ['McPherson', 'Cayton']

方式2:字符集限定匹配范围

因为姓名不会包含&,而结束引号的开头是&,直接匹配非&的所有字符,精准停止在结束引号前:

first_names = re.findall(r'"firstName":"([^&]+)"', st)
last_names = re.findall(r'"lastName":"([^&]+)"', st)

这种方式匹配效率更高,也能避免意外匹配。


二、转为JSON解析(更推荐)

正则处理JSON格式数据容易踩坑(比如字段顺序变化、特殊字符转义等),更稳妥的方式是先将字符串转为合法JSON,再用Python内置json模块解析:

import json

# 1. 替换转义引号为标准双引号,包装成JSON数组
processed_str = '[' + st.replace('"', '"') + ']'
# 2. 解析JSON数据
data_list = json.loads(processed_str)

# 3. 提取目标字段
first_names = [item['firstName'] for item in data_list]
last_names = [item['lastName'] for item in data_list]

print(first_names)  # 输出: ['Quaran', 'Micole']
print(last_names)   # 输出: ['McPherson', 'Cayton']

这种方法能应对各种复杂场景,比如姓名含特殊字符、JSON结构调整等,通用性更强。


内容的提问来源于stack exchange,提问作者Rithwik Sivadasan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 11:37:01