如何用Regex从字符串中精准提取指定键值对(如firstName)
问题
抓取到一段包含多组JSON键值对的网页字符串,想要提取其中firstName、lastName等字段生成列表。尝试用正则表达式re.findall('"firstName":"(.*)\S$",st)时,匹配结果附带了多余内容,如何让正则在姓名的结束引号位置停止匹配?
示例字符串代码:
st = '{"accountId":405266,"firstName":"Quaran","lastName":"McPherson","accountIdentifier":"StudentAthlete","profilePicUrl":"https://pbs.twimg.com/profile_images/1331475329014181888/4z19KrCf.jpg","networkProfileCode":"quaran-mcpherson","hasDeals":true,"activityMin":11,"sports":["Men\'s Basketball","Basketball"],"currentTeams":["Nebraska Cornhuskers"],"previousTeams":[],"facebookReach":null,"twitterReach":619,"instagramReach":0,"linkedInReach":null},{"accountId":375964,"firstName":"Micole","lastName":"Cayton","accountIdentifier":"StudentAthlete","profilePicUrl":"https://opendorsepr.blob.core.windows.net/media/375964/20220622223838_46dbe3fd-a683-436b-84d4-90c84a5af35f.jpg","networkProfileCode":"micole-cayton","hasDeals":true,"activityMin":16,"sports":["Basketball","Women\'s Basketball"],"currentTeams":["Minnesota Golden Gophers"],"previousTeams":["Cal Berkeley Golden Bears"],"facebookReach":0,"twitterReach":1273,"instagramReach":5700,"linkedInReach":null}'
解决方法
一、优化正则表达式
之前的正则用(.*)贪婪匹配,会尽可能抓取到最后一个符合条件的位置,导致多余内容。可以通过两种方式修正:
方式1:非贪婪匹配
将.*改为.*?,让正则匹配到第一个结束引号"就停止:
import re # 提取firstName first_names = re.findall(r'"firstName":"(.*?)"', st) # 提取lastName last_names = re.findall(r'"lastName":"(.*?)"', st) print(first_names) # 输出: ['Quaran', 'Micole'] print(last_names) # 输出: ['McPherson', 'Cayton']
方式2:字符集限定匹配范围
因为姓名不会包含&,而结束引号的开头是&,直接匹配非&的所有字符,精准停止在结束引号前:
first_names = re.findall(r'"firstName":"([^&]+)"', st) last_names = re.findall(r'"lastName":"([^&]+)"', st)
这种方式匹配效率更高,也能避免意外匹配。
二、转为JSON解析(更推荐)
正则处理JSON格式数据容易踩坑(比如字段顺序变化、特殊字符转义等),更稳妥的方式是先将字符串转为合法JSON,再用Python内置json模块解析:
import json # 1. 替换转义引号为标准双引号,包装成JSON数组 processed_str = '[' + st.replace('"', '"') + ']' # 2. 解析JSON数据 data_list = json.loads(processed_str) # 3. 提取目标字段 first_names = [item['firstName'] for item in data_list] last_names = [item['lastName'] for item in data_list] print(first_names) # 输出: ['Quaran', 'Micole'] print(last_names) # 输出: ['McPherson', 'Cayton']
这种方法能应对各种复杂场景,比如姓名含特殊字符、JSON结构调整等,通用性更强。
内容的提问来源于stack exchange,提问作者Rithwik Sivadasan
相关产品推荐
相关产品推荐

