如何用正则移除JSON内嵌套双引号解决json.loads解析报错问题
修复方案
1 移除无效的BeautifulSoup处理
你请求的接口返回的是纯JSON格式数据,并非HTML结构,用BeautifulSoup解析属于冗余操作,还可能引入额外转义问题,直接使用requests返回的原始响应文本即可。
2 精准替换rComments字段内部的未转义双引号
核心逻辑是用正则精准匹配rComments字段的完整键值范围,仅替换字段内容内部的双引号为单引号,不破坏JSON外层语法结构,代码如下:
import requests import re import json # 配置请求参数 tid = 124880 url = f'https://www.ratemyprofessors.com/paginate/professors/ratings?tid={tid}&filter=&courseCode=&page=1' # 加浏览器UA避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } # 发起请求拿原始文本 resp = requests.get(url, headers=headers) raw_json = resp.text # 替换rComments内部嵌套的双引号 def replace_nested_quotes(match): # 匹配到的完整片段格式为 "rComments":"[内容]",仅替换[内容]中的双引号 content = match.group(1) fixed_content = content.replace('"', "'") return f'"rComments":"{fixed_content}"' # 正则规则:非贪婪匹配rComments的内容,正向预查确保匹配到下一个JSON键之前停止 fixed_json = re.sub(r'"rComments":"(.*?)"(?=,")', replace_nested_quotes, raw_json) # 正常解析JSON result = json.loads(fixed_json) print(result)
其他说明
如果后续遇到其他字段也存在未转义双引号的问题,只需要把正则中的rComments替换为对应字段名即可。
内容的提问来源于stack exchange,提问作者Aaron Murz
相关产品推荐
相关产品推荐

