使用BeautifulSoup提取script标签JSON数据失败,请求排查问题
问题分析与修复方案
你的代码无法提取目标字段主要有以下几个原因:
- HTML内容中的标签、引号都是转义后的实体(如
<、"),未还原会导致BeautifulSoup解析异常 - Script标签内的JSON被
<!-- -->注释包裹,直接读取会包含注释标记,无法解析为有效JSON - 提供的原始JSON存在语法错误(比如
"queryState":numberFormatter缺少引号和结构、多余的括号逗号) - 提取脚本内容时重复使用BeautifulSoup解析属于冗余操作,且未处理注释内容
修正后的代码
from bs4 import BeautifulSoup import json import html def extract_property_info(html): # 还原HTML转义实体,将<、"转为正常标签和引号 unescaped_html = html.unescape(html) soup = BeautifulSoup(unescaped_html, 'html.parser') script_tag = soup.find('script', {'type': 'application/json', 'ss-data-keyy': 'mobileSearchPageStore'}) if script_tag: script_content = script_tag.string.strip() try: # 移除HTML注释标记,获取纯净JSON字符串 json_str = script_content.replace('<!--', '').replace('-->', '') # 修复原始JSON的语法错误:补全缺失的键结构,移除多余的结尾符号 json_str = json_str.replace('"queryState":numberFormatter', '"queryState":{},"numberFormatter"') json_str = json_str.rstrip('},]') + '}' data = json.loads(json_str) # 提取目标字段,用get()避免字段不存在报错 detail_url = data.get('detailUrl') address = data.get('address') beds = data.get('beds') baths = data.get('baths') price = data.get('price') return detail_url, address, beds, baths, price except (AttributeError, json.JSONDecodeError) as e: # 打印错误信息便于排查问题 print(f"解析错误: {str(e)}") pass return None # HTML内容 html_content = ''' <html> <body> <script type="application/json" ss-data-keyy="mobileSearchPageStore"> <!--{"queryState":numberFormatter":"0,0","inputFormatter":"0.[0]a"}},"sortOrder":1,"type":"Range","defaultValue":{"min":null,"max":null},"suggestedEnums":500","value":500},{"url":"https://photos.example.com/fp/e0ba89ff3ebcb9d689dcf6d0ffd87868-p_e.jpg"}],"detailUrl":"https://www.example.com/homedetails/111-S-1316th-St-Arizona-AZ-3324332/33327367_ppid/","statusType":"FOR_SALE","statusText":"home for buy","countryCurrency":"$","price":"$269,000","unformattedPrice":269000,"address":"111 S 1316th St Arizona AZ 33243322","addressStreet":"111 S 1316th St","addressCity":"Arizona","addressState":"OH","addressZipcode":"33243322","isUndisclosedAddress":false,"beds":4,"baths":2.0,"area":2467,"latLong":{"latitude":12.12345,"longitude":12.12345},"isexampleOwned":false,"variableData":{"type":"DAYS_ON","text":"2 days on example"},"badgeInfo":null,}],}--> </script> </body> </html> ''' # 调用函数并打印结果 result = extract_property_info(html_content) print("Result: ", result) if result is not None: detail_url, address, beds, baths, price = result print("Detail URL:", detail_url) print("Address:", address) print("Beds:", beds) print("Baths:", baths) print("Price:", price) else: print("Unable to extract property information.")
关键修改说明
- 还原HTML转义实体:用
html.unescape()将转义字符还原为正常格式,确保BeautifulSoup能正确识别标签和内容。 - 移除注释标记:清理脚本内容中的HTML注释,得到可解析的JSON字符串。
- 修复JSON语法:针对你提供的残缺JSON做针对性修复,实际场景中需根据真实返回的JSON结构调整。
- 增强容错性:用
get()方法提取字段,避免因字段缺失引发KeyError;添加错误打印,方便排查解析失败原因。
内容的提问来源于stack exchange,提问作者Lacer
相关产品推荐
相关产品推荐

