You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取script标签JSON数据失败,请求排查问题

问题分析与修复方案

你的代码无法提取目标字段主要有以下几个原因:

  • HTML内容中的标签、引号都是转义后的实体(如<、"),未还原会导致BeautifulSoup解析异常
  • Script标签内的JSON被<!-- -->注释包裹,直接读取会包含注释标记,无法解析为有效JSON
  • 提供的原始JSON存在语法错误(比如"queryState":numberFormatter缺少引号和结构、多余的括号逗号)
  • 提取脚本内容时重复使用BeautifulSoup解析属于冗余操作,且未处理注释内容

修正后的代码

from bs4 import BeautifulSoup
import json
import html

def extract_property_info(html):
    # 还原HTML转义实体,将&lt;、&quot;转为正常标签和引号
    unescaped_html = html.unescape(html)
    soup = BeautifulSoup(unescaped_html, 'html.parser')
    script_tag = soup.find('script', {'type': 'application/json', 'ss-data-keyy': 'mobileSearchPageStore'})

    if script_tag:
        script_content = script_tag.string.strip()
        try:
            # 移除HTML注释标记,获取纯净JSON字符串
            json_str = script_content.replace('<!--', '').replace('-->', '')
            # 修复原始JSON的语法错误:补全缺失的键结构,移除多余的结尾符号
            json_str = json_str.replace('"queryState":numberFormatter', '"queryState":{},"numberFormatter"')
            json_str = json_str.rstrip('},]') + '}'
            data = json.loads(json_str)
            
            # 提取目标字段,用get()避免字段不存在报错
            detail_url = data.get('detailUrl')
            address = data.get('address')
            beds = data.get('beds')
            baths = data.get('baths')
            price = data.get('price')
            
            return detail_url, address, beds, baths, price
        except (AttributeError, json.JSONDecodeError) as e:
            # 打印错误信息便于排查问题
            print(f"解析错误: {str(e)}")
            pass

    return None

# HTML内容 
html_content = '''
&lt;html&gt;
&lt;body&gt;
    &lt;script type=&quot;application/json&quot; ss-data-keyy=&quot;mobileSearchPageStore&quot;&gt;
    &lt;!--{&quot;queryState&quot;:numberFormatter&quot;:&quot;0,0&quot;,&quot;inputFormatter&quot;:&quot;0.[0]a&quot;}},&quot;sortOrder&quot;:1,&quot;type&quot;:&quot;Range&quot;,&quot;defaultValue&quot;:{&quot;min&quot;:null,&quot;max&quot;:null},&quot;suggestedEnums&quot;:500&quot;,&quot;value&quot;:500},{&quot;url&quot;:&quot;https://photos.example.com/fp/e0ba89ff3ebcb9d689dcf6d0ffd87868-p_e.jpg&quot;}],&quot;detailUrl&quot;:&quot;https://www.example.com/homedetails/111-S-1316th-St-Arizona-AZ-3324332/33327367_ppid/&quot;,&quot;statusType&quot;:&quot;FOR_SALE&quot;,&quot;statusText&quot;:&quot;home for buy&quot;,&quot;countryCurrency&quot;:&quot;$&quot;,&quot;price&quot;:&quot;$269,000&quot;,&quot;unformattedPrice&quot;:269000,&quot;address&quot;:&quot;111 S 1316th St Arizona AZ 33243322&quot;,&quot;addressStreet&quot;:&quot;111 S 1316th St&quot;,&quot;addressCity&quot;:&quot;Arizona&quot;,&quot;addressState&quot;:&quot;OH&quot;,&quot;addressZipcode&quot;:&quot;33243322&quot;,&quot;isUndisclosedAddress&quot;:false,&quot;beds&quot;:4,&quot;baths&quot;:2.0,&quot;area&quot;:2467,&quot;latLong&quot;:{&quot;latitude&quot;:12.12345,&quot;longitude&quot;:12.12345},&quot;isexampleOwned&quot;:false,&quot;variableData&quot;:{&quot;type&quot;:&quot;DAYS_ON&quot;,&quot;text&quot;:&quot;2 days on example&quot;},&quot;badgeInfo&quot;:null,}],}--&gt;
    &lt;/script&gt;
&lt;/body&gt;
&lt;/html&gt;
'''

# 调用函数并打印结果
result = extract_property_info(html_content)
print("Result: ", result)
if result is not None:
    detail_url, address, beds, baths, price = result
    print("Detail URL:", detail_url)
    print("Address:", address)
    print("Beds:", beds)
    print("Baths:", baths)
    print("Price:", price)
else:
    print("Unable to extract property information.")

关键修改说明

  1. 还原HTML转义实体:用html.unescape()将转义字符还原为正常格式,确保BeautifulSoup能正确识别标签和内容。
  2. 移除注释标记:清理脚本内容中的HTML注释,得到可解析的JSON字符串。
  3. 修复JSON语法:针对你提供的残缺JSON做针对性修复,实际场景中需根据真实返回的JSON结构调整。
  4. 增强容错性:用get()方法提取字段,避免因字段缺失引发KeyError;添加错误打印,方便排查解析失败原因。

内容的提问来源于stack exchange,提问作者Lacer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 06:44:58