You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取JSON格式script标签数据时遇解析错误

解决BeautifulSoup提取application/ld+json标签时的JSON解析错误

你遇到的Expecting value: line 1 column 1 (char 0)错误,核心问题有两个:

  • 直接将找到的script标签对象转成字符串,得到的是包含<script>标签本身的完整HTML代码,不是标签内的纯JSON文本,导致json.loads无法解析。
  • 多余的replace("'",'"')操作完全没必要,原JSON本身已经使用双引号,这个操作反而可能破坏合法的字符串内容。

修正后的代码如下:

import json
import re
from bs4 import BeautifulSoup

# 假设已完成HTML解析得到soup对象
script_tag = soup.find("script", type="application/ld+json")

# 先检查标签是否存在,避免None引发后续错误
if script_tag:
    # 获取标签内的纯JSON文本,strip()去除首尾空白和换行
    json_text = script_tag.string.strip()
    # 直接解析合法的JSON内容
    script_dict = json.loads(json_text)
    
    data = dict()
    data["author"] = script_dict["author"] 
    data["embed_url"] = script_dict["embedUrl"]
    data["duration"] = ":".join(re.findall(r"\d\d", script_dict["duration"]))
    data["upload_date"] = re.findall(r"\d{4}-\d{2}-\d{2}", script_dict["uploadDate"])[0]
    data["accurate_views"] = int(script_dict["interactionStatistic"][0]["userInteractionCount"].replace(",", ""))
    print(data)
else:
    print("未找到目标script标签")

关键修改说明:

  • 使用script_tag.string.strip()获取标签内的纯文本内容,而非整个标签的HTML代码,这是解析JSON的核心前提。
  • 移除了不必要的单引号替换操作,原JSON格式符合规范,直接解析即可。
  • 增加了标签存在性判断,避免找不到标签时触发空值解析错误。

内容的提问来源于stack exchange,提问作者MaxFrost

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 14:05:39