使用BeautifulSoup提取JSON格式script标签数据时遇解析错误
解决BeautifulSoup提取application/ld+json标签时的JSON解析错误
你遇到的Expecting value: line 1 column 1 (char 0)错误,核心问题有两个:
- 直接将找到的script标签对象转成字符串,得到的是包含
<script>标签本身的完整HTML代码,不是标签内的纯JSON文本,导致json.loads无法解析。 - 多余的
replace("'",'"')操作完全没必要,原JSON本身已经使用双引号,这个操作反而可能破坏合法的字符串内容。
修正后的代码如下:
import json import re from bs4 import BeautifulSoup # 假设已完成HTML解析得到soup对象 script_tag = soup.find("script", type="application/ld+json") # 先检查标签是否存在,避免None引发后续错误 if script_tag: # 获取标签内的纯JSON文本,strip()去除首尾空白和换行 json_text = script_tag.string.strip() # 直接解析合法的JSON内容 script_dict = json.loads(json_text) data = dict() data["author"] = script_dict["author"] data["embed_url"] = script_dict["embedUrl"] data["duration"] = ":".join(re.findall(r"\d\d", script_dict["duration"])) data["upload_date"] = re.findall(r"\d{4}-\d{2}-\d{2}", script_dict["uploadDate"])[0] data["accurate_views"] = int(script_dict["interactionStatistic"][0]["userInteractionCount"].replace(",", "")) print(data) else: print("未找到目标script标签")
关键修改说明:
- 使用
script_tag.string.strip()获取标签内的纯文本内容,而非整个标签的HTML代码,这是解析JSON的核心前提。 - 移除了不必要的单引号替换操作,原JSON格式符合规范,直接解析即可。
- 增加了标签存在性判断,避免找不到标签时触发空值解析错误。
内容的提问来源于stack exchange,提问作者MaxFrost
相关产品推荐
相关产品推荐

