You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将lxml提取的script中JSON字符串转为Python字典?

处理script标签内JSON数据并转为Python字典

基础实现方法

你可以直接用Python内置的json模块解析data[1].text,只要这段文本是合法的JSON格式就行。代码示例:

import json
from lxml import etree

# 假设已通过lxml获取到tree对象
data = tree.xpath('//script/text()')
try:
    # 将script文本转为Python字典
    json_data = json.loads(data[1].text)
    # 提取需要的内部数据,比如指定字段
    needed_data = json_data.get("your_target_key")
except json.JSONDecodeError as err:
    print(f"JSON解析出错: {err}")

更优实现方案

1. 精准定位目标script标签

靠索引data[1]定位很容易因为页面结构变动失效,建议通过script标签的属性(比如id、type)或者内容里的关键词来锁定:

# 定位包含特定关键词的script标签
target_script = tree.xpath('//script[contains(text(), "unique_json_key")]/text()')[0]
json_data = json.loads(target_script)

2. 兼容非标准JSON

如果网站返回的是非标准JSON(比如用单引号、末尾多逗号),可以用json5库解析,兼容性更强:

import json5

# 先安装:pip install json5
json_data = json5.loads(target_script)

3. 简化请求+解析流程

如果是自己发请求拿HTML,可以把requests和lxml结合起来,减少冗余步骤:

import requests
from lxml import etree
import json

url = "你要爬的网址"
resp = requests.get(url)
tree = etree.HTML(resp.text)
target_script = tree.xpath('//script[contains(text(), "target_keyword")]/text()')[0]
json_data = json.loads(target_script)

4. 增加鲁棒性避免索引越界

获取script列表后先判断长度,防止因标签数量不足报错:

scripts = tree.xpath('//script/text()')
if len(scripts) >= 2:
    json_data = json.loads(scripts[1])
else:
    print("找不到目标script标签")

内容的提问来源于stack exchange,提问作者The Dan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 18:15:47