如何将lxml提取的script中JSON字符串转为Python字典?
处理script标签内JSON数据并转为Python字典
基础实现方法
你可以直接用Python内置的json模块解析data[1].text,只要这段文本是合法的JSON格式就行。代码示例:
import json from lxml import etree # 假设已通过lxml获取到tree对象 data = tree.xpath('//script/text()') try: # 将script文本转为Python字典 json_data = json.loads(data[1].text) # 提取需要的内部数据,比如指定字段 needed_data = json_data.get("your_target_key") except json.JSONDecodeError as err: print(f"JSON解析出错: {err}")
更优实现方案
1. 精准定位目标script标签
靠索引data[1]定位很容易因为页面结构变动失效,建议通过script标签的属性(比如id、type)或者内容里的关键词来锁定:
# 定位包含特定关键词的script标签 target_script = tree.xpath('//script[contains(text(), "unique_json_key")]/text()')[0] json_data = json.loads(target_script)
2. 兼容非标准JSON
如果网站返回的是非标准JSON(比如用单引号、末尾多逗号),可以用json5库解析,兼容性更强:
import json5 # 先安装:pip install json5 json_data = json5.loads(target_script)
3. 简化请求+解析流程
如果是自己发请求拿HTML,可以把requests和lxml结合起来,减少冗余步骤:
import requests from lxml import etree import json url = "你要爬的网址" resp = requests.get(url) tree = etree.HTML(resp.text) target_script = tree.xpath('//script[contains(text(), "target_keyword")]/text()')[0] json_data = json.loads(target_script)
4. 增加鲁棒性避免索引越界
获取script列表后先判断长度,防止因标签数量不足报错:
scripts = tree.xpath('//script/text()') if len(scripts) >= 2: json_data = json.loads(scripts[1]) else: print("找不到目标script标签")
内容的提问来源于stack exchange,提问作者The Dan
相关产品推荐
相关产品推荐

