You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python requests-html无法定位script标签内容 如何提取EAN编码

EAN编码提取失败修复方案
  • 原有代码通过固定索引[3]定位script标签的方式稳定性极差,页面加载过程中script标签的数量、顺序会随页面渲染逻辑变动,极易定位到非目标标签导致提取失败。
  • 正确提取逻辑优先通过内容特征定位目标script块,而非依赖固定索引,参考代码如下:
import re

# 遍历所有script标签匹配内容特征
for script in r.html.find('script'):
    script_content = script.text
    # 匹配EAN字段特征或者目标编码,可根据实际页面字段调整关键词
    if 'ean' in script_content.lower() or 'gtin13' in script_content or '8806090571589' in script_content:
        # 正则提取13位EAN编码
        ean_res = re.search(r'"(ean|gtin13)"\s*:\s*"(\d{13})"', script_content)
        if ean_res:
            extracted_ean = ean_res.group(2)
            print(title, price, extracted_ean)
        break
  • 若页面为前端动态渲染,需要先执行js渲染再提取内容,调用r.html.render()即可加载动态生成的script内容。

内容的提问来源于stack exchange,提问作者lanavargas2002

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 04:27:03