Python requests-html无法定位script标签内容 如何提取EAN编码
EAN编码提取失败修复方案
- 原有代码通过固定索引
[3]定位script标签的方式稳定性极差,页面加载过程中script标签的数量、顺序会随页面渲染逻辑变动,极易定位到非目标标签导致提取失败。 - 正确提取逻辑优先通过内容特征定位目标script块,而非依赖固定索引,参考代码如下:
import re # 遍历所有script标签匹配内容特征 for script in r.html.find('script'): script_content = script.text # 匹配EAN字段特征或者目标编码,可根据实际页面字段调整关键词 if 'ean' in script_content.lower() or 'gtin13' in script_content or '8806090571589' in script_content: # 正则提取13位EAN编码 ean_res = re.search(r'"(ean|gtin13)"\s*:\s*"(\d{13})"', script_content) if ean_res: extracted_ean = ean_res.group(2) print(title, price, extracted_ean) break
- 若页面为前端动态渲染,需要先执行js渲染再提取内容,调用
r.html.render()即可加载动态生成的script内容。
内容的提问来源于stack exchange,提问作者lanavargas2002
相关产品推荐
相关产品推荐

