You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用while循环遍历lxml获取的全部script元素(而非手动指定索引)

使用while循环遍历lxml获取的所有script元素

保留你现有代码结构的基础上,我们可以通过先获取完整的script元素列表,再用while循环按索引遍历的方式实现自动输出所有元素,修改后的代码如下:

import requests
import lxml.html as lh
from lxml.etree import tostring

req = requests.get('https://www.dailymail.co.uk/debate/article-11113609/DAN-WOOTTON-intolerant-left-belittles-death-threats-against-JK-Rowling-peril.html#comments-11113609')
df = lh.fromstring(req.text)

# 先获取所有script元素的完整列表
scripts = df.xpath('//script')
# 初始化索引变量,从0开始遍历
index = 0
# while循环:只要索引小于列表长度,就继续执行
while index < len(scripts):
    elem = scripts[index]
    print(tostring(elem))
    # 索引自增,避免死循环
    index += 1

关键说明

  • 先通过df.xpath('//script')拿到所有script元素的列表,明确总元素数量,防止索引越界
  • 初始化index为0(对应第一个元素的索引),每次输出当前元素后将索引加1
  • while循环的终止条件设为index < len(scripts),当索引等于列表长度时,说明已遍历完所有元素,循环自动停止

内容的提问来源于stack exchange,提问作者Babiqowski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 12:20:40