如何用while循环遍历lxml获取的全部script元素(而非手动指定索引)
使用while循环遍历lxml获取的所有script元素
保留你现有代码结构的基础上,我们可以通过先获取完整的script元素列表,再用while循环按索引遍历的方式实现自动输出所有元素,修改后的代码如下:
import requests import lxml.html as lh from lxml.etree import tostring req = requests.get('https://www.dailymail.co.uk/debate/article-11113609/DAN-WOOTTON-intolerant-left-belittles-death-threats-against-JK-Rowling-peril.html#comments-11113609') df = lh.fromstring(req.text) # 先获取所有script元素的完整列表 scripts = df.xpath('//script') # 初始化索引变量,从0开始遍历 index = 0 # while循环:只要索引小于列表长度,就继续执行 while index < len(scripts): elem = scripts[index] print(tostring(elem)) # 索引自增,避免死循环 index += 1
关键说明
- 先通过
df.xpath('//script')拿到所有script元素的列表,明确总元素数量,防止索引越界 - 初始化
index为0(对应第一个元素的索引),每次输出当前元素后将索引加1 - while循环的终止条件设为
index < len(scripts),当索引等于列表长度时,说明已遍历完所有元素,循环自动停止
内容的提问来源于stack exchange,提问作者Babiqowski
相关产品推荐
相关产品推荐

