You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python for循环仅返回首个索引项 BeautifulSoup网页爬取问题求解

问题原因

  • HTML规范中id属性为全局唯一值,你使用soup.find_all("div",{"id":"content"})获取到的列表实际仅包含1个元素,也就是页面唯一的content容器
  • 后续遍历这个只有1个元素的列表时,每次都只取了容器下所有h2标签的第一个(item.find_all("h2")[0].text),自然只会输出首个标题

修复后代码

import requests
from bs4 import BeautifulSoup

r = requests.get("http://cpaleaks.com", headers={'User-agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:61.0) Gecko/20100101 Firefox/61.0'})
c = r.content
soup = BeautifulSoup(c,"html.parser")

# 直接定位到content容器后提取其下所有h2标签
h2_list = soup.find("div",{"id":"content"}).find_all("h2")

# 遍历所有h2标签输出文本内容
for h2 in h2_list:
    print(h2.text)

补充说明

如果修改后仍然无法拿到全部标题,说明网站内容为前端动态渲染,静态请求无法获取完整DOM结构,需要改用Selenium等支持动态渲染的爬取工具。

内容的提问来源于stack exchange,提问作者CoderToBeWon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 04:06:06