You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页解析问题求助:爬取pointercrate.com前十页数据失败

问题分析与解决

你的代码存在两个核心问题:

  1. URL构造逻辑错误:该网站的分页是路径后缀形式(如demonlist/1),并非查询参数。你用params传递参数会生成https://pointercrate.com/?demonlist/=1这类错误URL,而实际需要的是https://pointercrate.com/demonlist/1。
  2. 缺少必要请求头:部分网站会拦截无浏览器标识的请求,导致返回的页面内容不完整,最终BeautifulSoup找不到目标元素,输出None。

修正后的代码

import requests
from bs4 import BeautifulSoup

# 基础URL模板
base_url = "https://pointercrate.com/demonlist/{page}"
# 模拟浏览器请求头,避免被拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 遍历前10页
for page_num in range(1, 11):
    # 生成当前页的完整URL
    current_url = base_url.format(page=page_num)
    # 发送请求
    response = requests.get(current_url, headers=headers)
    # 解析页面内容
    soup = BeautifulSoup(response.text, 'html.parser')
    # 获取目标标题元素
    demon_title = soup.find('h1', id='demon-heading')
    # 输出结果,处理元素不存在的情况
    print(demon_title.get_text(strip=True) if demon_title else "页面未找到目标标题")

关键说明

  • URL拼接:用格式化字符串直接将页码嵌入路径,生成符合网站规则的分页地址。
  • 请求头设置:添加User-Agent模拟浏览器访问,确保返回完整的页面源码。
  • 异常防护:增加对目标元素是否存在的判断,避免因页面结构变化导致代码报错。

内容的提问来源于stack exchange,提问作者Visible Happiness

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 13:52:39