You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫爬取Stack Overflow问题返回空列表的求助

解决Stack Overflow爬虫返回空列表的问题

核心原因分析

  1. 反爬拦截:直接用requests.get发起请求会被识别为非浏览器请求,返回的内容可能不是正常页面结构。
  2. 页面结构变更:你使用的.question-summary类名已被Stack Overflow弃用,当前问题卡片的类名已更新。

解决步骤

1. 验证请求有效性

先打印请求状态码,确认是否被拦截:

print(response.status_code)

如果返回403,说明被反爬机制拦截,需要添加请求头模拟浏览器。

2. 添加请求头绕过基础反爬

给requests.get加上headers参数,模拟浏览器请求:

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
response = requests.get("https://www.stackoverflow.com/questions", headers=headers)

3. 更新选择器适配新页面结构

当前Stack Overflow的问题列表项使用.s-post-summary类,替换原有选择器:

questions = soup.select(".s-post-summary")

修正后的完整代码

import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
response = requests.get("https://www.stackoverflow.com/questions", headers=headers)

print("请求状态码:", response.status_code)
soup = BeautifulSoup(response.text, "html.parser")
questions = soup.select(".s-post-summary")

print("抓取到的问题数量:", len(questions))
if questions:
    print("第一个问题内容:", questions[0].get_text(strip=True))

新手排查技巧

  • 若仍为空列表,可将response.text写入本地HTML文件,打开查看实际返回内容是否正常。
  • 用浏览器开发者工具(F12)查看目标元素的最新类名或选择器,确保选择器与页面结构匹配。

内容的提问来源于stack exchange,提问作者David Henson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 03:46:04