Python网页爬虫返回空列表问题排查与解决方法
问题描述
我用Python构建网页爬虫,目标是爬取Hacker News上前10篇文章的标题和链接,但实际返回空列表。代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd def get_articles(): '''Gets the title and description of the first 10 articles on Hacker News.''' response = requests.get("https://news.ycombinator.com/") soup = BeautifulSoup(response.content, "html.parser") articles = soup.find_all("a", class_="storylink") return [{"title": article.text, "description": article["href"]} for article in articles[:10]] if __name__ == "__main__": articles = get_articles() print(articles) # df = pd.DataFrame(articles) # print(df) # df.to_csv("hacker_news_articles.csv")
预期返回包含前10篇文章信息的列表,但实际返回空列表,结果如下:
问题分析与解决方法
核心原因:请求被网站反爬机制拦截
Hacker News会校验请求的User-Agent标识,默认requests.get()使用的是python-requests/x.x.x这类标识,容易被识别为爬虫,导致返回的页面不包含目标内容。解决步骤:
- 在请求时添加模拟浏览器的User-Agent头部,示例代码修改如下:
import requests from bs4 import BeautifulSoup import pandas as pd def get_articles(): '''Gets the title and description of the first 10 articles on Hacker News.''' # 添加浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get("https://news.ycombinator.com/", headers=headers) # 先验证请求状态码,200表示请求成功 print(response.status_code) soup = BeautifulSoup(response.content, "html.parser") articles = soup.find_all("a", class_="storylink") return [{"title": article.text, "link": article["href"]} for article in articles[:10]] if __name__ == "__main__": articles = get_articles() print(articles) - 先打印
response.status_code确认请求状态,返回200说明请求正常;若返回403则需更换User-Agent值。
- 在请求时添加模拟浏览器的User-Agent头部,示例代码修改如下:
额外排查点:
- 用浏览器开发者工具(F12)查看Hacker News页面的最新元素结构,确认文章标题链接的类名是否仍为
storylink(网站可能会更新结构)。 - 打印
soup.prettify()查看返回的页面内容,确认是否包含目标元素,区分是请求被拦截还是解析逻辑出错。
- 用浏览器开发者工具(F12)查看Hacker News页面的最新元素结构,确认文章标题链接的类名是否仍为
内容的提问来源于stack exchange,提问作者Ibtihal Homadi
相关产品推荐
相关产品推荐

