You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫返回空列表问题排查与解决方法

问题描述

我用Python构建网页爬虫,目标是爬取Hacker News上前10篇文章的标题和链接,但实际返回空列表。代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd


def get_articles():
    '''Gets the title and description of the first 10 articles on Hacker News.'''
    response = requests.get("https://news.ycombinator.com/")
    soup = BeautifulSoup(response.content, "html.parser")
    articles = soup.find_all("a", class_="storylink")
    return [{"title": article.text, "description": article["href"]} for article in articles[:10]]


if __name__ == "__main__":
    articles = get_articles()
    print(articles)
    # df = pd.DataFrame(articles)
    # print(df)
    # df.to_csv("hacker_news_articles.csv")

预期返回包含前10篇文章信息的列表,但实际返回空列表,结果如下:
爬取结果为空列表

问题分析与解决方法
  • 核心原因:请求被网站反爬机制拦截
    Hacker News会校验请求的User-Agent标识,默认requests.get()使用的是python-requests/x.x.x这类标识,容易被识别为爬虫,导致返回的页面不包含目标内容。

  • 解决步骤:

    1. 在请求时添加模拟浏览器的User-Agent头部,示例代码修改如下:
      import requests
      from bs4 import BeautifulSoup
      import pandas as pd
      
      
      def get_articles():
          '''Gets the title and description of the first 10 articles on Hacker News.'''
          # 添加浏览器请求头
          headers = {
              "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
          }
          response = requests.get("https://news.ycombinator.com/", headers=headers)
          # 先验证请求状态码,200表示请求成功
          print(response.status_code)
          soup = BeautifulSoup(response.content, "html.parser")
          articles = soup.find_all("a", class_="storylink")
          return [{"title": article.text, "link": article["href"]} for article in articles[:10]]
      
      
      if __name__ == "__main__":
          articles = get_articles()
          print(articles)
      
    2. 先打印response.status_code确认请求状态,返回200说明请求正常;若返回403则需更换User-Agent值。
  • 额外排查点:

    • 用浏览器开发者工具(F12)查看Hacker News页面的最新元素结构,确认文章标题链接的类名是否仍为storylink(网站可能会更新结构)。
    • 打印soup.prettify()查看返回的页面内容,确认是否包含目标元素,区分是请求被拦截还是解析逻辑出错。

内容的提问来源于stack exchange,提问作者Ibtihal Homadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 10:25:04