You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫返回空值:爬取albumoftheyear.org得空DataFrame如何排查

问题排查顺序
  • 先确认是否触发反爬拦截:在原代码请求的位置新增打印响应状态码和前1000位响应内容的逻辑,若返回403/503状态码,或响应内容包含Cloudflare、人机验证相关字段,就是被反爬拦截了。原生requests的默认UA会被Cloudflare直接识别,这是你当前空数据最可能的原因。
  • 再确认DOM选择器是否正确:若响应状态码为200,且返回的是正常页面内容,直接在返回的HTML文本中搜索你使用的albumListRow、albumListTitle等类名是否存在,确认是否是站点前端更新导致类名变更。
可用实现代码

首先给请求添加模拟浏览器的请求头绕过基础反爬,同时新增空值判断避免部分字段缺失导致代码中断:

import pandas as pd
import requests
from bs4 import BeautifulSoup
import time

# 模拟浏览器请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_5; rv:109.0) Gecko/20100101 Firefox/116.0",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"
}

base_url = 'https://www.albumoftheyear.org/list/1500-rolling-stones-500-greatest-albums-of-all-time-2020/{}'
result = []

for page in range(1, 11):
    resp = requests.get(base_url.format(page), headers=headers, timeout=15)
    if resp.status_code != 200:
        print(f"第{page}页请求失败,跳过")
        continue
    soup = BeautifulSoup(resp.text, "html.parser")
    album_items = soup.find_all(class_="albumListRow")
    for item in album_items:
        # 提取字段,增加容错判断
        title_ele = item.find("h2", class_="albumListTitle")
        title = title_ele.find("a").get_text(strip=True) if title_ele and title_ele.find("a") else None
        date_ele = item.find("div", class_="albumListDate")
        release_year = date_ele.get_text(strip=True)[-4:] if date_ele else None
        genre_ele = item.find("div", class_="albumListGenre")
        genre = genre_ele.find("a").get_text(strip=True) if genre_ele and genre_ele.find("a") else None
        result.append({"title": title, "release_year": release_year, "genre": genre})
    # 增加延时,避免请求频率过高被封
    time.sleep(1)

df = pd.DataFrame(result).dropna(subset=["title"]).reset_index(drop=True)
print(f"共采集到{len(df)}条有效专辑数据")
print(df.head())
额外注意事项

如果添加请求头后仍然被Cloudflare拦截,可使用cloudscraper库替换requests,该库专门用于绕过Cloudflare基础验证,安装后仅需将请求部分替换为scraper = cloudscraper.create_scraper(),再用scraper.get()发起请求即可。
不要高频发起请求,建议单页请求间隔至少1秒,避免被站点封IP。

内容的提问来源于stack exchange,提问作者user12076260

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 20:54:03