网页爬虫返回空值:爬取albumoftheyear.org得空DataFrame如何排查
问题排查顺序
- 先确认是否触发反爬拦截:在原代码请求的位置新增打印响应状态码和前1000位响应内容的逻辑,若返回403/503状态码,或响应内容包含Cloudflare、人机验证相关字段,就是被反爬拦截了。原生requests的默认UA会被Cloudflare直接识别,这是你当前空数据最可能的原因。
- 再确认DOM选择器是否正确:若响应状态码为200,且返回的是正常页面内容,直接在返回的HTML文本中搜索你使用的
albumListRow、albumListTitle等类名是否存在,确认是否是站点前端更新导致类名变更。
可用实现代码
首先给请求添加模拟浏览器的请求头绕过基础反爬,同时新增空值判断避免部分字段缺失导致代码中断:
import pandas as pd import requests from bs4 import BeautifulSoup import time # 模拟浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_5; rv:109.0) Gecko/20100101 Firefox/116.0", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8" } base_url = 'https://www.albumoftheyear.org/list/1500-rolling-stones-500-greatest-albums-of-all-time-2020/{}' result = [] for page in range(1, 11): resp = requests.get(base_url.format(page), headers=headers, timeout=15) if resp.status_code != 200: print(f"第{page}页请求失败,跳过") continue soup = BeautifulSoup(resp.text, "html.parser") album_items = soup.find_all(class_="albumListRow") for item in album_items: # 提取字段,增加容错判断 title_ele = item.find("h2", class_="albumListTitle") title = title_ele.find("a").get_text(strip=True) if title_ele and title_ele.find("a") else None date_ele = item.find("div", class_="albumListDate") release_year = date_ele.get_text(strip=True)[-4:] if date_ele else None genre_ele = item.find("div", class_="albumListGenre") genre = genre_ele.find("a").get_text(strip=True) if genre_ele and genre_ele.find("a") else None result.append({"title": title, "release_year": release_year, "genre": genre}) # 增加延时,避免请求频率过高被封 time.sleep(1) df = pd.DataFrame(result).dropna(subset=["title"]).reset_index(drop=True) print(f"共采集到{len(df)}条有效专辑数据") print(df.head())
额外注意事项
如果添加请求头后仍然被Cloudflare拦截,可使用
cloudscraper库替换requests,该库专门用于绕过Cloudflare基础验证,安装后仅需将请求部分替换为scraper = cloudscraper.create_scraper(),再用scraper.get()发起请求即可。
不要高频发起请求,建议单页请求间隔至少1秒,避免被站点封IP。
内容的提问来源于stack exchange,提问作者user12076260
相关产品推荐
相关产品推荐

