使用BeautifulSoup爬取Billboard榜单返回空列表问题求助
解决Billboard Hot100榜单爬取返回空列表的问题
问题原因
- 页面结构变更:你使用的
chart-element__information__song类名已被Billboard更新,当前页面的歌曲名称元素结构已不同。 - 反爬拦截:直接用
requests.get请求可能被网站识别为非浏览器访问,返回的内容不包含实际榜单数据。
修复方案
1. 添加请求头模拟浏览器
网站会检查请求的User-Agent字段,添加浏览器标识可避免被拦截:
from bs4 import BeautifulSoup import requests date = input("Which year do you want to travel to? Type the date in this format YYYY-MM-DD: ") # 添加浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(f"https://www.billboard.com/charts/hot-100/{date}", headers=headers) # 先确认请求成功,状态码应为200 print(response.status_code)
2. 更新元素选择器
当前Billboard Hot100页面中,歌曲名称位于h3标签,带有c-title a-no-trucate类,修改解析逻辑:
soup = BeautifulSoup(response.text, 'html.parser') # 使用更新后的选择器获取歌曲名 song_names = [song.get_text(strip=True) for song in soup.find_all("h3", class_="c-title a-no-trucate")] # 打印前10首验证 print(song_names[:10])
验证步骤
- 运行代码时先查看
response.status_code,如果返回200说明请求正常;若返回403/404,检查日期格式是否正确,或更换User-Agent值。 - 若仍无结果,可打开浏览器对应日期的榜单页面,右键检查元素,确认最新的歌曲名称标签和类名。
内容的提问来源于stack exchange,提问作者Adekeye Joshua
相关产品推荐
相关产品推荐

