如何用BeautifulSoup爬取Billboard歌手单曲的峰值日期与在榜周数
解决Billboard榜单数据爬取及存入DataFrame问题
问题原因
你之前用find_next()的方式定位元素容易受DOM结构变化影响,导致无法准确抓取peak date和wks字段。改用元素类名直接定位能大幅提升稳定性,因为Billboard的榜单表格元素都有明确的类标识。
修改后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd # 发起请求并解析页面 url = 'https://www.billboard.com/artist/john-lennon/chart-history/hsi/' response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') rows = soup.find_all('div', 'o-chart-results-list-row') # 初始化列表存储数据 chart_data = [] for row in rows: # 抓取各字段,用类名定位更可靠 song = row.find('h3', class_='c-title').text.strip() artist = row.find('span', class_='c-label').text.strip() debut_date = row.find('span', class_='o-chart-results-list__item--date').text.strip() peak_rank = row.find('span', class_='o-chart-results-list__item--peak').text.strip() peak_date = row.find('span', class_='o-chart-results-list__item--peak-date').text.strip() weeks_on_chart = row.find('span', class_='o-chart-results-list__item--weeks-on-chart').text.strip() # 将单条数据存入字典,再添加到列表 chart_data.append({ '歌曲名': song, '歌手': artist, '上榜日期': debut_date, '峰值排名': peak_rank, '峰值日期': peak_date, '在榜周数': weeks_on_chart }) # 转换为DataFrame df = pd.DataFrame(chart_data) print(df.head()) # 可选:保存为CSV文件 # df.to_csv('john_lennon_billboard_chart.csv', index=False, encoding='utf-8-sig')
关键说明
- 所有字段都通过类名精准定位,避免了
find_next()链式调用的定位误差 - 用列表+字典的方式统一收集数据,最后转换为DataFrame,符合数据分析的常规流程
- 可直接通过
to_csv()方法将数据保存为本地文件,方便后续分析
内容的提问来源于stack exchange,提问作者rar9000
相关产品推荐
相关产品推荐

