You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup爬取Billboard歌手单曲的峰值日期与在榜周数

解决Billboard榜单数据爬取及存入DataFrame问题

问题原因

你之前用find_next()的方式定位元素容易受DOM结构变化影响,导致无法准确抓取peak date和wks字段。改用元素类名直接定位能大幅提升稳定性,因为Billboard的榜单表格元素都有明确的类标识。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 发起请求并解析页面
url = 'https://www.billboard.com/artist/john-lennon/chart-history/hsi/'
response = requests.get(url)
soup = BeautifulSoup(response.content, 'html.parser')
rows = soup.find_all('div', 'o-chart-results-list-row')

# 初始化列表存储数据
chart_data = []

for row in rows:
    # 抓取各字段,用类名定位更可靠
    song = row.find('h3', class_='c-title').text.strip()
    artist = row.find('span', class_='c-label').text.strip()
    debut_date = row.find('span', class_='o-chart-results-list__item--date').text.strip()
    peak_rank = row.find('span', class_='o-chart-results-list__item--peak').text.strip()
    peak_date = row.find('span', class_='o-chart-results-list__item--peak-date').text.strip()
    weeks_on_chart = row.find('span', class_='o-chart-results-list__item--weeks-on-chart').text.strip()
    
    # 将单条数据存入字典,再添加到列表
    chart_data.append({
        '歌曲名': song,
        '歌手': artist,
        '上榜日期': debut_date,
        '峰值排名': peak_rank,
        '峰值日期': peak_date,
        '在榜周数': weeks_on_chart
    })

# 转换为DataFrame
df = pd.DataFrame(chart_data)
print(df.head())
# 可选:保存为CSV文件
# df.to_csv('john_lennon_billboard_chart.csv', index=False, encoding='utf-8-sig')

关键说明

  • 所有字段都通过类名精准定位,避免了find_next()链式调用的定位误差
  • 用列表+字典的方式统一收集数据,最后转换为DataFrame,符合数据分析的常规流程
  • 可直接通过to_csv()方法将数据保存为本地文件,方便后续分析

内容的提问来源于stack exchange,提问作者rar9000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:02:37