You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup爬虫:批量处理Billboard榜li元素及清洗换行符

解决方案

首先先修正原有代码的核心错误:你的CleanBullet函数内部错误使用了未定义的all_bullets[0]做元素查找,所有查找操作都应该替换为传入的参数bullet,否则无论传入什么元素,都会重复返回第一条榜单数据。


问题1:批量处理100个li元素得到完整DataFrame

推荐用「先收集所有数据到字典列表,最后一次性转DataFrame」的方案,比循环拼接DataFrame性能高很多,实现逻辑如下:

  1. 修正CleanBullet函数,改为返回单条数据字典而非小DataFrame
  2. 遍历full_table所有元素,调用函数收集所有数据
  3. 最后用pd.DataFrame直接把整个数据列表转成100行的大表

如果你一定要用原有返回小DataFrame的写法,也可以用pd.concat把所有返回的小df拼接起来。


问题2:rank列换行符无法去除的解决方法

strip('\n')只能去掉字符串首尾的换行符,如果换行符出现在中间、或者同时混有其他空白字符,直接用replace('\n', '')全局替换所有换行符即可,再配合无参数的strip()自动去掉所有首尾空白(包括空格、制表符等),就能彻底清除多余格式。


修正后完整可运行代码

from bs4 import BeautifulSoup 
import requests
import pandas as pd

def CleanBullet(bullet):
    # 替换all_bullets[0]为bullet,用replace全局清除换行符
    this_rank = bullet.find("span", class_="chart-element__rank").get_text().replace('\n', '').strip('Rising').strip()
    this_song = bullet.find("span", class_="chart-element__information__song").get_text().replace('\n', '').strip()
    this_artist = bullet.find("span", class_="chart-element__information__artist").get_text().replace('\n', '').strip()
    this_last_week = bullet.find("span", class_="text--last").get_text().strip(' Last Week').strip()
    this_peak = bullet.find("span", class_="text--peak").get_text().strip(' Peak Rank').strip()
    this_weeks_on = bullet.find("span", class_="text--week").get_text().strip(' Weeks on Chart').strip()

    # 直接返回字典,不用每次生成小DataFrame
    return {
        'rank': this_rank,
        'song': this_song,
        'artist': this_artist,
        'last_week': this_last_week,
        'peak': this_peak,
        'weeks_on': this_weeks_on
    }


base_url = "https://www.billboard.com/charts/hot-100/2021-10-30"
response = requests.get(base_url)
web_page = response.text
soup = BeautifulSoup(web_page, "html.parser")    
full_table = soup.find("ol", class_="chart-list__elements").find_all("li")

# 收集所有数据
all_data = []
for li in full_table:
    all_data.append(CleanBullet(li))

# 一次性转成完整DataFrame
final_df = pd.DataFrame(all_data)
print(final_df)

内容的提问来源于stack exchange,提问作者Canovice

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 19:54:04