Python BeautifulSoup爬虫:批量处理Billboard榜li元素及清洗换行符
解决方案
首先先修正原有代码的核心错误:你的CleanBullet函数内部错误使用了未定义的all_bullets[0]做元素查找,所有查找操作都应该替换为传入的参数bullet,否则无论传入什么元素,都会重复返回第一条榜单数据。
问题1:批量处理100个li元素得到完整DataFrame
推荐用「先收集所有数据到字典列表,最后一次性转DataFrame」的方案,比循环拼接DataFrame性能高很多,实现逻辑如下:
- 修正CleanBullet函数,改为返回单条数据字典而非小DataFrame
- 遍历full_table所有元素,调用函数收集所有数据
- 最后用pd.DataFrame直接把整个数据列表转成100行的大表
如果你一定要用原有返回小DataFrame的写法,也可以用pd.concat把所有返回的小df拼接起来。
问题2:rank列换行符无法去除的解决方法
strip('\n')只能去掉字符串首尾的换行符,如果换行符出现在中间、或者同时混有其他空白字符,直接用replace('\n', '')全局替换所有换行符即可,再配合无参数的strip()自动去掉所有首尾空白(包括空格、制表符等),就能彻底清除多余格式。
修正后完整可运行代码
from bs4 import BeautifulSoup import requests import pandas as pd def CleanBullet(bullet): # 替换all_bullets[0]为bullet,用replace全局清除换行符 this_rank = bullet.find("span", class_="chart-element__rank").get_text().replace('\n', '').strip('Rising').strip() this_song = bullet.find("span", class_="chart-element__information__song").get_text().replace('\n', '').strip() this_artist = bullet.find("span", class_="chart-element__information__artist").get_text().replace('\n', '').strip() this_last_week = bullet.find("span", class_="text--last").get_text().strip(' Last Week').strip() this_peak = bullet.find("span", class_="text--peak").get_text().strip(' Peak Rank').strip() this_weeks_on = bullet.find("span", class_="text--week").get_text().strip(' Weeks on Chart').strip() # 直接返回字典,不用每次生成小DataFrame return { 'rank': this_rank, 'song': this_song, 'artist': this_artist, 'last_week': this_last_week, 'peak': this_peak, 'weeks_on': this_weeks_on } base_url = "https://www.billboard.com/charts/hot-100/2021-10-30" response = requests.get(base_url) web_page = response.text soup = BeautifulSoup(web_page, "html.parser") full_table = soup.find("ol", class_="chart-list__elements").find_all("li") # 收集所有数据 all_data = [] for li in full_table: all_data.append(CleanBullet(li)) # 一次性转成完整DataFrame final_df = pd.DataFrame(all_data) print(final_df)
内容的提问来源于stack exchange,提问作者Canovice
相关产品推荐
相关产品推荐

