使用BeautifulSoup爬取MAL季番仅获单条数据的技术问题
问题描述
我开发了一个Python程序,使用BeautifulSoup和Requests爬取MyAnimeList(MAL)的季番数据,提取标题、评分等信息并写入CSV文件。程序整体功能正常,但存在问题:尽管所有目标番剧数据都隶属于同一个父类别,却仅能爬取到单条番剧数据。
我尝试将代码中list.find(第58行)改为list.find_all,此时可获取全部数据集,但无法再使用.text方法处理数据。
以下是我的代码:
#Beautiful soup or BS4 is a package I will be using to allow me to parse the HTML data which I will be retrieving from a website. #parsing is te conversion of codes from machine language into a code which humans can understand and allow it to be structured. #(Converting data from one format to another) with BS4 from bs4 import BeautifulSoup #requests is an HTTP Library which allows me to send requests to websites the retrieve date using Python. This is helpful as #The website is writtin in a different language so it allows me to retrieve what I want and read it as well. import requests #import writer from csv import writer #defining the website which I will be retrieving my code url= "https://myanimelist.net/anime/season" #requesting to get data using 'requests' and gain acess as well. #hadve to check the response before moving forward to ensure there is no problem retrieving data. page= requests.get(url) #print(page) #<Response [200]> response was "200" meaning "Successful responses" soup = BeautifulSoup(page.content, 'html.parser') #here i retrieve my #for this to identify the html code and determine what we will be producing(retrieveing data) for each item on the page we had to #find the parent category which contains all the info we need to make our data categories. lists = soup.find_all('div', class_="js-seasonal-anime-list-key-1") #we add _ after class to make class_ because without the underscore the program identifies it as a python class # when really it is more of a cs class #this allows us to create and close a csv file. using 'w' to allow editing with open('shows.csv', 'w', encoding='utf8', newline='')as f: #will write onto our file 'f' writing=writer(f) #organizing chart header=['Title', 'Show Rating', 'Members', 'Release Date'] #use our writer to write a row in file writing.writerow(header) #must create loop to find titles seperate as there are alot that will come up for list in lists: #identify and find class which includes the title of the shows, show ratings, members watching, and episodes #added .text.replace in order to get rid of the|n spacing which was in html format title= list.find('a', class_="link-title").text.replace('\n', '') rating= list.find('div', class_="score").text.replace('\n', '') members= list.find('div', class_="scormem-item member").text.replace('\n', '') release_date= list.find('span', class_="item").text.replace('\n', '') #testing for errors and makins sure locations are correct to withdraw/request the data info= [title, rating, members, release_date] writing.writerow(info)
解决方案
你的代码存在两个核心问题:
- 循环缩进错误:
info的定义和writing.writerow(info)语句不在for循环的缩进块内,导致循环结束后只写入最后一条番剧的数据,而非每条循环都执行写入操作。 - 父元素选择偏差:
js-seasonal-anime-list-key-1是包含多组番剧的大容器,而非单个番剧的独立节点。正确的单番剧容器应为div.js-seasonal-anime,这样find_all才能获取到每个番剧的单独节点。
修正后的代码如下:
from bs4 import BeautifulSoup import requests from csv import writer url = "https://myanimelist.net/anime/season" page = requests.get(url) soup = BeautifulSoup(page.content, 'html.parser') # 选择每个番剧的独立容器节点 anime_items = soup.find_all('div', class_="js-seasonal-anime") with open('shows.csv', 'w', encoding='utf8', newline='') as f: writing = writer(f) header = ['Title', 'Show Rating', 'Members', 'Release Date'] writing.writerow(header) for item in anime_items: # 增加空值判断,避免页面元素缺失导致程序报错 title = item.find('a', class_="link-title").text.strip() if item.find('a', class_="link-title") else "N/A" rating = item.find('div', class_="score").text.strip() if item.find('div', class_="score") else "N/A" members = item.find('div', class_="scormem-item member").text.strip() if item.find('div', class_="scormem-item member") else "N/A" release_date = item.find('span', class_="item").text.strip() if item.find('span', class_="item") else "N/A" info = [title, rating, members, release_date] writing.writerow(info)
关键改进点
- 修正缩进逻辑,确保每条番剧数据都被写入CSV
- 更换准确的番剧容器选择器,保证
find_all获取所有单条番剧节点 - 添加空值判断,避免页面元素缺失导致程序崩溃
- 使用
.strip()替代replace('\n', ''),更高效清除多余换行和空格
内容的提问来源于stack exchange,提问作者user20784428
相关产品推荐
相关产品推荐

