You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取MAL季番仅获单条数据的技术问题

问题描述

我开发了一个Python程序,使用BeautifulSoup和Requests爬取MyAnimeList(MAL)的季番数据,提取标题、评分等信息并写入CSV文件。程序整体功能正常,但存在问题:尽管所有目标番剧数据都隶属于同一个父类别,却仅能爬取到单条番剧数据。

我尝试将代码中list.find(第58行)改为list.find_all,此时可获取全部数据集,但无法再使用.text方法处理数据。

以下是我的代码:

#Beautiful soup or BS4 is  a package I will be using to allow me to parse the HTML data which I will be retrieving from a website.
#parsing is te conversion of codes from machine language into a code which humans can understand and allow it to be structured. 
#(Converting data from one format to another) with BS4
from bs4 import BeautifulSoup


#requests is an HTTP Library which allows me to send requests to websites the retrieve date using Python. This is helpful as 
#The website is writtin in a different language so it allows me to retrieve what I want and read it as well. 
import requests
#import writer


from csv import writer
#defining the website which I will be retrieving my code

url= "https://myanimelist.net/anime/season"


#requesting to get data using 'requests' and gain acess as well. 
#hadve to check the response before moving forward to ensure there is no problem retrieving data. 

page= requests.get(url)
#print(page)
#<Response [200]> response was "200" meaning "Successful responses"



soup = BeautifulSoup(page.content, 'html.parser')
#here i retrieve my 
#for this to identify the html code and determine what we will be producing(retrieveing data) for each item on the page we had to 
#find the parent category which contains all the info we need to make our data categories. 
lists = soup.find_all('div', class_="js-seasonal-anime-list-key-1")
#we add _ after class  to make class_ because without the underscore the program identifies it as a python class 
# when really it is more of a cs class

#this allows us to create and close a csv file. using 'w' to allow editing
with open('shows.csv', 'w', encoding='utf8', newline='')as f:


#will write onto our file 'f'
    writing=writer(f)
#organizing chart
    header=['Title', 'Show Rating', 'Members', 'Release Date']
    
#use our writer to write a row in file
    writing.writerow(header)    
    
#must create loop to find titles seperate as there are alot that will come up
    for list in lists:
    #identify and find class which includes the title of the shows, show ratings, members watching, and episodes
    #added .text.replace in order to get rid of the|n spacing which was in html format
        title= list.find('a', class_="link-title").text.replace('\n', '')
        rating= list.find('div', class_="score").text.replace('\n', '')
        members= list.find('div', class_="scormem-item member").text.replace('\n', '')
        release_date= list.find('span', class_="item").text.replace('\n', '')
       
    #testing for errors and makins sure locations are correct to withdraw/request the data
    info= [title, rating, members, release_date]
    writing.writerow(info) 
解决方案

你的代码存在两个核心问题:

  1. 循环缩进错误:info的定义和writing.writerow(info)语句不在for循环的缩进块内,导致循环结束后只写入最后一条番剧的数据,而非每条循环都执行写入操作。
  2. 父元素选择偏差:js-seasonal-anime-list-key-1是包含多组番剧的大容器,而非单个番剧的独立节点。正确的单番剧容器应为div.js-seasonal-anime,这样find_all才能获取到每个番剧的单独节点。

修正后的代码如下:

from bs4 import BeautifulSoup
import requests
from csv import writer

url = "https://myanimelist.net/anime/season"
page = requests.get(url)
soup = BeautifulSoup(page.content, 'html.parser')

# 选择每个番剧的独立容器节点
anime_items = soup.find_all('div', class_="js-seasonal-anime")

with open('shows.csv', 'w', encoding='utf8', newline='') as f:
    writing = writer(f)
    header = ['Title', 'Show Rating', 'Members', 'Release Date']
    writing.writerow(header)    
    
    for item in anime_items:
        # 增加空值判断,避免页面元素缺失导致程序报错
        title = item.find('a', class_="link-title").text.strip() if item.find('a', class_="link-title") else "N/A"
        rating = item.find('div', class_="score").text.strip() if item.find('div', class_="score") else "N/A"
        members = item.find('div', class_="scormem-item member").text.strip() if item.find('div', class_="scormem-item member") else "N/A"
        release_date = item.find('span', class_="item").text.strip() if item.find('span', class_="item") else "N/A"
        
        info = [title, rating, members, release_date]
        writing.writerow(info) 

关键改进点

  • 修正缩进逻辑,确保每条番剧数据都被写入CSV
  • 更换准确的番剧容器选择器,保证find_all获取所有单条番剧节点
  • 添加空值判断,避免页面元素缺失导致程序崩溃
  • 使用.strip()替代replace('\n', ''),更高效清除多余换行和空格

内容的提问来源于stack exchange,提问作者user20784428

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 05:20:22