使用Requests和Beautiful Soup爬取MyAnimeList用户评分失败求助
问题原因
MyAnimeList的用户动画列表采用动态渲染机制,直接通过requests获取的静态HTML仅包含前端模板占位符(如${ item.title_localized || item.anime_title }),而非实际的动画标题和用户评分数据。这些真实数据是页面加载完成后,通过JavaScript从后端接口异步拉取的,所以静态爬取方式无法获取有效内容。
解决方案
推荐两种可行方案:
方案一:调用官方API(稳定合规)
MyAnimeList提供官方API,可合法获取用户动画列表数据,无需处理动态渲染问题。需先注册开发者账号获取API密钥。
示例代码:
import requests import csv CLIENT_ID = "你的API密钥" # 替换为你的开发者客户端ID def scrape_user_profile(username): url = f"https://api.myanimelist.net/v2/users/{username}/animelist?fields=list_status&limit=1000" headers = {"X-MAL-CLIENT-ID": CLIENT_ID} response = requests.get(url, headers=headers) if response.status_code == 200: data = response.json() anime_list = data.get("data", []) result = [] for item in anime_list: anime_title = item["node"]["title"] score = item["list_status"]["score"] score = score if score != 0 else "-" result.append([username, anime_title, score]) if result: with open('user_score.csv', 'w', newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(["用户名", "动画标题", "评分"]) writer.writerows(result) print(f"已保存用户 {username} 的评分数据") else: print(f"未找到用户 {username} 的动画列表") else: print(f"获取用户 {username} 数据失败,状态码: {response.status_code}") usernames = ["Arcane"] for username in usernames: scrape_user_profile(username)
方案二:用Selenium模拟浏览器加载
通过模拟浏览器行为,等待页面动态渲染完成后再提取数据。需安装selenium库及对应浏览器驱动(如ChromeDriver)。
示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import csv def scrape_user_profile(username): url = f"https://myanimelist.net/animelist/{username}" driver = webdriver.Chrome() # 确保ChromeDriver路径配置正确 driver.get(url) try: # 等待列表元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "list-table-data")) ) anime_rows = driver.find_elements(By.CLASS_NAME, "list-table-data") result = [] for row in anime_rows: title_element = row.find_element(By.CSS_SELECTOR, "td.data.title.clearfix a.link.sort") score_element = row.find_element(By.CSS_SELECTOR, "td.data.score span.score-label") title = title_element.text.strip() score = score_element.text.strip() if score_element.text.strip() else "-" result.append([username, title, score]) if result: with open('user_score.csv', 'w', newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(["用户名", "动画标题", "评分"]) writer.writerows(result) print(f"已保存用户 {username} 的评分数据") else: print(f"未找到用户 {username} 的动画列表") finally: driver.quit() usernames = ["Arcane"] for username in usernames: scrape_user_profile(username)
注意事项
- 使用官方API需遵守平台使用条款,控制请求频率,避免触发限制。
- 使用Selenium时,需保证浏览器驱动版本与本地浏览器版本匹配。
内容的提问来源于stack exchange,提问作者DBD Mobile
相关产品推荐
相关产品推荐

