You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬虫获取MyAnimeList悬浮窗中的动漫流派信息?

解决MyAnimeList悬浮窗流派抓取问题

你用requests获取的是页面初始HTML,但鼠标悬停显示的悬浮窗内容是动态加载的——只有当鼠标触发hover事件时,网站才会通过AJAX请求拉取对应动漫的详情数据,所以初始HTML里根本没有<span class="dark_text">这类元素,自然用BeautifulSoup找不到。

下面给两种可行的解决方案:

方案一:模拟AJAX请求抓取数据

MyAnimeList的悬浮窗数据来自专门的AJAX接口,每部动漫的详情可通过其ID直接请求:

  1. 先从top页面提取每部动漫的ID(从标题链接中截取,比如https://myanimelist.net/anime/1/Cowboy_Bebop里的1就是ID)
  2. 构造请求URL:https://myanimelist.net/anime/{id}/_ajax/synopsis
  3. 对每个ID发送请求,解析返回的HTML提取流派

代码示例:

import requests
from bs4 import BeautifulSoup

# 请求头,模拟浏览器访问
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 1. 获取top页面的动漫ID列表
top_url = "https://myanimelist.net/topanime.php"
response = requests.get(top_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

anime_ids = []
for title_link in soup.select(".ranking-list .title a"):
    anime_id = title_link["href"].split("/")[4]
    anime_ids.append(anime_id)

# 2. 遍历ID,请求悬浮窗数据并提取流派
anime_genres = []
for anime_id in anime_ids[:5]:  # 先测试前5个,避免请求过于频繁
    ajax_url = f"https://myanimelist.net/anime/{anime_id}/_ajax/synopsis"
    ajax_response = requests.get(ajax_url, headers=headers)
    ajax_soup = BeautifulSoup(ajax_response.text, "html.parser")
    
    # 定位流派标签并提取内容
    genre_label = ajax_soup.find("span", class_="dark_text", string="Genres:")
    if genre_label:
        genre_text = genre_label.next_sibling.strip()
        anime_genres.append({"anime_id": anime_id, "genres": genre_text})

print(anime_genres)

方案二:用Selenium模拟浏览器行为

如果不想分析AJAX接口,可以用Selenium模拟鼠标悬停操作,等待悬浮窗加载后再抓取:

from selenium import webdriver
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# 初始化浏览器驱动(需提前安装对应浏览器的驱动,比如ChromeDriver)
driver = webdriver.Chrome()
driver.get("https://myanimelist.net/topanime.php")
wait = WebDriverWait(driver, 10)

anime_genres = []
# 定位所有动漫标题元素
title_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".ranking-list .title a")))

for title_elem in title_elements[:5]:  # 测试前5个
    # 模拟鼠标悬停触发悬浮窗加载
    ActionChains(driver).move_to_element(title_elem).perform()
    time.sleep(0.5)  # 给悬浮窗加载留时间
    
    # 获取悬浮窗的HTML内容
    popup = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, "tooltipster-content")))
    popup_html = popup.get_attribute("innerHTML")
    soup = BeautifulSoup(popup_html, "html.parser")
    
    # 提取流派信息
    genre_label = soup.find("span", class_="dark_text", string="Genres:")
    if genre_label:
        genre_text = genre_label.next_sibling.strip()
        anime_genres.append({"title": title_elem.text, "genres": genre_text})

# 关闭浏览器
driver.quit()
print(anime_genres)

注意事项

  • 务必添加User-Agent请求头,避免被网站识别为爬虫拦截
  • 批量请求时建议添加延迟(比如time.sleep(1)),避免触发反爬机制
  • 使用Selenium时,需确保浏览器驱动版本与浏览器版本一致

内容的提问来源于stack exchange,提问作者codeledger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 02:49:55