You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python BeautifulSoup4抓取Reddit动漫讨论存档的关联链接信息

Got it, let's tackle this problem step by step! I'll walk you through how to traverse that Reddit wiki page's structure properly to link each discussion URL with its corresponding year, quarter, anime title, and episode info.

Step 1: Break Down the Page Structure

First, let's map out how the content is organized on that 2018 archive page:

  • The page is split into quarter blocks (e.g., Q1 2018, Q2 2018), marked by <h3> headers.
  • Each quarter block has a sibling <ul> containing all anime entries for that period.
  • Each anime entry is a <li> with a bolded title, followed by a nested <ul> holding all episode discussion links.

Step 2: Complete Working Code

Here's the full code with comments explaining each part. I've added safeguards for edge cases (like missing tags) and handled relative URLs:

import requests
from bs4 import BeautifulSoup

# Target Reddit archive page
url = "https://www.reddit.com/r/anime/wiki/discussion_archive/2018"

# Add a User-Agent to avoid being blocked by Reddit's anti-bot measures
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

# Fetch and parse the page
try:
    response = requests.get(url, headers=headers)
    response.raise_for_status()  # Throw error if request fails
    soup = BeautifulSoup(response.text, "html.parser")
except requests.exceptions.RequestException as e:
    print(f"Request failed: {e}")
    exit()

# Store all collected discussion data here
discussion_data = []

# Traverse each quarter block
for quarter_header in soup.find_all("h3"):
    # Extract quarter and year from the header text (e.g., "Q1 2018" → ("Q1", 2018))
    quarter_text = quarter_header.get_text(strip=True)
    if len(quarter_text.split()) != 2:
        continue  # Skip any malformed headers
    quarter, year = quarter_text.split()
    year = int(year)

    # Get the list of anime entries for this quarter (next sibling <ul>)
    anime_list = quarter_header.find_next_sibling("ul")
    if not anime_list:
        continue  # Skip if no anime list exists for this quarter

    # Traverse each anime entry in the quarter
    for anime_item in anime_list.find_all("li", recursive=False):
        # Extract the anime title from the bolded tag
        title_tag = anime_item.find("strong")
        if not title_tag:
            continue
        anime_title = title_tag.get_text(strip=True)

        # Get the nested list of episode discussions
        episode_list = anime_item.find("ul")
        if not episode_list:
            continue

        # Traverse each episode discussion link
        for episode_item in episode_list.find_all("li"):
            link_tag = episode_item.find("a")
            if not link_tag:
                continue

            # Extract episode info and clean up the URL
            episode_info = link_tag.get_text(strip=True)
            discussion_url = link_tag["href"]
            # Convert relative URLs to full Reddit URLs
            if not discussion_url.startswith("http"):
                discussion_url = f"https://www.reddit.com{discussion_url}"

            # Add all data to our list
            discussion_data.append({
                "year": year,
                "quarter": quarter,
                "anime_title": anime_title,
                "episode_info": episode_info,
                "discussion_url": discussion_url
            })

# Example: Print the first 5 entries to verify
print("Sample Results:")
for entry in discussion_data[:5]:
    print(f"\nYear: {entry['year']} | Quarter: {entry['quarter']}")
    print(f"Anime: {entry['anime_title']}")
    print(f"Episode: {entry['episode_info']}")
    print(f"Link: {entry['discussion_url']}")

Step 3: Key Explanations

  • User-Agent Header: Reddit blocks requests without a proper user agent, so we mimic a browser to avoid being flagged as a bot.
  • Recursive=False: When fetching anime entries, we use recursive=False to ensure we only grab direct child <li> elements (not nested episode links).
  • Relative URL Handling: Some links on the page are relative, so we prepend the Reddit domain to make them full, usable URLs.
  • Error Handling: Basic try/except blocks and checks for missing tags ensure the code doesn't crash if the page has unexpected formatting.

Notes for Maintenance

If the page structure changes in the future, you'll need to re-inspect the HTML (using your browser's dev tools) and adjust the tag selectors (e.g., if headers switch from <h3> to <h2>).

内容的提问来源于stack exchange,提问作者alpacafondue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:02:56