使用Python BeautifulSoup4抓取Reddit动漫讨论存档的关联链接信息
Got it, let's tackle this problem step by step! I'll walk you through how to traverse that Reddit wiki page's structure properly to link each discussion URL with its corresponding year, quarter, anime title, and episode info.
Step 1: Break Down the Page Structure
First, let's map out how the content is organized on that 2018 archive page:
- The page is split into quarter blocks (e.g., Q1 2018, Q2 2018), marked by
<h3>headers. - Each quarter block has a sibling
<ul>containing all anime entries for that period. - Each anime entry is a
<li>with a bolded title, followed by a nested<ul>holding all episode discussion links.
Step 2: Complete Working Code
Here's the full code with comments explaining each part. I've added safeguards for edge cases (like missing tags) and handled relative URLs:
import requests from bs4 import BeautifulSoup # Target Reddit archive page url = "https://www.reddit.com/r/anime/wiki/discussion_archive/2018" # Add a User-Agent to avoid being blocked by Reddit's anti-bot measures headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # Fetch and parse the page try: response = requests.get(url, headers=headers) response.raise_for_status() # Throw error if request fails soup = BeautifulSoup(response.text, "html.parser") except requests.exceptions.RequestException as e: print(f"Request failed: {e}") exit() # Store all collected discussion data here discussion_data = [] # Traverse each quarter block for quarter_header in soup.find_all("h3"): # Extract quarter and year from the header text (e.g., "Q1 2018" → ("Q1", 2018)) quarter_text = quarter_header.get_text(strip=True) if len(quarter_text.split()) != 2: continue # Skip any malformed headers quarter, year = quarter_text.split() year = int(year) # Get the list of anime entries for this quarter (next sibling <ul>) anime_list = quarter_header.find_next_sibling("ul") if not anime_list: continue # Skip if no anime list exists for this quarter # Traverse each anime entry in the quarter for anime_item in anime_list.find_all("li", recursive=False): # Extract the anime title from the bolded tag title_tag = anime_item.find("strong") if not title_tag: continue anime_title = title_tag.get_text(strip=True) # Get the nested list of episode discussions episode_list = anime_item.find("ul") if not episode_list: continue # Traverse each episode discussion link for episode_item in episode_list.find_all("li"): link_tag = episode_item.find("a") if not link_tag: continue # Extract episode info and clean up the URL episode_info = link_tag.get_text(strip=True) discussion_url = link_tag["href"] # Convert relative URLs to full Reddit URLs if not discussion_url.startswith("http"): discussion_url = f"https://www.reddit.com{discussion_url}" # Add all data to our list discussion_data.append({ "year": year, "quarter": quarter, "anime_title": anime_title, "episode_info": episode_info, "discussion_url": discussion_url }) # Example: Print the first 5 entries to verify print("Sample Results:") for entry in discussion_data[:5]: print(f"\nYear: {entry['year']} | Quarter: {entry['quarter']}") print(f"Anime: {entry['anime_title']}") print(f"Episode: {entry['episode_info']}") print(f"Link: {entry['discussion_url']}")
Step 3: Key Explanations
- User-Agent Header: Reddit blocks requests without a proper user agent, so we mimic a browser to avoid being flagged as a bot.
- Recursive=False: When fetching anime entries, we use
recursive=Falseto ensure we only grab direct child<li>elements (not nested episode links). - Relative URL Handling: Some links on the page are relative, so we prepend the Reddit domain to make them full, usable URLs.
- Error Handling: Basic try/except blocks and checks for missing tags ensure the code doesn't crash if the page has unexpected formatting.
Notes for Maintenance
If the page structure changes in the future, you'll need to re-inspect the HTML (using your browser's dev tools) and adjust the tag selectors (e.g., if headers switch from <h3> to <h2>).
内容的提问来源于stack exchange,提问作者alpacafondue
相关产品推荐
相关产品推荐

