HTML表格数据爬取:选取标题间元素及MobyGames页面爬取问题
Hey there! Let's get that MobyGames credit scraping working properly, especially focusing on grabbing elements between those section headers. Here's a polished version of your code that does exactly that, with explanations to walk you through it:
Full Refined Scraping Code
import requests from bs4 import BeautifulSoup # Target URL for the game credits credit_link = "http://www.mobygames.com/game/xbox360/wheelman/credits" # Add a user-agent to avoid being blocked by anti-scraping measures headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(credit_link, headers=headers) response.raise_for_status() # Catch HTTP errors immediately # Parse the HTML content soup = BeautifulSoup(response.text, "lxml") # Locate the main credit container and target table credit_container = soup.find("div", class_="col-md-8 col-lg-8") credit_table = credit_container.select('table[summary="List of Credits"]')[0] all_rows = credit_table.find_all('tr') # Dictionary to store credits grouped by their section headers credits_by_section = {} current_section = None # Iterate through rows to separate headers and associated personnel for row in all_rows: # Identify section headers (they use a colspan to span the table width) header_cell = row.find('td', colspan=True) if header_cell: # Clean up header text and set as the active section current_section = header_cell.get_text(strip=True) credits_by_section[current_section] = [] continue # Add personnel to the current section if we have an active header if current_section: cells = row.find_all('td') if len(cells) >= 2: name = cells[0].get_text(strip=True) role = cells[1].get_text(strip=True) credits_by_section[current_section].append({"name": name, "role": role}) # Example: Print organized results for section, people in credits_by_section.items(): print(f"* {section}:") for person in people: print(f" - {person['name']}: {person['role']}")
Key Improvements & Explanations
- Anti-Blocking Measure: Added a realistic user-agent to bypass MobyGames' default block of generic
requestslibrary headers. - Section Grouping: Automatically detects section headers (via the
colspanattribute) and groups all subsequent personnel entries under that header—this is exactly how you capture elements between titles. - Error Handling:
response.raise_for_status()ensures you catch issues like broken links or forbidden access early instead of silent failures. - Clean Data Extraction: Strips excess whitespace from names and roles to keep your data tidy.
Quick Best Practices
- Always check a site's
robots.txtbefore scraping to confirm you're allowed (MobyGames permits non-commercial personal use scraping, but avoid hitting their servers too frequently). - If the page structure changes later, adjust the header detection logic (e.g., check for a specific class instead of
colspan).
内容的提问来源于stack exchange,提问作者edyvedy13
相关产品推荐
相关产品推荐

