You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML表格数据爬取:选取标题间元素及MobyGames页面爬取问题

Hey there! Let's get that MobyGames credit scraping working properly, especially focusing on grabbing elements between those section headers. Here's a polished version of your code that does exactly that, with explanations to walk you through it:

Full Refined Scraping Code
import requests
from bs4 import BeautifulSoup

# Target URL for the game credits
credit_link = "http://www.mobygames.com/game/xbox360/wheelman/credits"

# Add a user-agent to avoid being blocked by anti-scraping measures
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
response = requests.get(credit_link, headers=headers)
response.raise_for_status()  # Catch HTTP errors immediately

# Parse the HTML content
soup = BeautifulSoup(response.text, "lxml")

# Locate the main credit container and target table
credit_container = soup.find("div", class_="col-md-8 col-lg-8")
credit_table = credit_container.select('table[summary="List of Credits"]')[0]
all_rows = credit_table.find_all('tr')

# Dictionary to store credits grouped by their section headers
credits_by_section = {}
current_section = None

# Iterate through rows to separate headers and associated personnel
for row in all_rows:
    # Identify section headers (they use a colspan to span the table width)
    header_cell = row.find('td', colspan=True)
    if header_cell:
        # Clean up header text and set as the active section
        current_section = header_cell.get_text(strip=True)
        credits_by_section[current_section] = []
        continue
    
    # Add personnel to the current section if we have an active header
    if current_section:
        cells = row.find_all('td')
        if len(cells) >= 2:
            name = cells[0].get_text(strip=True)
            role = cells[1].get_text(strip=True)
            credits_by_section[current_section].append({"name": name, "role": role})

# Example: Print organized results
for section, people in credits_by_section.items():
    print(f"* {section}:")
    for person in people:
        print(f"  - {person['name']}: {person['role']}")
Key Improvements & Explanations
  • Anti-Blocking Measure: Added a realistic user-agent to bypass MobyGames' default block of generic requests library headers.
  • Section Grouping: Automatically detects section headers (via the colspan attribute) and groups all subsequent personnel entries under that header—this is exactly how you capture elements between titles.
  • Error Handling: response.raise_for_status() ensures you catch issues like broken links or forbidden access early instead of silent failures.
  • Clean Data Extraction: Strips excess whitespace from names and roles to keep your data tidy.
Quick Best Practices
  • Always check a site's robots.txt before scraping to confirm you're allowed (MobyGames permits non-commercial personal use scraping, but avoid hitting their servers too frequently).
  • If the page structure changes later, adjust the header detection logic (e.g., check for a specific class instead of colspan).

内容的提问来源于stack exchange,提问作者edyvedy13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:18:35