如何从HTML表格抓取含特定属性的行?附MobyGames爬取代码
Got it, let's walk through how to filter rows with specific attributes and pull the data you need from that MobyGames review table. Here's a practical breakdown based on your existing code:
First, let's recap your current setup
You've already grabbed the table rows with valign="top" — great start! Here's your code with necessary imports added for context:
import requests from bs4 import BeautifulSoup # Don't forget to define your headers (example below) headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} url_credit = "http://www.mobygames.com/game/wheelman/view-moby-score" response = requests.get(url_credit, headers=headers) soup = BeautifulSoup(response.text, "lxml") table_rows = soup.find("table", class_="reviewList table table-striped table-condensed table-hover").select('tr[valign="top"]')
Now, let's filter rows by specific attributes
The approach depends on what "specific attribute" you're targeting. Below are common scenarios with code examples:
Scenario 1: Filter rows with a specific HTML attribute
If your target rows have a unique attribute (like data-review-id), you can check for its existence directly:
# Filter all rows that have a data-review-id attribute filtered_rows = [row for row in table_rows[1:] if row.has_attr("data-review-id")]
Scenario 2: Filter rows with specific content (e.g., reviews from a particular outlet)
Suppose you want only reviews from IGN. You can check the text in the media name cell:
for row in table_rows[1:]: # Grab the first TD (which contains the media outlet name) media_outlet = row.find("td").get_text(strip=True) if media_outlet == "IGN": # Extract relevant data from the row score = row.find("td", class_="reviewScore").get_text(strip=True) review_date = row.find("td", class_="reviewDate").get_text(strip=True) review_excerpt = row.find("td", class_="reviewExcerpt").get_text(strip=True) print(f"IGN Review: {score}/10 | {review_date}\nExcerpt: {review_excerpt}\n")
Scenario 3: Filter rows by numerical score (e.g., scores above 8)
If you want only high-scoring reviews, parse the score value and filter accordingly:
for row in table_rows[1:]: score_cell = row.find("td", class_="reviewScore") if score_cell: try: # Convert score text to a float score = float(score_cell.get_text(strip=True)) if score > 8: media = row.find("td").get_text(strip=True) print(f"High-scoring review: {media} - {score}/10") except ValueError: # Skip rows where the score isn't a number (like "N/A") continue
Key Tips
- Use your browser's developer tools (F12) to inspect the table rows and confirm:
- The exact attributes or class names of your target rows/cells
- The structure of the data you want to extract
- Always handle edge cases (like non-numeric scores or missing cells) to avoid errors in your scraper.
内容的提问来源于stack exchange,提问作者edyvedy13

