如何用Python扩展代码提取维基百科信息表指定行的列表?
Hey there! Sounds like you're looking to extend your existing Wikipedia infobox scraper to pull out list items from specific rows (like the "Products" row for Apple) while keeping the code flexible for other pages. Let's break this down with a Python example using BeautifulSoup—super common for web scraping tasks like this.
Step 1: The Core Reusable Function
First, let's build a generic function that takes a parsed Wikipedia page (as a BeautifulSoup object) and a target heading (like "Products") and returns the list items from that row. This is where the reusability comes in: you can pass any valid infobox heading and any Wikipedia page soup to get the corresponding list.
from bs4 import BeautifulSoup import requests def extract_infobox_list_items(soup, target_heading): # Find the infobox (covers common Wikipedia infobox class names) infobox = soup.find("table", class_=["infobox", "infobox_v2", "infobox_v3"]) if not infobox: return [] # Iterate through each row in the infobox for row in infobox.find_all("tr"): # Grab the header cell for the current row header_cell = row.find("th") if header_cell and header_cell.get_text(strip=True) == target_heading: # Get the data cell paired with the header data_cell = row.find("td") if not data_cell: return [] # Extract and clean all list items from the data cell cleaned_items = [] for li in data_cell.find_all("li"): raw_text = li.get_text(strip=True) # Optional: Remove citation markers like [1], [2] cleaned_text = ''.join([char for char in raw_text if not (char.isdigit() or char in '[]')]) cleaned_items.append(cleaned_text) return cleaned_items # Return empty list if target heading isn't found return []
Step 2: Integrate with Your Existing Code
Here's how to plug this function into your existing scraper workflow. We'll use Apple's Wikipedia page as an example, but you can swap in any URL and target heading:
def scrape_wikipedia_data(wiki_url, target_heading): # Fetch and parse the Wikipedia page response = requests.get(wiki_url) soup = BeautifulSoup(response.text, "html.parser") # Insert your existing infobox scraping code here (if you need full infobox data) # ... # Use our reusable function to get the target list target_list = extract_infobox_list_items(soup, target_heading) return target_list # Test with Apple's Products apple_products = scrape_wikipedia_data( "https://en.wikipedia.org/wiki/Apple_Inc.", "Products" ) print("Apple's Products:") for product in apple_products: print(f"- {product}")
Step 3: Reuse for Other Pages
Want to scrape Google's Products instead? Just swap the URL and keep the target heading—no need to rewrite core logic:
google_products = scrape_wikipedia_data( "https://en.wikipedia.org/wiki/Google", "Products" ) print("\nGoogle's Products:") for product in google_products: print(f"- {product}")
Key Reusability Tips
- Infobox Compatibility: The function checks for multiple common infobox class names used across Wikipedia, so it works for most pages.
- Robust Error Handling: Returns an empty list if the infobox or target heading is missing, preventing crashes in your main code.
- Customizable Cleaning: The text cleanup step is optional—tweak it if you need to preserve citations or adjust formatting.
This setup keeps your code modular: you can use extract_infobox_list_items on its own, or integrate it seamlessly with your existing infobox scraping logic.
内容的提问来源于stack exchange,提问作者Ric

