You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python扩展代码提取维基百科信息表指定行的列表?

Reusable Solution to Extract Specific List Items from Wikipedia Infoboxes

Hey there! Sounds like you're looking to extend your existing Wikipedia infobox scraper to pull out list items from specific rows (like the "Products" row for Apple) while keeping the code flexible for other pages. Let's break this down with a Python example using BeautifulSoup—super common for web scraping tasks like this.

Step 1: The Core Reusable Function

First, let's build a generic function that takes a parsed Wikipedia page (as a BeautifulSoup object) and a target heading (like "Products") and returns the list items from that row. This is where the reusability comes in: you can pass any valid infobox heading and any Wikipedia page soup to get the corresponding list.

from bs4 import BeautifulSoup
import requests

def extract_infobox_list_items(soup, target_heading):
    # Find the infobox (covers common Wikipedia infobox class names)
    infobox = soup.find("table", class_=["infobox", "infobox_v2", "infobox_v3"])
    if not infobox:
        return []
    
    # Iterate through each row in the infobox
    for row in infobox.find_all("tr"):
        # Grab the header cell for the current row
        header_cell = row.find("th")
        if header_cell and header_cell.get_text(strip=True) == target_heading:
            # Get the data cell paired with the header
            data_cell = row.find("td")
            if not data_cell:
                return []
            
            # Extract and clean all list items from the data cell
            cleaned_items = []
            for li in data_cell.find_all("li"):
                raw_text = li.get_text(strip=True)
                # Optional: Remove citation markers like [1], [2]
                cleaned_text = ''.join([char for char in raw_text if not (char.isdigit() or char in '[]')])
                cleaned_items.append(cleaned_text)
            
            return cleaned_items
    
    # Return empty list if target heading isn't found
    return []

Step 2: Integrate with Your Existing Code

Here's how to plug this function into your existing scraper workflow. We'll use Apple's Wikipedia page as an example, but you can swap in any URL and target heading:

def scrape_wikipedia_data(wiki_url, target_heading):
    # Fetch and parse the Wikipedia page
    response = requests.get(wiki_url)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # Insert your existing infobox scraping code here (if you need full infobox data)
    # ...
    
    # Use our reusable function to get the target list
    target_list = extract_infobox_list_items(soup, target_heading)
    
    return target_list

# Test with Apple's Products
apple_products = scrape_wikipedia_data(
    "https://en.wikipedia.org/wiki/Apple_Inc.",
    "Products"
)

print("Apple's Products:")
for product in apple_products:
    print(f"- {product}")

Step 3: Reuse for Other Pages

Want to scrape Google's Products instead? Just swap the URL and keep the target heading—no need to rewrite core logic:

google_products = scrape_wikipedia_data(
    "https://en.wikipedia.org/wiki/Google",
    "Products"
)

print("\nGoogle's Products:")
for product in google_products:
    print(f"- {product}")

Key Reusability Tips

  • Infobox Compatibility: The function checks for multiple common infobox class names used across Wikipedia, so it works for most pages.
  • Robust Error Handling: Returns an empty list if the infobox or target heading is missing, preventing crashes in your main code.
  • Customizable Cleaning: The text cleanup step is optional—tweak it if you need to preserve citations or adjust formatting.

This setup keeps your code modular: you can use extract_infobox_list_items on its own, or integrate it seamlessly with your existing infobox scraping logic.

内容的提问来源于stack exchange,提问作者Ric

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:15:30