You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按顺序从HTML提取数据:如何从维基百科页面提取表格及前置数据

Extracting Order/Family Data and Associated Species Tables from the Wikipedia Birds Page

Here's a practical, step-by-step solution to pull the repeated Order/Family headers and their corresponding species tables from the Trinidad and Tobago birds Wikipedia page. We'll use Python with requests and BeautifulSoup—tools commonly used for web scraping tasks like this.

Prerequisites

First, install the required libraries if you haven't already:

pip install requests beautifulsoup4

Full Implementation Code

import requests
from bs4 import BeautifulSoup

# Set a user-agent to avoid being blocked by Wikipedia
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}

# Fetch the page content
url = "https://en.wikipedia.org/wiki/List_of_birds_of_Trinidad_and_Tobago"
response = requests.get(url, headers=headers)
response.raise_for_status()  # Raise an error if the request fails
soup = BeautifulSoup(response.text, 'html.parser')

# Find all <p> elements that contain both Order and Family headers
order_family_sections = soup.find_all(
    'p',
    lambda tag: tag.find('b', string=lambda s: s and 'Order' in s.strip()) 
    and tag.find('b', string=lambda s: s and 'Family' in s.strip())
)

# Process each section
for section in order_family_sections:
    # Extract Order name
    order_bold = section.find('b', string=lambda s: s and 'Order' in s.strip())
    order_name = order_bold.find_next('a').text.strip()
    
    # Extract Family name
    family_bold = section.find('b', string=lambda s: s and 'Family' in s.strip())
    family_name = family_bold.find_next('a').text.strip()
    
    # Get the associated species table
    species_table = section.find_next('table', class_='wikitable')
    
    if not species_table:
        print(f"No table found for Order: {order_name}, Family: {family_name}\n")
        continue
    
    # Extract table rows and cells
    table_data = []
    for row in species_table.find_all('tr'):
        cells = row.find_all(['th', 'td'])
        row_data = [cell.text.strip() for cell in cells]
        table_data.append(row_data)
    
    # Print or save the results
    print(f"=== Order: {order_name} | Family: {family_name} ===")
    for row in table_data:
        print(row)
    print("\n" + "-"*50 + "\n")

How It Works

  1. Fetching the Page: We use requests to get the HTML content, with a proper user-agent to comply with Wikipedia's terms of service.
  2. Locating Sections: We filter <p> elements that contain both "Order" and "Family" bolded labels—these are our section headers.
  3. Extracting Metadata: For each section, we find the linked text next to the "Order" and "Family" labels to get their names.
  4. Grabbing the Table: We find the immediately following wikitable (the standard class for Wikipedia data tables) associated with each section.
  5. Parsing Table Data: We convert each table row into a list of cleaned cell text, making it easy to process or save as CSV/JSON.

Notes

  • Edge Cases: The code includes a check for missing tables (though rare on this page) to avoid errors.
  • Data Cleaning: Depending on your needs, you might want to further clean the table data (e.g., remove superscripts, handle merged cells).
  • Alternative: If you prefer using pandas for table handling, you can replace the table parsing part with pd.read_html(str(species_table))[0]—just make sure to map each DataFrame to its corresponding Order/Family pair.

内容的提问来源于stack exchange,提问作者jasmaar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:19:17