按顺序从HTML提取数据:如何从维基百科页面提取表格及前置数据
Extracting Order/Family Data and Associated Species Tables from the Wikipedia Birds Page
Here's a practical, step-by-step solution to pull the repeated Order/Family headers and their corresponding species tables from the Trinidad and Tobago birds Wikipedia page. We'll use Python with requests and BeautifulSoup—tools commonly used for web scraping tasks like this.
Prerequisites
First, install the required libraries if you haven't already:
pip install requests beautifulsoup4
Full Implementation Code
import requests from bs4 import BeautifulSoup # Set a user-agent to avoid being blocked by Wikipedia headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } # Fetch the page content url = "https://en.wikipedia.org/wiki/List_of_birds_of_Trinidad_and_Tobago" response = requests.get(url, headers=headers) response.raise_for_status() # Raise an error if the request fails soup = BeautifulSoup(response.text, 'html.parser') # Find all <p> elements that contain both Order and Family headers order_family_sections = soup.find_all( 'p', lambda tag: tag.find('b', string=lambda s: s and 'Order' in s.strip()) and tag.find('b', string=lambda s: s and 'Family' in s.strip()) ) # Process each section for section in order_family_sections: # Extract Order name order_bold = section.find('b', string=lambda s: s and 'Order' in s.strip()) order_name = order_bold.find_next('a').text.strip() # Extract Family name family_bold = section.find('b', string=lambda s: s and 'Family' in s.strip()) family_name = family_bold.find_next('a').text.strip() # Get the associated species table species_table = section.find_next('table', class_='wikitable') if not species_table: print(f"No table found for Order: {order_name}, Family: {family_name}\n") continue # Extract table rows and cells table_data = [] for row in species_table.find_all('tr'): cells = row.find_all(['th', 'td']) row_data = [cell.text.strip() for cell in cells] table_data.append(row_data) # Print or save the results print(f"=== Order: {order_name} | Family: {family_name} ===") for row in table_data: print(row) print("\n" + "-"*50 + "\n")
How It Works
- Fetching the Page: We use
requeststo get the HTML content, with a proper user-agent to comply with Wikipedia's terms of service. - Locating Sections: We filter
<p>elements that contain both "Order" and "Family" bolded labels—these are our section headers. - Extracting Metadata: For each section, we find the linked text next to the "Order" and "Family" labels to get their names.
- Grabbing the Table: We find the immediately following
wikitable(the standard class for Wikipedia data tables) associated with each section. - Parsing Table Data: We convert each table row into a list of cleaned cell text, making it easy to process or save as CSV/JSON.
Notes
- Edge Cases: The code includes a check for missing tables (though rare on this page) to avoid errors.
- Data Cleaning: Depending on your needs, you might want to further clean the table data (e.g., remove superscripts, handle merged cells).
- Alternative: If you prefer using
pandasfor table handling, you can replace the table parsing part withpd.read_html(str(species_table))[0]—just make sure to map each DataFrame to its corresponding Order/Family pair.
内容的提问来源于stack exchange,提问作者jasmaar
相关产品推荐
相关产品推荐

