Python网页爬虫:从同一页面多链接提取指定数据
Got it, let's break this down step by step. Even without fancy class names in the HTML, we can leverage the page's consistent structural patterns—like <dt>/<dd> pairs for metadata fields and section headings to target specific blocks—to pull the exact data you need. Here's a practical Python implementation using requests and BeautifulSoup:
Step 1: Extract Core Metadata from the Main Compound Page
First, we'll grab the Name, Formula, reaction, and pathway fields from the main compound page (C10453). The page uses a definition list (<dl>) structure for top-level metadata, and a table under the "Reactions" heading for reaction entries.
import requests from bs4 import BeautifulSoup import time # Base URL for KEGG base_url = "http://www.genome.jp" compound_url = f"{base_url}/dbget-bin/www_bget?cpd:C10453" # Fetch and parse the main compound page response = requests.get(compound_url) response.encoding = "utf-8" # Ensure proper character encoding to avoid garbled text soup = BeautifulSoup(response.text, "html.parser") # Helper function to pull metadata from <dt>/<dd> pairs def get_metadata(soup_obj, field_name): dt_tag = soup_obj.find("dt", string=field_name) if dt_tag: return dt_tag.find_next_sibling("dd").text.strip() return "N/A" # Extract top-level compound data compound_name = get_metadata(soup, "Name") compound_formula = get_metadata(soup, "Formula") # Extract reactions, pathways, and their links reactions_data = [] reactions_heading = soup.find("h3", string="Reactions") if reactions_heading: reactions_table = reactions_heading.find_next("table") # Skip the header row, process each reaction entry for row in reactions_table.find_all("tr")[1:]: cols = row.find_all("td") if len(cols) >= 3: reaction_link = base_url + cols[0].find("a")["href"] reaction_text = cols[1].text.strip() pathway = cols[2].text.strip() reactions_data.append({ "reaction_link": reaction_link, "reaction": reaction_text, "pathway": pathway }) # Print main page summary (optional) print(f"Compound Name: {compound_name}") print(f"Formula: {compound_formula}") print(f"Found {len(reactions_data)} reactions to process\n")
Step 2: Extract Details from Each Reaction Link
Next, we'll loop through each reaction link we collected and pull the Name, definition, and reaction class fields. These reaction pages use the same <dt>/<dd> pattern for metadata, so we can reuse our helper logic.
def extract_reaction_details(reaction_url): """Pull Name, Definition, and Reaction Class from a single reaction page""" try: response = requests.get(reaction_url, timeout=10) response.raise_for_status() # Raise error for HTTP issues like 404s response.encoding = "utf-8" soup = BeautifulSoup(response.text, "html.parser") return { "Name": get_metadata(soup, "Name"), "Definition": get_metadata(soup, "Definition"), "Reaction Class": get_metadata(soup, "Reaction class") } except requests.exceptions.RequestException as e: print(f"Failed to fetch {reaction_url}: {str(e)}") return {"Name": "N/A", "Definition": "N/A", "Reaction Class": "N/A"} # Process each reaction and merge details for idx, entry in enumerate(reactions_data, 1): print(f"Processing reaction {idx}...") details = extract_reaction_details(entry["reaction_link"]) entry.update(details) # Print the extracted details (or save to a file/database) print(f"Reaction Name: {details['Name']}") print(f"Definition: {details['Definition']}") print(f"Reaction Class: {details['Reaction Class']}\n") # Add a small delay to avoid overwhelming KEGG's servers time.sleep(1)
Key Notes & Troubleshooting
- Structural Reliance: This solution depends on KEGG's current HTML layout (e.g.,
<dt>tags matching exact field names, reactions stored in a table under an<h3>heading). If KEGG updates their site, you may need to adjust selectors (use browser dev tools to inspect new elements). - Error Handling: The code includes basic error handling for failed requests, but you can expand it to handle missing fields gracefully.
- Rate Limiting: Adding
time.sleep(1)between requests helps avoid getting blocked by KEGG's servers—always respect a site's rate limits when scraping.
内容的提问来源于stack exchange,提问作者Hemachandra Ghanta

