You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫:从同一页面多链接提取指定数据

Solution to Extract Reaction Data from KEGG Compound Page

Got it, let's break this down step by step. Even without fancy class names in the HTML, we can leverage the page's consistent structural patterns—like <dt>/<dd> pairs for metadata fields and section headings to target specific blocks—to pull the exact data you need. Here's a practical Python implementation using requests and BeautifulSoup:

Step 1: Extract Core Metadata from the Main Compound Page

First, we'll grab the Name, Formula, reaction, and pathway fields from the main compound page (C10453). The page uses a definition list (<dl>) structure for top-level metadata, and a table under the "Reactions" heading for reaction entries.

import requests
from bs4 import BeautifulSoup
import time

# Base URL for KEGG
base_url = "http://www.genome.jp"
compound_url = f"{base_url}/dbget-bin/www_bget?cpd:C10453"

# Fetch and parse the main compound page
response = requests.get(compound_url)
response.encoding = "utf-8"  # Ensure proper character encoding to avoid garbled text
soup = BeautifulSoup(response.text, "html.parser")

# Helper function to pull metadata from <dt>/<dd> pairs
def get_metadata(soup_obj, field_name):
    dt_tag = soup_obj.find("dt", string=field_name)
    if dt_tag:
        return dt_tag.find_next_sibling("dd").text.strip()
    return "N/A"

# Extract top-level compound data
compound_name = get_metadata(soup, "Name")
compound_formula = get_metadata(soup, "Formula")

# Extract reactions, pathways, and their links
reactions_data = []
reactions_heading = soup.find("h3", string="Reactions")
if reactions_heading:
    reactions_table = reactions_heading.find_next("table")
    # Skip the header row, process each reaction entry
    for row in reactions_table.find_all("tr")[1:]:
        cols = row.find_all("td")
        if len(cols) >= 3:
            reaction_link = base_url + cols[0].find("a")["href"]
            reaction_text = cols[1].text.strip()
            pathway = cols[2].text.strip()
            
            reactions_data.append({
                "reaction_link": reaction_link,
                "reaction": reaction_text,
                "pathway": pathway
            })

# Print main page summary (optional)
print(f"Compound Name: {compound_name}")
print(f"Formula: {compound_formula}")
print(f"Found {len(reactions_data)} reactions to process\n")

Next, we'll loop through each reaction link we collected and pull the Name, definition, and reaction class fields. These reaction pages use the same <dt>/<dd> pattern for metadata, so we can reuse our helper logic.

def extract_reaction_details(reaction_url):
    """Pull Name, Definition, and Reaction Class from a single reaction page"""
    try:
        response = requests.get(reaction_url, timeout=10)
        response.raise_for_status()  # Raise error for HTTP issues like 404s
        response.encoding = "utf-8"
        soup = BeautifulSoup(response.text, "html.parser")
        
        return {
            "Name": get_metadata(soup, "Name"),
            "Definition": get_metadata(soup, "Definition"),
            "Reaction Class": get_metadata(soup, "Reaction class")
        }
    except requests.exceptions.RequestException as e:
        print(f"Failed to fetch {reaction_url}: {str(e)}")
        return {"Name": "N/A", "Definition": "N/A", "Reaction Class": "N/A"}

# Process each reaction and merge details
for idx, entry in enumerate(reactions_data, 1):
    print(f"Processing reaction {idx}...")
    details = extract_reaction_details(entry["reaction_link"])
    entry.update(details)
    
    # Print the extracted details (or save to a file/database)
    print(f"Reaction Name: {details['Name']}")
    print(f"Definition: {details['Definition']}")
    print(f"Reaction Class: {details['Reaction Class']}\n")
    
    # Add a small delay to avoid overwhelming KEGG's servers
    time.sleep(1)

Key Notes & Troubleshooting

  • Structural Reliance: This solution depends on KEGG's current HTML layout (e.g., <dt> tags matching exact field names, reactions stored in a table under an <h3> heading). If KEGG updates their site, you may need to adjust selectors (use browser dev tools to inspect new elements).
  • Error Handling: The code includes basic error handling for failed requests, but you can expand it to handle missing fields gracefully.
  • Rate Limiting: Adding time.sleep(1) between requests helps avoid getting blocked by KEGG's servers—always respect a site's rate limits when scraping.

内容的提问来源于stack exchange,提问作者Hemachandra Ghanta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:56:39