You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从指定网页HTML table生成指定JSON并去除UMLS标识?

Yes, this is absolutely achievable with Python!

Here's a practical, step-by-step implementation using common web scraping and data processing tools:

Required Libraries

First, install the necessary packages if you haven’t already:

pip install requests beautifulsoup4

Complete Code

import requests
from bs4 import BeautifulSoup
import json
import re

def scrape_disease_symptoms():
    # Target webpage URL
    url = "http://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/"
    
    # Fetch the page content
    response = requests.get(url)
    response.raise_for_status()  # Trigger error if request fails
    
    # Parse HTML content
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # Locate the disease-symptom table (assumed to be the only table on the page)
    table = soup.find('table')
    if not table:
        raise ValueError("Could not locate the target table on the webpage")
    
    # Initialize the result list
    data = []
    
    # Regex pattern to match UMLS identifiers (e.g., UMLS:C00080)
    umls_pattern = re.compile(r'\bUMLS:[A-Z0-9]+\b')
    
    # Iterate through table rows (skip the header row)
    rows = table.find_all('tr')[1:]
    for row in rows:
        cols = row.find_all('td')
        if len(cols) != 2:
            continue  # Skip malformed rows
        
        # Clean and extract disease name
        disease_name = cols[0].get_text(strip=True)
        
        # Process symptoms: split, remove UMLS tags, and clean
        symptoms_raw = cols[1].get_text(strip=True)
        symptoms_list = [s.strip() for s in symptoms_raw.split(',')]
        
        cleaned_symptoms = []
        for symptom in symptoms_list:
            # Remove UMLS identifier and trim extra whitespace
            cleaned = umls_pattern.sub('', symptom).strip()
            # Skip empty strings resulting from cleaning
            if cleaned:
                cleaned_symptoms.append(cleaned)
        
        # Add to the result dataset
        data.append({
            "disease_name": disease_name,
            "symptoms": cleaned_symptoms
        })
    
    return data

# Execute the function to get the formatted data
data = scrape_disease_symptoms()

# Optional: Save the output to a JSON file
with open('disease_symptoms.json', 'w', encoding='utf-8') as f:
    json.dump(data, f, indent=2, ensure_ascii=False)

# Print a sample of the output to verify
print(json.dumps(data[:2], indent=2))

Key Features

  • UMLS Identifier Removal: The regex r'\bUMLS:[A-Z0-9]+\b' targets and removes all UMLS tags, leaving only clean symptom text.
  • Robust Data Cleaning: We handle whitespace, split symptoms correctly, and filter out empty entries that might come from tag removal.
  • Error Handling: Basic checks for failed requests and missing tables ensure the code doesn’t crash unexpectedly.
  • Exact Output Format: The final data variable matches your requested structure exactly: a list of dictionaries with disease_name (string) and symptoms (list of strings).

Notes

  • If the webpage’s HTML structure changes (e.g., table classes, row formatting), you may need to adjust the selectors in the code.
  • Always respect the website’s robots.txt and terms of service when scraping. This academic site appears publicly accessible, but it’s good practice to confirm before scraping.

内容的提问来源于stack exchange,提问作者Khan Asfi Reza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 17:33:13