如何用Python从指定网页HTML table生成指定JSON并去除UMLS标识?
Yes, this is absolutely achievable with Python!
Here's a practical, step-by-step implementation using common web scraping and data processing tools:
Required Libraries
First, install the necessary packages if you haven’t already:
pip install requests beautifulsoup4
Complete Code
import requests from bs4 import BeautifulSoup import json import re def scrape_disease_symptoms(): # Target webpage URL url = "http://people.dbmi.columbia.edu/~friedma/Projects/DiseaseSymptomKB/" # Fetch the page content response = requests.get(url) response.raise_for_status() # Trigger error if request fails # Parse HTML content soup = BeautifulSoup(response.text, 'html.parser') # Locate the disease-symptom table (assumed to be the only table on the page) table = soup.find('table') if not table: raise ValueError("Could not locate the target table on the webpage") # Initialize the result list data = [] # Regex pattern to match UMLS identifiers (e.g., UMLS:C00080) umls_pattern = re.compile(r'\bUMLS:[A-Z0-9]+\b') # Iterate through table rows (skip the header row) rows = table.find_all('tr')[1:] for row in rows: cols = row.find_all('td') if len(cols) != 2: continue # Skip malformed rows # Clean and extract disease name disease_name = cols[0].get_text(strip=True) # Process symptoms: split, remove UMLS tags, and clean symptoms_raw = cols[1].get_text(strip=True) symptoms_list = [s.strip() for s in symptoms_raw.split(',')] cleaned_symptoms = [] for symptom in symptoms_list: # Remove UMLS identifier and trim extra whitespace cleaned = umls_pattern.sub('', symptom).strip() # Skip empty strings resulting from cleaning if cleaned: cleaned_symptoms.append(cleaned) # Add to the result dataset data.append({ "disease_name": disease_name, "symptoms": cleaned_symptoms }) return data # Execute the function to get the formatted data data = scrape_disease_symptoms() # Optional: Save the output to a JSON file with open('disease_symptoms.json', 'w', encoding='utf-8') as f: json.dump(data, f, indent=2, ensure_ascii=False) # Print a sample of the output to verify print(json.dumps(data[:2], indent=2))
Key Features
- UMLS Identifier Removal: The regex
r'\bUMLS:[A-Z0-9]+\b'targets and removes all UMLS tags, leaving only clean symptom text. - Robust Data Cleaning: We handle whitespace, split symptoms correctly, and filter out empty entries that might come from tag removal.
- Error Handling: Basic checks for failed requests and missing tables ensure the code doesn’t crash unexpectedly.
- Exact Output Format: The final
datavariable matches your requested structure exactly: a list of dictionaries withdisease_name(string) andsymptoms(list of strings).
Notes
- If the webpage’s HTML structure changes (e.g., table classes, row formatting), you may need to adjust the selectors in the code.
- Always respect the website’s
robots.txtand terms of service when scraping. This academic site appears publicly accessible, but it’s good practice to confirm before scraping.
内容的提问来源于stack exchange,提问作者Khan Asfi Reza
相关产品推荐
相关产品推荐

