寻求Python+Beautiful Soup构建药物数据Pandas DataFrame的方案
Hey there! I’ve dealt with similar issues scraping structured data into Pandas DataFrames before, so let’s work through this together.
The core problem here is that splitting on --- often leaves messy, inconsistent chunks of text—some entries might be missing fields, have extra whitespace, or have values that contain colons/line breaks. This makes it hard for Pandas to create a clean, aligned DataFrame. Let’s fix this with two approaches, depending on how you’re extracting the raw text.
Problem Breakdown
When you split using split("---"), you probably end up with a list where each item looks something like this (messy, inconsistent):
Name: Aspirin
Type: Small MoleculeName: Ibuprofen
Type: Small Molecule
Target: COX-2
(extra blank lines here)
Blank lines, missing fields, and unstructured text throw off Pandas’ ability to map entries to columns. We need to standardize each drug’s data into a consistent dictionary first.
Solution 1: Clean Up Split Text & Build a Structured Dictionary List
If you’re stuck using the --- split method, here’s how to turn those messy chunks into a clean DataFrame:
import pandas as pd # Example raw text (replace with your scraped content) raw_drug_content = """ Name: Aspirin Type: Small Molecule --- Name: Ibuprofen Type: Small Molecule Target: COX-2 --- Name: Paracetamol Type: Small Molecule Indication: Pain Relief """ # Step 1: Split and clean entries drug_entries = [entry.strip() for entry in raw_drug_content.split("---") if entry.strip()] # Step 2: Define all possible fields you want to capture (expand this based on Drugbank's data) required_fields = ["Name", "Type", "Target", "Indication", "Mechanism of Action"] # Step 3: Convert each entry to a standardized dictionary drug_records = [] for entry in drug_entries: # Initialize dict with all fields set to None (handles missing data) drug_dict = {field: None for field in required_fields} # Split entry into individual lines, skip empty ones lines = [line.strip() for line in entry.split("\n") if line.strip()] for line in lines: # Split only on the FIRST colon (avoids breaking values with colons) if ": " in line: key, value = line.split(": ", 1) # Only add the key if it's in our predefined fields if key in required_fields: drug_dict[key] = value drug_records.append(drug_dict) # Step 4: Convert to DataFrame clean_df = pd.DataFrame(drug_records) print(clean_df)
Key Fixes Here:
- Whitespace Cleaning:
strip()removes extra newlines and spaces from each entry and line. - Controlled Splitting:
split(": ", 1)ensures we don’t split values that contain colons (e.g., a mechanism of action like "Inhibits: COX-1"). - Standardized Fields: Predefining fields guarantees every entry has the same columns, even if some values are missing (filled with
NaNin the DataFrame).
Solution 2: Use BeautifulSoup Directly (More Reliable)
If you’re scraping HTML pages (which Drugbank uses), it’s way better to extract data directly from HTML elements instead of relying on text splits. This avoids issues with inconsistent --- placement or messy text formatting.
Here’s a sample workflow:
import pandas as pd from bs4 import BeautifulSoup import requests # Fetch the subpage (replace with your actual URL) url = "https://www.drugbank.ca/drugs/DB00945" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Example: Extract data from specific HTML elements (adjust selectors to match Drugbank's structure) drug_records = [] # Assume each drug entry is in a div with class "drug-card" (adjust based on actual page) drug_cards = soup.find_all("div", class_="drug-card") for card in drug_cards: drug_dict = {} # Extract name (use the correct tag/class for Drugbank) name_elem = card.find("h3", class_="drug-name") drug_dict["Name"] = name_elem.get_text(strip=True) if name_elem else None # Extract type type_elem = card.find("span", class_="drug-type") drug_dict["Type"] = type_elem.get_text(strip=True) if type_elem else None # Extract target (handle cases where it might not exist) target_elem = card.find("div", class_="target-info") drug_dict["Target"] = target_elem.get_text(strip=True) if target_elem else None # Add more fields as needed (indication, mechanism, etc.) drug_records.append(drug_dict) # Convert to DataFrame clean_df = pd.DataFrame(drug_records)
Why This Is Better:
- Structured Extraction: You’re pulling data directly from HTML elements (like
<h3>or<span>) that are designed to hold specific information, so you avoid text parsing errors. - Robustness: Even if Drugbank changes minor text formatting, as long as the HTML structure stays similar, your code will keep working.
Final Notes
- Expand Fields: Make sure your
required_fieldslist includes every piece of data you want to capture from Drugbank (e.g., "CAS Number", "Molecular Weight", "Indication"). - Handle Missing Data: Using
Nonefor missing fields ensures your DataFrame has consistent columns, which makes it easier to analyze later. - Rate Limiting: Remember to respect Drugbank’s terms of service—add delays between requests to avoid getting blocked.
内容的提问来源于stack exchange,提问作者Lizou

