基于BeautifulSoup实现Drugbank药品专利信息分字段深度解析需求
Solution: Parse Drugbank Patent Details into Separate Fields
Hey there! I get what you're dealing with—those patent tables in Drugbank pages get mashed into a single block of text with your current code, which isn't ideal for extracting structured data like patent number, approval info, and country. Let's fix that by targeting the patent table specifically and pulling out each field separately.
Here's the updated code with focused patent parsing logic:
import requests from bs4 import BeautifulSoup def get_details(url): print('details:', url) r = requests.get(url) soup = BeautifulSoup(r.text, "lxml") # Iterate through all dt-dd pairs as before, but handle Patent section specially for dt, dd in zip(soup.findAll('dt'), soup.findAll('dd')): dt_text = dt.text.strip() print(f'* {dt_text}') # Check if we're in the Patent section if dt_text == "Patent": # Find the patent table inside the dd patent_table = dd.find('table') if patent_table: # Iterate through each row in the table body for row in patent_table.tbody.find_all('tr'): cells = row.find_all('td') if len(cells) >= 3: # Extract country from the flag image's alt attribute flag_img = cells[0].find('img') country = flag_img['alt'].strip() if flag_img else "Unknown" # Extract patent number patent_number = cells[1].text.strip() # Extract approval information approval_info = cells[2].text.strip() # Print structured patent data print(f" - Country: {country}") print(f" Patent Number: {patent_number}") print(f" Approval Info: {approval_info}") print(" ---") else: print(" No patent table found.") else: # For other sections, keep the original display logic print(dd.text.strip()) print('---------------------------') def drug_data(): url = 'https://www.drugbank.ca/drugs/' while url: print(url) r = requests.get(url) soup = BeautifulSoup(r.text, "lxml") links = soup.select('strong a') for link in links: get_details('https://www.drugbank.ca' + link['href']) # Get next page URL next_link = soup.find('a', {'class': 'page-link', 'rel': 'next'}) url = 'https://www.drugbank.ca' + next_link['href'] if next_link else None drug_data()
Key Changes Explained:
- Targeted Patent Section Detection: We check if the current
dttext is exactly "Patent" to trigger the special parsing logic. - Table Row Iteration: Inside the Patent
dd, we locate the embedded table and loop through each row in its tbody. - Field Extraction:
- Country: Pulled from the flag image's
altattribute (Drugbank consistently uses this to label country flags). - Patent Number: Directly extracted from the second table cell.
- Approval Info: Extracted from the third table cell, which includes details like approval dates or statuses.
- Country: Pulled from the flag image's
- Fallback Handling: Added checks for missing tables or cells to avoid errors if the page structure varies slightly.
This way, instead of getting a blob of merged text for patents, you'll get clean, structured output for each patent entry with separate fields for country, patent number, and approval information.
内容的提问来源于stack exchange,提问作者Lizou
相关产品推荐
相关产品推荐

