You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup实现Drugbank药品专利信息分字段深度解析需求

Solution: Parse Drugbank Patent Details into Separate Fields

Hey there! I get what you're dealing with—those patent tables in Drugbank pages get mashed into a single block of text with your current code, which isn't ideal for extracting structured data like patent number, approval info, and country. Let's fix that by targeting the patent table specifically and pulling out each field separately.

Here's the updated code with focused patent parsing logic:

import requests
from bs4 import BeautifulSoup

def get_details(url):
    print('details:', url)
    r = requests.get(url)
    soup = BeautifulSoup(r.text, "lxml")
    
    # Iterate through all dt-dd pairs as before, but handle Patent section specially
    for dt, dd in zip(soup.findAll('dt'), soup.findAll('dd')):
        dt_text = dt.text.strip()
        print(f'* {dt_text}')
        
        # Check if we're in the Patent section
        if dt_text == "Patent":
            # Find the patent table inside the dd
            patent_table = dd.find('table')
            if patent_table:
                # Iterate through each row in the table body
                for row in patent_table.tbody.find_all('tr'):
                    cells = row.find_all('td')
                    if len(cells) >= 3:
                        # Extract country from the flag image's alt attribute
                        flag_img = cells[0].find('img')
                        country = flag_img['alt'].strip() if flag_img else "Unknown"
                        
                        # Extract patent number
                        patent_number = cells[1].text.strip()
                        
                        # Extract approval information
                        approval_info = cells[2].text.strip()
                        
                        # Print structured patent data
                        print(f"  - Country: {country}")
                        print(f"    Patent Number: {patent_number}")
                        print(f"    Approval Info: {approval_info}")
                        print("    ---")
            else:
                print("  No patent table found.")
        else:
            # For other sections, keep the original display logic
            print(dd.text.strip())
        print('---------------------------')

def drug_data():
    url = 'https://www.drugbank.ca/drugs/'
    while url:
        print(url)
        r = requests.get(url)
        soup = BeautifulSoup(r.text, "lxml")
        links = soup.select('strong a')
        for link in links:
            get_details('https://www.drugbank.ca' + link['href'])
        # Get next page URL
        next_link = soup.find('a', {'class': 'page-link', 'rel': 'next'})
        url = 'https://www.drugbank.ca' + next_link['href'] if next_link else None

drug_data()

Key Changes Explained:

  • Targeted Patent Section Detection: We check if the current dt text is exactly "Patent" to trigger the special parsing logic.
  • Table Row Iteration: Inside the Patent dd, we locate the embedded table and loop through each row in its tbody.
  • Field Extraction:
    • Country: Pulled from the flag image's alt attribute (Drugbank consistently uses this to label country flags).
    • Patent Number: Directly extracted from the second table cell.
    • Approval Info: Extracted from the third table cell, which includes details like approval dates or statuses.
  • Fallback Handling: Added checks for missing tables or cells to avoid errors if the page structure varies slightly.

This way, instead of getting a blob of merged text for patents, you'll get clean, structured output for each patent entry with separate fields for country, patent number, and approval information.

内容的提问来源于stack exchange,提问作者Lizou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:27:10