You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup爬取加拿大药房页面:信息提取及价格表问题

加拿大药房药品详情页爬取问题解决方案

一、精准提取

标识的板块内容

放弃依赖固定标签位置的写法,改用标题标签的兄弟节点遍历法,基于每个

的位置,收集其后直到下一个

前的所有内容元素,适配不同页面的结构差异。

实现代码示例(基于BeautifulSoup)

from bs4 import BeautifulSoup

def extract_sections(html_content):
    soup = BeautifulSoup(html_content, 'html.parser')
    sections = {}
    section_headers = soup.find_all('h4')

    for header in section_headers:
        section_name = header.get_text(strip=True)
        content_elements = []
        next_sibling = header.next_sibling

        # 遍历当前标题后的所有兄弟节点,直到遇到下一个<h4>
        while next_sibling is not None:
            # 终止条件:遇到下一个板块标题
            if next_sibling.name == 'h4':
                break
            # 只提取有实际内容的有效标签(可根据页面结构补充标签类型)
            if next_sibling.name in ['p', 'ul', 'ol'] and next_sibling.get_text(strip=True):
                # 处理列表类内容,转成易读的格式
                if next_sibling.name in ['ul', 'ol']:
                    list_items = [li.get_text(strip=True) for li in next_sibling.find_all('li')]
                    content_elements.append('\n- ' + '\n- '.join(list_items))
                else:
                    content_elements.append(next_sibling.get_text(strip=True))
            next_sibling = next_sibling.next_sibling
        
        # 将板块内容拼接后存入字典
        sections[section_name] = '\n'.join(content_elements)
    return sections

核心逻辑说明

  • 先定位所有

    标签作为板块标识;

  • 对每个标题,遍历其后的兄弟节点,过滤掉空白文本和无效标签;
  • 区分普通文本(

    )和列表(

      /
        ),分别处理格式,确保内容可读性;
      1. 遇到下一个

        时停止当前板块的内容收集,避免跨板块抓取。

    二、价格表结构化提取

    针对价格表格式混乱问题,直接定位页面中的价格表标签(通常为),提取表头和每行数据,转成结构化的字典列表,避免原始HTML格式的混乱。

    实现代码示例

    def extract_price_table(html_content):
        soup = BeautifulSoup(html_content, 'html.parser')
        # 根据页面实际的价格表class/id调整定位器
        price_table = soup.find('table', class_='price-table')
        price_data = []
    
        if not price_table:
            return price_data
        
        # 提取表头
        header_cells = price_table.find('thead').find_all('th')
        headers = [cell.get_text(strip=True) for cell in header_cells]
        
        # 提取每行数据
        rows = price_table.find('tbody').find_all('tr')
        for row in rows:
            cell_texts = [cell.get_text(strip=True) for cell in row.find_all('td')]
            # 表头与行数据对应,生成结构化字典
            price_row = dict(zip(headers, cell_texts))
            price_data.append(price_row)
        
        return price_data
    

    补充:动态加载内容处理

    如果页面的板块内容或价格表是通过JavaScript动态渲染的,静态HTTP请求无法获取完整内容,需使用浏览器自动化工具(如Selenium)获取渲染后的HTML:

    from selenium import webdriver
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    driver = webdriver.Chrome()
    driver.get('药品详情页URL')
    
    # 等待关键元素加载完成(以价格表为例)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, 'price-table'))
    )
    # 获取渲染后的完整页面源码
    html_content = driver.page_source
    driver.quit()
    
    # 调用上面的提取函数处理html_content
    sections = extract_sections(html_content)
    price_data = extract_price_table(html_content)
    

    内容的提问来源于stack exchange,提问作者Lalit Joshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 02:20:47