BeautifulSoup爬取加拿大药房页面:信息提取及价格表问题
加拿大药房药品详情页爬取问题解决方案
一、精准提取
标识的板块内容
放弃依赖固定标签位置的写法,改用标题标签的兄弟节点遍历法,基于每个
的位置,收集其后直到下一个
前的所有内容元素,适配不同页面的结构差异。
实现代码示例(基于BeautifulSoup)
from bs4 import BeautifulSoup def extract_sections(html_content): soup = BeautifulSoup(html_content, 'html.parser') sections = {} section_headers = soup.find_all('h4') for header in section_headers: section_name = header.get_text(strip=True) content_elements = [] next_sibling = header.next_sibling # 遍历当前标题后的所有兄弟节点,直到遇到下一个<h4> while next_sibling is not None: # 终止条件:遇到下一个板块标题 if next_sibling.name == 'h4': break # 只提取有实际内容的有效标签(可根据页面结构补充标签类型) if next_sibling.name in ['p', 'ul', 'ol'] and next_sibling.get_text(strip=True): # 处理列表类内容,转成易读的格式 if next_sibling.name in ['ul', 'ol']: list_items = [li.get_text(strip=True) for li in next_sibling.find_all('li')] content_elements.append('\n- ' + '\n- '.join(list_items)) else: content_elements.append(next_sibling.get_text(strip=True)) next_sibling = next_sibling.next_sibling # 将板块内容拼接后存入字典 sections[section_name] = '\n'.join(content_elements) return sections
核心逻辑说明
- 先定位所有
标签作为板块标识;
- 对每个标题,遍历其后的兄弟节点,过滤掉空白文本和无效标签;
- 区分普通文本(
)和列表(
- /
- 遇到下一个
时停止当前板块的内容收集,避免跨板块抓取。
- ),分别处理格式,确保内容可读性;
二、价格表结构化提取
针对价格表格式混乱问题,直接定位页面中的价格表标签(通常为),提取表头和每行数据,转成结构化的字典列表,避免原始HTML格式的混乱。
实现代码示例
def extract_price_table(html_content): soup = BeautifulSoup(html_content, 'html.parser') # 根据页面实际的价格表class/id调整定位器 price_table = soup.find('table', class_='price-table') price_data = [] if not price_table: return price_data # 提取表头 header_cells = price_table.find('thead').find_all('th') headers = [cell.get_text(strip=True) for cell in header_cells] # 提取每行数据 rows = price_table.find('tbody').find_all('tr') for row in rows: cell_texts = [cell.get_text(strip=True) for cell in row.find_all('td')] # 表头与行数据对应,生成结构化字典 price_row = dict(zip(headers, cell_texts)) price_data.append(price_row) return price_data补充:动态加载内容处理
如果页面的板块内容或价格表是通过JavaScript动态渲染的,静态HTTP请求无法获取完整内容,需使用浏览器自动化工具(如Selenium)获取渲染后的HTML:
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get('药品详情页URL') # 等待关键元素加载完成(以价格表为例) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'price-table')) ) # 获取渲染后的完整页面源码 html_content = driver.page_source driver.quit() # 调用上面的提取函数处理html_content sections = extract_sections(html_content) price_data = extract_price_table(html_content)内容的提问来源于stack exchange,提问作者Lalit Joshi
- 遇到下一个
相关产品推荐
相关产品推荐

