You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从医疗网站HTML中按标题提取药物分段信息?

按标题归类提取药物信息的解决方案

Got it, let's break down how to solve this—this is a classic structured scraping problem, and the key is to tie each section heading to its corresponding content block instead of grabbing everything at once. Here are two reliable approaches based on the HTML structure you described:

方法1:基于锚点的兄弟元素遍历

Since you have <a name="Warning"></a> style anchors marking section headers, you can target each anchor, then collect all content that comes after it until you hit the next section anchor. This works great if the content flows directly after the anchor without a wrapping container.

Here's a concrete example using Python's BeautifulSoup:

from bs4 import BeautifulSoup

# 假设你的HTML内容存储在html变量中
soup = BeautifulSoup(html, 'html.parser')

# 先定义网站锚点名称和你要归类的字段的映射(根据实际网站调整)
section_mapping = {
    "Warning": "Precautions",
    "Contraindications": "Contraindications",
    "Dosage": "Dosage",
    "Uses": "Uses"
}

drug_details = {}

for anchor_name, target_field in section_mapping.items():
    # 找到对应的锚点标签
    anchor = soup.find('a', attrs={"name": anchor_name})
    if not anchor:
        drug_details[target_field] = None
        continue
    
    # 收集锚点之后的所有内容,直到遇到下一个目标锚点
    content_pieces = []
    current_sibling = anchor.next_sibling
    
    while current_sibling:
        # 停止条件:遇到下一个我们关心的section锚点
        if current_sibling.name == "a" and current_sibling.get("name") in section_mapping:
            break
        
        # 提取内容块(匹配你提到的p.drug-content或div.report-content等标签)
        if current_sibling.name in ["p", "div"]:
            # 检查是否是目标内容类
            if "drug-content" in current_sibling.get("class", []) or "report-content" in current_sibling.get("class", []):
                # 清理文本,去掉多余换行和空格
                clean_text = current_sibling.get_text(strip=True, separator=" ")
                if clean_text:
                    content_pieces.append(clean_text)
        
        current_sibling = current_sibling.next_sibling
    
    # 合并内容并存入结果
    drug_details[target_field] = " ".join(content_pieces) if content_pieces else None

# 输出结果
print(drug_details)

方法2:基于容器的区块提取

If each section (heading + content) is wrapped in a parent container (like a <div class="section">), this is even simpler—just iterate over each container and extract the heading and content separately.

For example, if the HTML looks like:

<div class="section">
    <a name="Uses"></a>
    <div class="report-content">...</div>
</div>

You can adjust the code to target these containers directly:

sections = soup.find_all("div", class_="section")
for section in sections:
    anchor = section.find("a")
    if not anchor or anchor.get("name") not in section_mapping:
        continue
    target_field = section_mapping[anchor.get("name")]
    content = section.find("div", class_="report-content") or section.find("p", class_="drug-content")
    drug_details[target_field] = content.get_text(strip=True, separator=" ") if content else None

关键注意事项

  • Adjust selectors: Tweak the tag names, class names, and anchor names to match the exact structure of the medical site you're scraping—sites often have minor variations in class names.
  • Handle edge cases: Some sections might have content split across multiple <p> or <div> tags; the sibling traversal method will collect all of them until the next heading.
  • Clean text: Always strip extra whitespace and newlines to keep the content readable (the get_text(strip=True, separator=" ") method helps with this).

Hope this gets you sorted! If you run into weird edge cases (like nested sections or dynamic content), just adjust the traversal logic to match the site's structure.

内容的提问来源于stack exchange,提问作者serendipity

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:47:41