You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python抓取指定<div>内的<p>标签并过滤无关内容——以Adobe安全公告页面为例

Solution to Filter Unwanted Content

To exclude the "Platform:..." line, you can add a precise check to only keep <p> tags where the <strong> element contains one of your target labels. Here's how to adjust your code:

from bs4 import BeautifulSoup

# Assuming you've already parsed the page into the 'soup' object
div = soup.find("div", attrs={"id": "L0C1-body"})

# Define the exact labels you want to extract
target_labels = {
    "Release date:",
    "Last updated:",
    "Vulnerability identifier:",
    "CVE number:"
}

for p in div.findAll("p"):
    strong_element = p.find('strong')
    # Only process the tag if it has a strong element with a target label
    if strong_element and strong_element.text.strip() in target_labels:
        print(p.text.strip())

How This Works:

  1. We create a set of the exact labels we care about—this makes lookups fast and avoids accidental matches.
  2. For each <p> tag, we first check if it contains a <strong> child element.
  3. We then verify if the text inside that <strong> tag matches one of our target labels.
  4. Only when both conditions are met do we print the full text of the <p> tag.

This will skip the "Platform:..." entry entirely since its <strong> text isn’t in our target set, leaving you with just the four pieces of information you need.

内容的提问来源于stack exchange,提问作者tas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 12:02:38