如何使用Python抓取指定<div>内的<p>标签并过滤无关内容——以Adobe安全公告页面为例
Solution to Filter Unwanted Content
To exclude the "Platform:..." line, you can add a precise check to only keep <p> tags where the <strong> element contains one of your target labels. Here's how to adjust your code:
from bs4 import BeautifulSoup # Assuming you've already parsed the page into the 'soup' object div = soup.find("div", attrs={"id": "L0C1-body"}) # Define the exact labels you want to extract target_labels = { "Release date:", "Last updated:", "Vulnerability identifier:", "CVE number:" } for p in div.findAll("p"): strong_element = p.find('strong') # Only process the tag if it has a strong element with a target label if strong_element and strong_element.text.strip() in target_labels: print(p.text.strip())
How This Works:
- We create a set of the exact labels we care about—this makes lookups fast and avoids accidental matches.
- For each
<p>tag, we first check if it contains a<strong>child element. - We then verify if the text inside that
<strong>tag matches one of our target labels. - Only when both conditions are met do we print the full text of the
<p>tag.
This will skip the "Platform:..." entry entirely since its <strong> text isn’t in our target set, leaving you with just the four pieces of information you need.
内容的提问来源于stack exchange,提问作者tas
相关产品推荐
相关产品推荐

