You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3中用BeautifulSoup解析HTML时如何排除特定class的<p>标签?

Hey there! The issue here is that your current code grabs all <p> tags, including the one with the notice class you want to exclude. Let's fix this with a couple of clean approaches using BeautifulSoup:

Method 1: Use a lambda function with find_all

You can pass a lambda to the class_ parameter to filter out any <p> tags that have the notice class. This works even if the tag has multiple classes (like class="notice some-other-class"):

def getBody(url):
    html_page = requests.get(url)
    soup = BeautifulSoup(html_page.content, 'html.parser')
    # Exclude <p> tags with "notice" in their class
    Con = "".join([
        p.text for p in soup.find_all(
            "p", 
            class_=lambda x: x is None or "notice" not in x.split()
        )
    ])
    return Con

Method 2: Use CSS selectors (more concise)

BeautifulSoup supports CSS selectors, and the :not() pseudo-class makes this super straightforward. This will directly target all <p> tags that don't have the notice class:

def getBody(url):
    html_page = requests.get(url)
    soup = BeautifulSoup(html_page.content, 'html.parser')
    # Use CSS selector to exclude .notice class on <p> tags
    Con = "".join([p.text for p in soup.select("p:not(.notice)")])
    return Con

Both methods will skip that unwanted "Comments are closed for this article." text. The CSS selector approach is usually more readable if you're familiar with CSS syntax!

内容的提问来源于stack exchange,提问作者Parth Parikh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:10:38