You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取时排除<p>内<div class='ads'>内容的方法

如何用BeautifulSoup提取

标签内容并排除嵌套的广告

我懂你现在的困扰——用getText()提取<p>内容时,连嵌套在里面的广告div文本也一起抓进来了,确实头疼。问题出在getText()会提取当前标签下所有子节点的文本内容,包括那个<div class="ads">里的广告内容。下面给你两种可行的解决方案:

方法一:先移除广告节点再提取文本(推荐)

这种方法直接把广告节点从DOM树中删除,之后再提取<p>的文本就不会包含广告内容了,逻辑清晰且易维护:

import bs4

soup = bs4.BeautifulSoup(page.content, 'html.parser')
news = soup.find('div',{'class': 'col-md-8 left-container details'})
News_article = news.find_all('div',{'class': 'news-article'})
news_body = ""

for fd in News_article:
    find1 = fd.findAll("p")
    for i in find1:
        # 找到当前p标签下所有class为ads的div
        ad_elements = i.find_all('div', class_='ads')
        # 逐个移除这些广告节点
        for ad in ad_elements:
            ad.decompose()
        # 提取干净的文本,strip=True去除多余的空格和换行
        news_body += i.get_text(strip=True) + " "

# 去除末尾多余的空格
news_body = news_body.strip()
print(news_body)

方法二:直接筛选非广告的文本节点

如果你不想修改DOM结构,可以直接提取<p>中不属于广告div的文本节点:

import bs4

soup = bs4.BeautifulSoup(page.content, 'html.parser')
news = soup.find('div',{'class': 'col-md-8 left-container details'})
News_article = news.find_all('div',{'class': 'news-article'})
news_body = ""

for fd in News_article:
    find1 = fd.findAll("p")
    for i in find1:
        # 获取所有非广告div下的文本片段
        text_segments = []
        for text in i.stripped_strings:
            # 检查文本的父节点是不是广告div
            if text.parent.get('class') != ['ads']:
                text_segments.append(text)
        # 拼接文本片段
        news_body += " ".join(text_segments) + " "

news_body = news_body.strip()
print(news_body)

两种方法都能得到你期望的结果:

Location News Content 1 News Content 2 News Content 3 News Content 4

内容的提问来源于stack exchange,提问作者RajatRaja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 19:02:34