You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网页时,如何排除父div内指定class的子div?

解决思路与修正代码

问题核心是你只筛选了p标签,但entry-content下的.ymae是独立的div元素,这些元素里的文本并没有被排除。正确的做法是先把所有.ymae子元素从父容器中移除,再提取剩余内容。

修正后的代码如下:

import requests
from bs4 import BeautifulSoup

def scrape_minimalism():
    base_url = "https://www.theminimalists.com/minimalism/"
    
    response = requests.get(base_url)
    
    if response.status_code == 200:
        soup = BeautifulSoup(response.text, 'html.parser')
        
        title = soup.find('h1', class_='entry-title').text.strip()
        print('\n', title)
        
        # 定位主内容容器
        entry_content = soup.find('div', class_='entry-content')
        # 移除所有class为ymae的子元素
        for ymae_div in entry_content.find_all('div', class_='ymae'):
            ymae_div.extract()
        
        # 提取所有剩余的p标签内容
        main_paragraphs = entry_content.find_all('p')
        main_pcontent = '\n'.join(paragraph.text.strip() for paragraph in main_paragraphs)
        print('\n', main_pcontent)
            

scrape_minimalism()

关键改动说明

  • 先定位entry-content容器,再遍历所有.ymae子div,用extract()方法彻底移除这些元素(该方法会将元素从DOM树中删除,同时返回被删除的元素,这里我们不需要返回值)
  • 移除干扰元素后,直接提取所有p标签即可,无需再过滤class,因为干扰项已经被清除

内容的提问来源于stack exchange,提问作者Leona Raine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 09:12:49