使用BeautifulSoup爬取网页时,如何排除父div内指定class的子div?
解决思路与修正代码
问题核心是你只筛选了p标签,但entry-content下的.ymae是独立的div元素,这些元素里的文本并没有被排除。正确的做法是先把所有.ymae子元素从父容器中移除,再提取剩余内容。
修正后的代码如下:
import requests from bs4 import BeautifulSoup def scrape_minimalism(): base_url = "https://www.theminimalists.com/minimalism/" response = requests.get(base_url) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') title = soup.find('h1', class_='entry-title').text.strip() print('\n', title) # 定位主内容容器 entry_content = soup.find('div', class_='entry-content') # 移除所有class为ymae的子元素 for ymae_div in entry_content.find_all('div', class_='ymae'): ymae_div.extract() # 提取所有剩余的p标签内容 main_paragraphs = entry_content.find_all('p') main_pcontent = '\n'.join(paragraph.text.strip() for paragraph in main_paragraphs) print('\n', main_pcontent) scrape_minimalism()
关键改动说明
- 先定位
entry-content容器,再遍历所有.ymae子div,用extract()方法彻底移除这些元素(该方法会将元素从DOM树中删除,同时返回被删除的元素,这里我们不需要返回值) - 移除干扰元素后,直接提取所有
p标签即可,无需再过滤class,因为干扰项已经被清除
内容的提问来源于stack exchange,提问作者Leona Raine
相关产品推荐
相关产品推荐

