You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从BS4 ResultSet中移除属于另一ResultSet的元素?

新闻爬取工具:移除嵌套的Twitter推文元素问题

现有爬取代码

import requests
from bs4 import BeautifulSoup

url = 'https://www.moneyreview.gr/life-and-arts/86916/mia-apli-lysi-gia-to-rochalito-to-kolpo-poy-sozei-chiliades-gamoys/'

r1 = requests.get(url)
coverpage = r1.content
soup1 = BeautifulSoup(coverpage, 'html5lib')
title = soup1.find('h1').get_text()
article = requests.get(url)
article_content = article.content

soup_article = BeautifulSoup(article_content, 'html5lib')
body = soup_article.find_all('div', class_='entry-content')

待移除元素说明

文章内容中嵌套了带有twitter-tweet类的<blockquote>元素(推文内容),需要移除这些元素及内部所有内容,以获取干净的正文。

原用来定位推文元素的代码:

for elements in body:
   quote = soup1.find_all('blockquote', class_= "twitter-tweet")
   print(quote)

段落列表生成代码

通过提取<p>标签生成段落文本列表:

x = body[0].find_all('p')
list_paragraphs = []

for p in np.arange(0, len(x)):
    paragraph = x[p].text.replace("\n", " ")
    list_paragraphs.append(paragraph)

问题与尝试方案

尝试从list_paragraphs中移除推文内容,但以下两种方法均失败:

尝试方案1

l3 = [x for x in list_paragraphs if x not in my_list]
print(l3)

尝试方案2

for element in my_list:
    if element in list_paragraphs:
        list_paragraphs.remove(element)

可行解决方案

失败原因是:quote是BeautifulSoup的元素对象,而list_paragraphs是文本字符串,两者类型不匹配,无法直接匹配删除。正确做法是在提取段落前就从HTML结构中移除推文元素,步骤如下:

  1. 定位所有推文元素后,使用decompose()方法从DOM中彻底删除;
  2. 再提取<p>标签生成干净的段落列表。

修改后的完整代码:

import requests
from bs4 import BeautifulSoup

url = 'https://www.moneyreview.gr/life-and-arts/86916/mia-apli-lysi-gia-to-rochalito-to-kolpo-poy-sozei-chiliades-gamoys/'

# 统一请求页面,避免重复请求
r = requests.get(url)
soup = BeautifulSoup(r.content, 'html5lib')

# 提取标题
title = soup.find('h1').get_text()

# 获取文章主体容器
body = soup.find('div', class_='entry-content')

# 移除所有Twitter推文元素
twitter_tweets = body.find_all('blockquote', class_='twitter-tweet')
for tweet in twitter_tweets:
    tweet.decompose()  # 从HTML结构中删除元素及其子内容

# 提取干净的段落文本
list_paragraphs = []
for p in body.find_all('p'):
    paragraph = p.text.replace("\n", " ").strip()
    if paragraph:  # 过滤空段落
        list_paragraphs.append(paragraph)

# 输出结果
print(title)
print("\n".join(list_paragraphs))

内容的提问来源于stack exchange,提问作者Elli Kafritsa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 06:15:40