You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取目标段落并去除冗余HTML?

解决方案

1. 过滤冗余段落并提取纯文本

你的问题核心是要剔除带elementor-icon-box-description类的无效<p>标签,同时剥离HTML标签获取纯文本。可以通过CSS伪类选择器精准筛选目标段落,再用BeautifulSoup的文本提取功能处理内容。

修正后的完整代码如下:

from urllib.request import urlopen
from bs4 import BeautifulSoup

# 补充页面内容获取步骤(你之前的代码缺少这部分)
url = "https://holypython.com/100-python-tips-tricks/"
html = urlopen(url).read()
soup = BeautifulSoup(html, features="html.parser")

# 选择所有不带elementor-icon-box-description类的p标签
target_paragraphs = soup.select('p:not(.elementor-icon-box-description)')

# 提取纯文本并过滤空内容
cleaned_content = []
for p in target_paragraphs:
    text = p.get_text(strip=True)
    if text:  # 跳过仅含空格的空段落
        cleaned_content.append(text)

# 输出整理后的内容
for item in cleaned_content:
    print(item)

2. 精准定位单个技巧内容(进阶)

如果要针对性提取每个Python技巧对应的说明段落,可以观察页面结构:每个技巧都包裹在特定的容器内,通常包含标题(<h3>)和对应说明。可以先定位技巧容器,再提取内部内容:

# 定位每个技巧的独立容器
tip_containers = soup.select('div.elementor-element[data-id]')

organized_tips = []
for container in tip_containers:
    # 提取技巧标题
    title_tag = container.select_one('h3')
    if not title_tag:
        continue
    title = title_tag.get_text(strip=True)
    
    # 提取技巧对应的有效段落
    content_tag = container.select_one('p:not(.elementor-icon-box-description)')
    content = content_tag.get_text(strip=True) if content_tag else ""
    
    organized_tips.append({
        "序号": len(organized_tips)+1,
        "技巧标题": title,
        "内容说明": content
    })

# 输出结构化的技巧内容
for tip in organized_tips:
    print(f"{tip['序号']}. {tip['技巧标题']}")
    print(f"   {tip['内容说明']}\n")

关键知识点说明

  • p:not(.elementor-icon-box-description):CSS选择器语法,直接排除带指定类的冗余<p>标签,减少无效数据。
  • .get_text(strip=True):自动剥离HTML标签,提取元素内纯文本,同时去除首尾空格和换行符,让内容更整洁。
  • 空段落过滤:避免提取到仅含空白字符的无效内容,提升结果质量。

内容的提问来源于stack exchange,提问作者Max Allen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 07:35:17