You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping:如何过滤冗余内容并拆分目标网页的100条Python技巧

解决方法

要过滤冗余信息并提取100个独立的Python技巧,核心是精准定位存放技巧的HTML元素,而非提取整个页面的文本。以下是优化后的代码:

from urllib.request import urlopen
from bs4 import BeautifulSoup

url = "https://holypython.com/100-python-tips-tricks/"
html = urlopen(url).read()
soup = BeautifulSoup(html, features="html.parser")

# 定位所有技巧的标题(页面中每个技巧标题都用<h3>标签包裹)
tip_headings = soup.find_all("h3")

# 遍历标题,提取对应技巧内容
tips = []
for heading in tip_headings:
    # 提取技巧编号和标题
    tip_title = heading.get_text(strip=True)
    # 收集标题后的所有内容,直到遇到下一个<h3>标签
    content = []
    next_sibling = heading.next_sibling
    while next_sibling and next_sibling.name != "h3":
        # 技巧内容通常存放在p、pre、div标签中
        if next_sibling.name in ["p", "pre", "div"]:
            content.append(next_sibling.get_text(strip=True))
        next_sibling = next_sibling.next_sibling
    # 合并标题与内容,组成完整技巧项
    full_tip = f"{tip_title}\n{' '.join(content)}"
    tips.append(full_tip)

# 输出前5个技巧示例
for i, tip in enumerate(tips[:5], 1):
    print(f"技巧{i}:\n{tip}\n")

# 可选:将所有技巧保存到本地文件
with open("python_tips.txt", "w", encoding="utf-8") as f:
    for i, tip in enumerate(tips, 1):
        f.write(f"技巧{i}:\n{tip}\n\n")

代码说明

  • 精准定位元素:页面中每个Python技巧的标题都用<h3>标签包裹,通过soup.find_all("h3")直接获取所有技巧标题节点,自动避开导航栏、侧边栏等冗余区块。
  • 提取对应内容:每个标题后的兄弟元素(<p>、<pre>等)是技巧的具体内容,遍历兄弟节点直到遇到下一个<h3>,即可收集当前技巧的完整内容。
  • 格式整理:将标题与内容合并,最终得到100个独立的技巧项,方便查看或保存。

原代码冗余原因

原代码直接提取整个页面的文本,未区分页面功能区块,导致导航、菜单等无关文本被一并提取。通过定位特定标签,就能只获取目标内容。

内容的提问来源于stack exchange,提问作者Max Allen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 04:15:35