You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫技术需求:提取文章正文并排除首尾特定<p>标签

解决方案

核心思路

别靠字符串拆分来过滤内容,直接在解析HTML标签的阶段就精准筛选。针对需求制定以下规则:

  • 保留所有标题标签(h1-h6)和列表项(li)
  • 过滤包含「分享文章」「最后更新」字样的<p>标签
  • 过滤页面底部非正文的<p>标签(比如社区邀请、邮件订阅相关内容)

修改后的代码

from bs4 import BeautifulSoup
import requests

def get_clean_article(url):
    # 注意:实际使用时建议添加请求头模拟浏览器,避免被反爬
    resp = requests.get(url)
    soup = BeautifulSoup(resp.text, 'html.parser')
    
    article_box = soup.find(class_="article")
    if not article_box:
        return "未找到文章内容"
    
    clean_content = []
    # 遍历目标标签类型
    for tag in article_box.find_all(["h1","h2","h3","h4","h5","h6","p","li"]):
        if tag.name == "p":
            text = tag.text.strip()
            # 跳过不符合要求的p标签
            if any(keyword in text for keyword in ["分享文章", "最后更新", "↓ Join the community ↓", "Email", "ago"]):
                continue
            if not text:  # 跳过空段落
                continue
        clean_content.append(tag.text.strip())
    
    # 用空行分隔不同内容块,还原文章排版逻辑
    return '\n\n'.join(clean_content)

# 批量处理目标链接
target_urls = [
    "https://www.traveloffpath.com/covid-19-travel-insurance-everything-you-need-to-know/",
    "https://www.traveloffpath.com/what-to-do-if-your-flight-is-delayed-or-canceled/?swcfpc=1"
]

for idx, url in enumerate(target_urls, 1):
    print(f"=== 第{idx}篇文章内容 ===")
    print(get_clean_article(url))
    print("\n" + "-"*50 + "\n")

代码优势

  • 比原代码的字符串拆分更可靠,不会因为页面文本微调导致索引越界或内容截取错误
  • 过滤规则清晰,后续要新增排除关键词,直接在any()括号里加就行
  • 保留了文章的基本排版逻辑,内容可读性更强

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 08:05:20