网页爬虫技术需求:提取文章正文并排除首尾特定<p>标签
解决方案
核心思路
别靠字符串拆分来过滤内容,直接在解析HTML标签的阶段就精准筛选。针对需求制定以下规则:
- 保留所有标题标签(h1-h6)和列表项(li)
- 过滤包含「分享文章」「最后更新」字样的
<p>标签 - 过滤页面底部非正文的
<p>标签(比如社区邀请、邮件订阅相关内容)
修改后的代码
from bs4 import BeautifulSoup import requests def get_clean_article(url): # 注意:实际使用时建议添加请求头模拟浏览器,避免被反爬 resp = requests.get(url) soup = BeautifulSoup(resp.text, 'html.parser') article_box = soup.find(class_="article") if not article_box: return "未找到文章内容" clean_content = [] # 遍历目标标签类型 for tag in article_box.find_all(["h1","h2","h3","h4","h5","h6","p","li"]): if tag.name == "p": text = tag.text.strip() # 跳过不符合要求的p标签 if any(keyword in text for keyword in ["分享文章", "最后更新", "↓ Join the community ↓", "Email", "ago"]): continue if not text: # 跳过空段落 continue clean_content.append(tag.text.strip()) # 用空行分隔不同内容块,还原文章排版逻辑 return '\n\n'.join(clean_content) # 批量处理目标链接 target_urls = [ "https://www.traveloffpath.com/covid-19-travel-insurance-everything-you-need-to-know/", "https://www.traveloffpath.com/what-to-do-if-your-flight-is-delayed-or-canceled/?swcfpc=1" ] for idx, url in enumerate(target_urls, 1): print(f"=== 第{idx}篇文章内容 ===") print(get_clean_article(url)) print("\n" + "-"*50 + "\n")
代码优势
- 比原代码的字符串拆分更可靠,不会因为页面文本微调导致索引越界或内容截取错误
- 过滤规则清晰,后续要新增排除关键词,直接在
any()括号里加就行 - 保留了文章的基本排版逻辑,内容可读性更强
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

