如何从各类RSS Feed中正确提取指定内容?
处理非标准RSS Feed的字段提取方案
一、抛弃正则,用专业解析库替代硬匹配
正则处理XML/HTML嵌套结构极易出错,直接用成熟解析工具构建基础流程:
- 先用XML解析器(如Python的
lxml或xml.etree.ElementTree)解析RSS结构,优先抓取标准标签:- 标题:优先取
<title>(RSS2.0)或<entry><title>(Atom) - 日期:优先匹配
<pubDate>(RSS2.0)、<updated>(Atom),用dateutil.parser自动兼容各种日期格式 - 描述:先抓
<description>(RSS2.0)或<summary>(Atom),若内容含HTML,用HTML解析器(如BeautifulSoup)做后续处理
- 标题:优先取
二、非标准内容的分层提取逻辑
1. 图片提取
- 先检查RSS规范内的
<enclosure>标签(筛选type为image/*的项),这是官方定义的媒体附件字段 - 若无符合项,将描述/扩展内容(如
<content:encoded>)的HTML用BeautifulSoup解析,提取所有<img>标签的src属性,可通过尺寸判断或域名过滤排除小图标类资源 - 注意部分Feed会把图片放在
<content:encoded>扩展标签中,需优先解析该字段
2. 视频提取
- 同图片逻辑,先匹配
<enclosure>标签(筛选type为video/*的项) - 再解析HTML内容中的
<video>标签src,或<iframe>中的视频嵌入链接(如YouTube、B站链接,可先存储链接,后续按需解析真实地址)
3. 字段兜底策略
若标准标签全无,遍历整个XML节点聚合所有文本与HTML内容,再定向提取:
- 标题:从聚合内容中筛选长度在10-200字符区间、格式类似标题的文本(如
<h1>/<h2>标签内的内容) - 日期:用
dateutil.parser尝试解析所有可能的日期字符串,匹配成功则取第一个有效结果
三、解决HTML显示错乱问题
- 提取到的HTML必须做清理:用
BeautifulSoup的get_text()提取纯文本,或保留结构化HTML但过滤<script>/<style>/广告类冗余标签 - 将相对路径的图片/链接转为绝对路径:以RSS的
<link>字段为根域名,拼接相对路径 - 修复HTML嵌套错误,统一标签闭合格式
四、简化代码的实用技巧
- 封装通用提取函数,比如
extract_field(tree, standard_tags, fallback_parser),传入标准标签列表与兜底解析逻辑,避免重复代码 - 用XPath查询一次性定位多个可能标签,比如
//title | //entry/title,无需逐个判断 - 将日期解析、HTML清理封装为工具函数,实现逻辑复用
核心逻辑示例(Python)
from lxml import etree from bs4 import BeautifulSoup from dateutil import parser def parse_rss(xml_content): tree = etree.fromstring(xml_content) ns = {'content': 'http://purl.org/rss/1.0/modules/content/'} # 提取标题 title_nodes = tree.xpath('//title/text() | //entry/title/text()') title = title_nodes[0].strip() if title_nodes else '' # 提取日期 date_nodes = tree.xpath('//pubDate/text() | //updated/text()') post_date = parser.parse(date_nodes[0]).isoformat() if date_nodes else '' # 提取描述/内容并清理 content_nodes = tree.xpath('//description/text() | //summary/text() | //content:encoded/text()', namespaces=ns) content_html = content_nodes[0] if content_nodes else '' soup = BeautifulSoup(content_html, 'html.parser') clean_desc = soup.get_text(strip=True) # 提取图片 images = [] # 优先处理enclosure标签 image_enclosures = tree.xpath('//enclosure[@type="image/jpeg" or @type="image/png"]') images.extend([enc.get('url') for enc in image_enclosures]) # 从HTML中兜底提取 if not images: img_tags = soup.find_all('img') images = [img.get('src') for img in img_tags if img.get('src')] # 提取视频 videos = [] # 优先处理enclosure标签 video_enclosures = tree.xpath('//enclosure[@type="video/mp4" or @type="video/webm"]') videos.extend([enc.get('url') for enc in video_enclosures]) # 从HTML中兜底提取 if not videos: video_tags = soup.find_all('video') videos.extend([vid.get('src') for vid in video_tags if vid.get('src')]) iframe_tags = soup.find_all('iframe') videos.extend([iframe.get('src') for iframe in iframe_tags if 'youtube' in iframe.get('src', '') or 'bilibili' in iframe.get('src', '')]) return { 'title': title, 'post_date': post_date, 'description': clean_desc, 'images': images, 'videos': videos }
五、极端非标准Feed的应对补充
- 若Feed不符合XML规范导致解析报错,直接用
BeautifulSoup将其当作HTML解析,再从中提取类RSS结构 - 针对WordPress、Substack等特定平台Feed,添加针对性规则(如WordFeed优先取
<content:encoded>,Substack图片多在<figure><img>内) - 增加容错机制:解析失败时返回空字段而非崩溃,同时记录错误日志以便后续优化规则
内容的提问来源于stack exchange,提问作者Giuseppe De Paola
相关产品推荐
相关产品推荐

