You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从各类RSS Feed中正确提取指定内容?

处理非标准RSS Feed的字段提取方案

一、抛弃正则,用专业解析库替代硬匹配

正则处理XML/HTML嵌套结构极易出错,直接用成熟解析工具构建基础流程:

  • 先用XML解析器(如Python的lxml或xml.etree.ElementTree)解析RSS结构,优先抓取标准标签:
    • 标题:优先取<title>(RSS2.0)或<entry><title>(Atom)
    • 日期:优先匹配<pubDate>(RSS2.0)、<updated>(Atom),用dateutil.parser自动兼容各种日期格式
    • 描述:先抓<description>(RSS2.0)或<summary>(Atom),若内容含HTML,用HTML解析器(如BeautifulSoup)做后续处理

二、非标准内容的分层提取逻辑

1. 图片提取

  • 先检查RSS规范内的<enclosure>标签(筛选type为image/*的项),这是官方定义的媒体附件字段
  • 若无符合项,将描述/扩展内容(如<content:encoded>)的HTML用BeautifulSoup解析,提取所有<img>标签的src属性,可通过尺寸判断或域名过滤排除小图标类资源
  • 注意部分Feed会把图片放在<content:encoded>扩展标签中,需优先解析该字段

2. 视频提取

  • 同图片逻辑,先匹配<enclosure>标签(筛选type为video/*的项)
  • 再解析HTML内容中的<video>标签src,或<iframe>中的视频嵌入链接(如YouTube、B站链接,可先存储链接,后续按需解析真实地址)

3. 字段兜底策略

若标准标签全无,遍历整个XML节点聚合所有文本与HTML内容,再定向提取:

  • 标题:从聚合内容中筛选长度在10-200字符区间、格式类似标题的文本(如<h1>/<h2>标签内的内容)
  • 日期:用dateutil.parser尝试解析所有可能的日期字符串,匹配成功则取第一个有效结果

三、解决HTML显示错乱问题

  • 提取到的HTML必须做清理:用BeautifulSoup的get_text()提取纯文本,或保留结构化HTML但过滤<script>/<style>/广告类冗余标签
  • 将相对路径的图片/链接转为绝对路径:以RSS的<link>字段为根域名,拼接相对路径
  • 修复HTML嵌套错误,统一标签闭合格式

四、简化代码的实用技巧

  • 封装通用提取函数,比如extract_field(tree, standard_tags, fallback_parser),传入标准标签列表与兜底解析逻辑,避免重复代码
  • 用XPath查询一次性定位多个可能标签,比如//title | //entry/title,无需逐个判断
  • 将日期解析、HTML清理封装为工具函数,实现逻辑复用

核心逻辑示例(Python)

from lxml import etree
from bs4 import BeautifulSoup
from dateutil import parser

def parse_rss(xml_content):
    tree = etree.fromstring(xml_content)
    ns = {'content': 'http://purl.org/rss/1.0/modules/content/'}
    
    # 提取标题
    title_nodes = tree.xpath('//title/text() | //entry/title/text()')
    title = title_nodes[0].strip() if title_nodes else ''
    
    # 提取日期
    date_nodes = tree.xpath('//pubDate/text() | //updated/text()')
    post_date = parser.parse(date_nodes[0]).isoformat() if date_nodes else ''
    
    # 提取描述/内容并清理
    content_nodes = tree.xpath('//description/text() | //summary/text() | //content:encoded/text()', namespaces=ns)
    content_html = content_nodes[0] if content_nodes else ''
    soup = BeautifulSoup(content_html, 'html.parser')
    clean_desc = soup.get_text(strip=True)
    
    # 提取图片
    images = []
    # 优先处理enclosure标签
    image_enclosures = tree.xpath('//enclosure[@type="image/jpeg" or @type="image/png"]')
    images.extend([enc.get('url') for enc in image_enclosures])
    # 从HTML中兜底提取
    if not images:
        img_tags = soup.find_all('img')
        images = [img.get('src') for img in img_tags if img.get('src')]
    
    # 提取视频
    videos = []
    # 优先处理enclosure标签
    video_enclosures = tree.xpath('//enclosure[@type="video/mp4" or @type="video/webm"]')
    videos.extend([enc.get('url') for enc in video_enclosures])
    # 从HTML中兜底提取
    if not videos:
        video_tags = soup.find_all('video')
        videos.extend([vid.get('src') for vid in video_tags if vid.get('src')])
        iframe_tags = soup.find_all('iframe')
        videos.extend([iframe.get('src') for iframe in iframe_tags if 'youtube' in iframe.get('src', '') or 'bilibili' in iframe.get('src', '')])
    
    return {
        'title': title,
        'post_date': post_date,
        'description': clean_desc,
        'images': images,
        'videos': videos
    }

五、极端非标准Feed的应对补充

  • 若Feed不符合XML规范导致解析报错,直接用BeautifulSoup将其当作HTML解析,再从中提取类RSS结构
  • 针对WordPress、Substack等特定平台Feed,添加针对性规则(如WordFeed优先取<content:encoded>,Substack图片多在<figure><img>内)
  • 增加容错机制:解析失败时返回空字段而非崩溃,同时记录错误日志以便后续优化规则

内容的提问来源于stack exchange,提问作者Giuseppe De Paola

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 20:20:24