RSS XML指定标签对比Python脚本修复请求
问题排查与修复方案
常见问题根源
你的脚本无报错但无预期结果,大概率是以下几个核心问题:
- XML命名空间未处理:多数RSS/Atom Feed带命名空间(如Atom的
http://www.w3.org/2005/Atom),直接用标签名查找会返回空结果。 - 标签路径不匹配:不同RSS格式的层级结构不同(RSS2.0是
channel/item/title,Atom是feed/entry/title),硬编码路径会导致读不到内容。 - 内容未归一化:相同数据可能存在空格、HTML转义字符(如
&→&)、大小写差异,直接对比会误判为不同。 - 数据源读取静默失败:URL请求未处理HTTPS验证、重定向,文件读取未处理编码(如UTF-8 BOM),但未抛出异常。
修复后的完整脚本
以下是支持文件/URL数据源、处理命名空间与格式差异的修复版代码:
import requests import xml.etree.ElementTree as ET import html import re from urllib.parse import urlparse # 注册常见RSS/Atom命名空间,避免查找失败 NAMESPACES = { 'atom': 'http://www.w3.org/2005/Atom', 'rss': 'http://purl.org/rss/1.0/', 'dc': 'http://purl.org/dc/elements/1.1/' } def normalize_content(text): """归一化内容:去除多余空格、转义HTML实体、统一小写(可选)""" if not text: return '' # 转义HTML实体(如& → &) text = html.unescape(text) # 去除首尾空格、合并中间多余空格 text = re.sub(r'\s+', ' ', text).strip() # 可选:统一大小写,避免大小写差异导致误判 # text = text.lower() return text def parse_rss(root): """自动识别RSS格式,提取指定标签(以title和link为例)""" items = [] # 尝试RSS2.0格式(无命名空间) rss_items = root.findall('.//item') if rss_items: for item in rss_items: title = normalize_content(item.find('title').text if item.find('title') is not None else '') link = normalize_content(item.find('link').text if item.find('link') is not None else '') items.append((title, link)) return items # 尝试Atom格式(带命名空间) atom_entries = root.findall('.//atom:entry', NAMESPACES) if atom_entries: for entry in atom_entries: title = normalize_content(entry.find('atom:title', NAMESPACES).text if entry.find('atom:title', NAMESPACES) is not None else '') link_elem = entry.find('atom:link', NAMESPACES) link = normalize_content(link_elem.get('href') if link_elem is not None else '') items.append((title, link)) return items # 尝试其他RSS 1.0格式 rss1_items = root.findall('.//rss:item', NAMESPACES) if rss1_items: for item in rss1_items: title = normalize_content(item.find('rss:title', NAMESPACES).text if item.find('rss:title', NAMESPACES) is not None else '') link = normalize_content(item.find('rss:link', NAMESPACES).text if item.find('rss:link', NAMESPACES) is not None else '') items.append((title, link)) return items return [] def get_rss_content(source): """读取文件或URL的RSS内容,返回归一化后的(title, link)列表""" try: if urlparse(source).scheme in ('http', 'https'): # 处理URL请求:允许重定向、跳过HTTPS验证(可选)、设置超时 response = requests.get(source, allow_redirects=True, verify=False, timeout=10) response.encoding = response.apparent_encoding # 自动识别编码 root = ET.fromstring(response.content) else: # 处理本地文件:指定UTF-8编码,避免BOM问题 tree = ET.parse(source, ET.XMLParser(encoding='utf-8')) root = tree.getroot() return parse_rss(root) except Exception as e: # 抛出异常,避免静默失败 raise RuntimeError(f"读取数据源失败:{str(e)}") def compare_rss(source1, source2, output_file='rss_diff.txt'): """对比两个RSS数据源的指定标签,生成差异文件""" content1 = dict(get_rss_content(source1)) content2 = dict(get_rss_content(source2)) diffs = [] # 找出source1有但source2没有的内容 for title, link in content1.items(): if title not in content2: diffs.append(f"仅在数据源1存在:\n标题: {title}\n链接: {link}\n") elif content2[title] != link: diffs.append(f"内容不一致:\n标题: {title}\n数据源1链接: {link}\n数据源2链接: {content2[title]}\n") # 找出source2有但source1没有的内容 for title, link in content2.items(): if title not in content1: diffs.append(f"仅在数据源2存在:\n标题: {title}\n链接: {link}\n") # 写入差异文件,无差异也会生成空文件或提示 with open(output_file, 'w', encoding='utf-8') as f: if diffs: f.write('\n'.join(diffs)) else: f.write("两个RSS数据源指定标签内容完全一致") if __name__ == '__main__': # 示例用法:支持本地文件或URL compare_rss('https://example.com/feed1.xml', './feed2.xml', 'rss_diff.txt')
关键修复点说明
- 命名空间处理:注册了常见的RSS/Atom命名空间,自动识别不同格式的Feed结构,避免标签查找失败。
- 内容归一化:通过
normalize_content函数处理HTML转义字符、多余空格,消除格式差异导致的误判。 - 数据源容错:URL请求处理重定向、编码识别,文件读取指定UTF-8编码,同时抛出异常避免静默失败。
- 自动格式识别:自动适配RSS2.0、Atom、RSS1.0等常见格式,无需硬编码标签路径。
- 完善的对比逻辑:不仅检测内容存在性差异,还会对比相同标题下的内容(如链接)是否不一致,无差异时也会生成提示文件。
测试验证步骤
- 替换示例中的数据源路径/URL为你的实际地址。
- 运行脚本后,查看生成的
rss_diff.txt文件:- 若内容为提示文本,说明两个数据源指定标签无差异。
- 若有差异,会明确标注差异类型(仅某数据源存在、内容不一致)。
内容的提问来源于stack exchange,提问作者signorz
相关产品推荐
相关产品推荐

