Python遍历URL生成RSS Feed时内层循环未执行问题求助
问题:RSS聚合时内层循环未生成匹配条目
我正在聚合多个RSS Feed,根据关键词匹配提取条目,同时希望获取源RSS Feed的标题。但运行代码后发现仅执行了外层循环,内层循环未生成对应条目,以下是代码及输出XML:
运行代码
import feedparser from feedgen.feed import FeedGenerator from dateutil import parser from dateutil import tz from pytz import timezone # Define a custom time zone mapping for EDT tzinfos = {"PST": tz.gettz("America/Los_Angeles"), "EDT": tz.gettz("America/New_York")} feeds = [ 'https://vancouversun.com/feed/?x=1', 'https://rss.cbc.ca/lineup/canada-britishcolumbia.xml', 'https://rss.cbc.ca/lineup/topstories.xml', # Add more feeds as needed ] keywords = ['autism', 'autistic', 'autisme', 'autistique', 'asperger', '#autism', 'neurodiversity'] # Create a new feed generator object. fg = FeedGenerator() fg.title('Autism News') fg.link(href='https://rcastonguay.github.io/autismfeeds/index.xml', rel='self') fg.description('An RSS feed filtered by autism keywords.') # Iterate over feeds for feed_url in feeds: d = feedparser.parse(feed_url) # Retrieve the source title if d.feed.get('title'): source_title = d.feed.title # Use d.feed.title to get the feed title else: source_title = feed_url # If feed title is not available, use the feed URL as the source title # Add source title as a regular entry in the feed fe = fg.add_entry() fe.title(f"{source_title}\n") # Iterate over entries for each feed and add them to the feed for entry in d.entries: if any(keyword in entry.title.lower() or keyword in entry.summary.lower() or keyword in entry.description.lower() for keyword in keywords): fe = fg.add_entry() fe.title(entry.title) fe.link(href=entry.link) fe.description(entry.description) date = parser.parse(entry.published, fuzzy=True, tzinfos=tzinfos) if date.tzinfo is None or date.tzinfo.utcoffset(date) is None: date = date.replace(tzinfo=tz.gettz('UTC')).astimezone(tz.gettz('PST')) else: date = date.astimezone(tz.gettz('PST')) fe.pubDate(date) # Generate the RSS feed XML file. rssfeed = fg.rss_str(pretty=True) print(rssfeed.decode())
输出XML
<?xml version='1.0' encoding='UTF-8'?> <rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0"> <channel> <title>Autism News</title> <link>https://rcastonguay.github.io/autismfeeds/index.xml</link> <description>An RSS feed filtered by autism keywords.</description> <atom:link href="https://rcastonguay.github.io/autismfeeds/index.xml" rel="self"/> <docs>http://www.rssboard.org/rss-specification</docs> <generator>python-feedgen</generator> <lastBuildDate>Thu, 23 Nov 2023 20:36:23 +0000</lastBuildDate> <item> <title>CBC | Top Stories News </title> </item> <item> <title>CBC | British Columbia News </title> </item> <item> <title>Vancouver Sun </title> </item> </channel> </rss>
解决建议
修复关键词匹配的健壮性问题
原代码直接访问entry.summary和entry.description,如果RSS条目缺少这些字段会触发AttributeError,导致匹配逻辑直接失败。修改为先判断字段是否存在,再收集文本进行匹配:# 替换内层循环的if判断部分 entry_texts = [] if hasattr(entry, 'title'): entry_texts.append(entry.title.lower()) if hasattr(entry, 'summary'): entry_texts.append(entry.summary.lower()) if hasattr(entry, 'description'): entry_texts.append(entry.description.lower()) if any(keyword in text for text in entry_texts for keyword in keywords): # 后续添加条目的代码保持不变验证目标RSS Feed是否含有关键词内容
手动访问你列出的RSS链接,确认其中是否存在包含指定关键词的条目。如果当前这些Feed里没有匹配内容,内层循环自然不会生成任何条目。添加调试日志排查问题
在内层循环中加入打印语句,查看每个条目的处理情况,是否触发异常:for entry in d.entries: print(f"处理条目:{entry.title}") try: # 原匹配逻辑和条目添加代码 except Exception as e: print(f"处理条目出错:{str(e)}")优化RSS结构(可选)
原代码把源Feed标题作为普通条目添加到结果中,这会干扰最终的RSS内容结构。如果不需要展示源标题,直接删除以下两行代码:fe = fg.add_entry() fe.title(f"{source_title}\n")
内容的提问来源于stack exchange,提问作者Remi Castonguay
相关产品推荐
相关产品推荐

