添加第5个URL后Python RSS阅读器的Pandas DataFrame为空问题排查
问题根源与解决方案
问题既不是Pandas DataFrame的问题,也不是Excel的问题,完全是代码的循环逻辑错误导致的。
错误分析
你现在的代码里,遍历URL列表时,每循环一次就把items变量重新赋值为当前URL返回的结果,这意味着最后只有最后一个URL的items会被保留。当你解开第五个URL的注释后,如果这个URL请求失败(比如被网站拦截)、或者返回的XML里没有item标签,items就会是空列表,后续提取标题、日期、链接的循环自然拿不到数据,最终生成的DataFrame就是空的。
另外,第五个URL大概率会拦截无请求标识的爬虫请求,直接返回空或者错误页面,这也会导致items为空。
修正后的代码
from bs4 import BeautifulSoup import requests import pandas as pd # 初始化列表,存放所有URL的item数据 items = [] urls=[ 'https://www.ccn-cert.cni.es/component/obrss/rss-ultimas-vulnerabilidades.feed', 'https://feeds.english.ncsc.nl/news.rss', 'https://www.cert.ssi.gouv.fr/feed/', 'https://cert.be/en/rss', 'https://www.cshub.com/rss/categories/attacks' ] # 遍历每个URL,把返回的item追加到总列表里 for url in urls: # 添加请求头,模拟浏览器访问,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: response = requests.get(url, headers=headers) response.raise_for_status() # 检查请求是否成功 content = BeautifulSoup(response.content, 'xml') current_items = content.find_all('item') items.extend(current_items) # 追加,而不是覆盖 except Exception as e: print(f"处理URL {url} 时出错: {e}") title = [] date = [] link = [] for item in items: # 处理可能缺失的字段,避免报错 title.append(item.title.text if item.title else '无标题') date.append(item.pubDate.text if item.pubDate else '无日期') link.append(item.link.text if item.link else '无链接') data_frame = pd.DataFrame({'Title': title, 'Date published':date, 'Link':link}) print(data_frame.to_string()) with pd.ExcelWriter('RSS.xlsx') as writer: data_frame.to_excel(writer, sheet_name='mysheet') workbook = writer.book worksheet = writer.sheets['mysheet'] worksheet.set_column(1,1,130) worksheet.set_column(2,2,40) worksheet.set_column(3,3,120)
关键修改点
- 把
items的初始化移到循环外,每次用extend()追加当前URL的item,而不是覆盖 - 添加了
User-Agent请求头,避免被目标网站拦截 - 增加了异常处理,单个URL出错不会导致整个程序崩溃
- 对item的字段做了空值判断,避免因某个item缺失字段而报错
内容的提问来源于stack exchange,提问作者Hugo
相关产品推荐
相关产品推荐

