Pandas DataFrame添加行报错:append属性不存在,concat也无效
解决pandas AttributeError: 'DataFrame' object has no attribute 'append'问题
问题根源
- pandas 2.0及以上版本正式移除了
DataFrame.append()方法,这是报错的直接原因。 - 原代码中
append的用法本身存在逻辑错误:传入多个单键字典组成的列表,会导致每个字典单独成为一行,最终生成的news_df存在大量空值,数据结构完全错误。
修复方案
推荐先收集所有新闻条目到列表,最后一次性转换为DataFrame的方式,既规避了append方法的问题,又比循环添加更高效。
修改后的完整代码:
import feedparser import pandas as pd from datetime import datetime # 读取存档数据,处理文件不存在的情况 try: archive = pd.read_csv("national_news_scrape.csv") except FileNotFoundError: archive = pd.DataFrame(columns=['source', 'title', 'date', 'summary', "url"]) pd.set_option('display.max_colwidth', None) # 新闻源列表 feeds = [{"type": "news","title": "BBC", "url": "http://feeds.bbci.co.uk/news/uk/rss.xml"}, {"type": "news","title": "The Economist", "url": "https://www.economist.com/international/rss.xml"}, {"type": "news","title": "The New Statesman", "url": "https://www.newstatesman.com/feed"}, {"type": "news","title": "The New York Times", "url": "https://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml"}, {"type": "news","title": "Metro UK","url": "https://metro.co.uk/feed/"}, {"type": "news", "title": "Evening Standard", "url": "https://www.standard.co.uk/rss.xml"}, {"type": "news","title": "Daily Mail", "url": "https://www.dailymail.co.uk/articles.rss"}, {"type": "news","title": "Sky News", "url": "https://news.sky.com/feeds/rss/home.xml"}, {"type": "news", "title": "The Mirror", "url": "https://www.mirror.co.uk/news/?service=rss"}, {"type": "news", "title": "The Sun", "url": "https://www.thesun.co.uk/news/feed/"}, {"type": "news", "title": "Sky News", "url": "https://news.sky.com/feeds/rss/home.xml"}, {"type": "news", "title": "The Guardian", "url": "https://www.theguardian.com/uk/rss"}, {"type": "news", "title": "The Independent", "url": "https://www.independent.co.uk/news/uk/rss"}, #{"type": "news", "title": "The Telegraph", "url": "https://www.telegraph.co.uk/news/rss.xml"}, {"type": "news", "title": "The Times", "url": "https://www.thetimes.co.uk/?service=rss"}] # 用列表统一收集新闻条目 news_entries = [] for feed in feeds: print(f"Scraping: {feed['title']}") d = feedparser.parse(feed['url']) for entry in d.entries: # 处理缺失的summary字段 summary = entry.summary if hasattr(entry, 'summary') else '' # 组装完整的单条新闻字典 news_item = { 'source': feed['title'], 'title': entry.title, 'url': entry.link, 'date': datetime(*entry.published_parsed[:6]), 'summary': summary } news_entries.append(news_item) # 一次性转换为DataFrame news_df = pd.DataFrame(news_entries) # 合并、去重、排序 combined = pd.concat([news_df, archive]).drop_duplicates(subset=['summary']) combined['date'] = pd.to_datetime(combined['date']) combined = combined.sort_values("date", ascending=False) # 保存结果 combined.to_csv('national_news_scrape.csv', index=False)
关键修改点说明
- 替换append为列表收集:彻底规避append方法被移除的问题,同时提升代码运行效率(循环操作列表比循环操作DataFrame开销小得多)。
- 修复数据结构错误:原代码将每个字段拆分为单独字典append,导致每行仅一个字段有值;现在每条新闻是完整字典,保证数据结构正常。
- 增加异常处理:首次运行无存档文件时自动初始化空DataFrame,避免报错。
- 规范列名访问:用
combined['date']替代combined.date,避免列名与DataFrame内置方法重名时的冲突。
内容的提问来源于stack exchange,提问作者elksie5000
相关产品推荐
相关产品推荐

