为何用Pandas无法将英国新闻爬虫数据存入SQL数据库?
英国新闻爬虫数据无法存入数据库的问题分析
你的代码存在多个逻辑错误,导致数据无法正确构建并写入数据库,具体问题如下:
1. 数据收集逻辑完全混乱
你在循环处理新闻条目时,没有按每条新闻对应一个包含title、url、summary、date的结构来构建数据,而是零散地把单个字段值塞进data列表,甚至出现data.append(data)这种递归添加自身的错误操作,最终data的结构完全不符合DataFrame的要求。
比如你当前的错误代码片段:
data = [] for entry in parsed_feed.entries: title = entry.title url = entry.link print(entry.summary) if entry.summary: summary = entry.summary data.append(summary) else: entry.summary = "No summary available" if entry.published: date = entry.published data.append(data) # 致命错误:把data自身添加进去 else: data.append("No data available")
这段代码只会生成一个杂乱的一维列表,根本无法对应你要设置的四列字段。
2. DataFrame列数不匹配
你创建DataFrame后强行设置四列['title', 'url', 'summary', 'date'],但此时data是一维结构,DataFrame只有1列,这会直接抛出ValueError,导致后续代码无法执行。
3. 字符串与DataFrame拼接错误
print("data" + df)这行代码试图把字符串和DataFrame直接拼接,会触发TypeError,中断程序执行,后续的数据库写入代码根本没机会运行。
4. 冗余函数定义
你重复定义了两次poll_rss函数,虽然没被调用,但属于不必要的冗余代码。
修正后的代码示例
以下是修复了上述问题的完整代码:
import feedparser import pandas as pd from sqlalchemy import create_engine def poll_rss(rss_url): feed = feedparser.parse(rss_url) entries_data = [] for entry in feed.entries: # 每条新闻构建一个字典,统一收集所有需要的字段 news_item = { "title": entry.get("title", "No title available"), "url": entry.get("link", "No url available"), "summary": entry.get("summary", "No summary available"), "date": entry.get("published", "No date available") } entries_data.append(news_item) return entries_data # 新闻源列表(已去重重复的Sky News和The Mirror) feeds = [ {"type": "news","title": "BBC", "url": "http://feeds.bbci.co.uk/news/uk/rss.xml"}, {"type": "news","title": "The Economist", "url": "https://www.economist.com/international/rss.xml"}, {"type": "news","title": "The New Statesman", "url": "https://www.newstatesman.com/feed"}, {"type": "news","title": "The New York Times", "url": "https://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml"}, {"type": "news","title": "Metro UK","url": "https://metro.co.uk/feed/"}, {"type": "news", "title": "Evening Standard", "url": "https://www.standard.co.uk/rss.xml"}, {"type": "news","title": "Daily Mail", "url": "https://www.dailymail.co.uk/articles.rss"}, {"type": "news","title": "Sky News", "url": "https://news.sky.com/feeds/rss/home.xml"}, {"type": "news","title": "The Mirror", "url": "https://www.mirror.co.uk/news/rss.xml"}, {"type": "news","title": "The Sun", "url": "https://www.thesun.co.uk/news/feed/"}, {"type": "news","title": "The Guardian", "url": "https://www.theguardian.com/uk/rss"}, {"type": "news","title": "The Independent", "url": "https://www.independent.co.uk/news/uk/rss"}, {"type": "news","title": "The Telegraph", "url": "https://www.telegraph.co.uk/news/rss.xml"}, {"type": "news","title": "The Times", "url": "https://www.thetimes.co.uk/?service=rss"} ] # 收集所有新闻数据 all_data = [] for feed in feeds: print(f"Processing: {feed['title']}") feed_data = poll_rss(feed['url']) all_data.extend(feed_data) print(f"Collected {len(feed_data)} articles\n") # 转换为DataFrame df = pd.DataFrame(all_data) print("Data preview:") print(df.head()) # 写入数据库 engine = create_engine('mysql+pymysql://root:password_thingbob@localhost/somedatabase') df.to_sql('nationals', con=engine, if_exists='append', index=False) print("Data successfully written to database")
内容的提问来源于stack exchange,提问作者elksie5000
相关产品推荐
相关产品推荐

