You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何用Pandas无法将英国新闻爬虫数据存入SQL数据库?

英国新闻爬虫数据无法存入数据库的问题分析

你的代码存在多个逻辑错误,导致数据无法正确构建并写入数据库,具体问题如下:

1. 数据收集逻辑完全混乱

你在循环处理新闻条目时,没有按每条新闻对应一个包含title、url、summary、date的结构来构建数据,而是零散地把单个字段值塞进data列表,甚至出现data.append(data)这种递归添加自身的错误操作,最终data的结构完全不符合DataFrame的要求。

比如你当前的错误代码片段:

data = []
for entry in parsed_feed.entries:
    title = entry.title
    url = entry.link
    print(entry.summary)
    if entry.summary:
        summary = entry.summary
        data.append(summary)
    else:
        entry.summary = "No summary available"
    if entry.published:
        date = entry.published
        data.append(data)  # 致命错误:把data自身添加进去
    else:
        data.append("No data available")

这段代码只会生成一个杂乱的一维列表,根本无法对应你要设置的四列字段。

2. DataFrame列数不匹配

你创建DataFrame后强行设置四列['title', 'url', 'summary', 'date'],但此时data是一维结构,DataFrame只有1列,这会直接抛出ValueError,导致后续代码无法执行。

3. 字符串与DataFrame拼接错误

print("data" + df)这行代码试图把字符串和DataFrame直接拼接,会触发TypeError,中断程序执行,后续的数据库写入代码根本没机会运行。

4. 冗余函数定义

你重复定义了两次poll_rss函数,虽然没被调用,但属于不必要的冗余代码。


修正后的代码示例

以下是修复了上述问题的完整代码:

import feedparser
import pandas as pd
from sqlalchemy import create_engine

def poll_rss(rss_url):
    feed = feedparser.parse(rss_url)
    entries_data = []
    for entry in feed.entries:
        # 每条新闻构建一个字典,统一收集所有需要的字段
        news_item = {
            "title": entry.get("title", "No title available"),
            "url": entry.get("link", "No url available"),
            "summary": entry.get("summary", "No summary available"),
            "date": entry.get("published", "No date available")
        }
        entries_data.append(news_item)
    return entries_data

# 新闻源列表(已去重重复的Sky News和The Mirror)
feeds = [
    {"type": "news","title": "BBC", "url": "http://feeds.bbci.co.uk/news/uk/rss.xml"},
    {"type": "news","title": "The Economist", "url": "https://www.economist.com/international/rss.xml"},    
    {"type": "news","title": "The New Statesman", "url": "https://www.newstatesman.com/feed"},    
    {"type": "news","title": "The New York Times", "url": "https://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml"},
    {"type": "news","title": "Metro UK","url": "https://metro.co.uk/feed/"},
    {"type": "news", "title": "Evening Standard", "url": "https://www.standard.co.uk/rss.xml"},
    {"type": "news","title": "Daily Mail", "url": "https://www.dailymail.co.uk/articles.rss"},
    {"type": "news","title": "Sky News", "url": "https://news.sky.com/feeds/rss/home.xml"},
    {"type": "news","title": "The Mirror", "url": "https://www.mirror.co.uk/news/rss.xml"},
    {"type": "news","title": "The Sun", "url": "https://www.thesun.co.uk/news/feed/"},
    {"type": "news","title": "The Guardian", "url": "https://www.theguardian.com/uk/rss"},
    {"type": "news","title": "The Independent", "url": "https://www.independent.co.uk/news/uk/rss"},
    {"type": "news","title": "The Telegraph", "url": "https://www.telegraph.co.uk/news/rss.xml"},
    {"type": "news","title": "The Times", "url": "https://www.thetimes.co.uk/?service=rss"}
]

# 收集所有新闻数据
all_data = []
for feed in feeds:
    print(f"Processing: {feed['title']}")
    feed_data = poll_rss(feed['url'])
    all_data.extend(feed_data)
    print(f"Collected {len(feed_data)} articles\n")

# 转换为DataFrame
df = pd.DataFrame(all_data)
print("Data preview:")
print(df.head())

# 写入数据库
engine = create_engine('mysql+pymysql://root:password_thingbob@localhost/somedatabase')  
df.to_sql('nationals', con=engine, if_exists='append', index=False)
print("Data successfully written to database")

内容的提问来源于stack exchange,提问作者elksie5000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 10:05:28