You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地运行正常,PythonAnywhere数据库无法填充的问题排查

问题

我在PythonAnywhere上已有数据库及对应凭证,目标是爬取多个新闻网站的数据,并通过Flask将数据展示到新网站。以下是网站的一段代码(已省略可正常运行的导入部分):

@app.route("/nationals")
def scrape_nationals():
    feeds = [{"type": "news","title": "BBC", "url": "http://feeds.bbci.co.uk/news/uk/rss.xml"},
        {"type": "news","title": "The Economist", "url": "https://www.economist.com/international/rss.xml"},
        {"type": "news","title": "The New Statesman", "url": "https://www.newstatesman.com/feed"},
        {"type": "news","title": "The New York Times", "url": "https://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml"},
        {"type": "news","title": "Metro UK","url": "https://metro.co.uk/feed/"},
        {"type": "news", "title": "Evening Standard", "url": "https://www.standard.co.uk/rss.xml"},
        {"type": "news","title": "Daily Mail", "url": "https://www.dailymail.co.uk/articles.rss"},
        {"type": "news","title": "Sky News", "url": "https://news.sky.com/feeds/rss/home.xml"},
        {"type": "news","title": "The Mirror", "url": "https://www.mirror.co.uk/news/?service=rss"},
        {"type": "news","title": "The Sun", "url": "https://www.thesun.co.uk/news/feed/"},
        {"type": "news", "title": "Sky News", "url": "https://news.sky.com/feeds/rss/home.xml"},
        {"type": "news", "title": "The Guardian", "url": "https://www.theguardian.com/uk/rss"},
        {"type": "news", "title": "The Independent", "url": "https://www.independent.co.uk/news/uk/rss"},
        #{"type": "news", "title": "The Telegraph", "url": "https://www.telegraph.co.uk/news/rss.xml"},
        {"type": "news", "title": "The Times", "url": "https://www.thetimes.co.uk/?service=rss"}]
    print(feeds)

    data = []                               # <---- initialize empty list here
    for feed in feeds:
        parsed_feed = feedparser.parse(feed['url'])
        #print("Title:", feed['title'])
        #print("Number of Articles:", len(parsed_feed.entries))
        #print("
")
        for entry in parsed_feed.entries:

            title = entry.title
            print(title)
            url = entry.link
            #print(entry.summary)
            try:
                summary = entry.summary[:400] or "No summary available" # I simplified the ternary operators here
            except:
                #print("no summary")
                summary = "none"
            try:
                date = pd.to_datetime(entry.published)#
                #or "No data available"     # I simplified the ternary operators here
            except:
                #print("date")
                date = pd.to_datetime("01-01-1970")
            data.append([title, url, summary, date])          # <---- append data from each entry here

    df = pd.DataFrame(data, columns = ['title', 'url', 'summary', 'date'])
    articles = pd.read_sql('nationals', con = engine)
    articles = articles.drop_duplicates()
    df = df.append(articles)
    df = df.drop_duplicates()
    df.to_sql('nationals', con = engine, if_exists = 'replace', index = False)

这段代码在本地VSCode中运行正常,但无法填充PythonAnywhere上的nationals表,请问问题出在哪里?

可能的问题及解决思路
  • 数据库连接配置错误
    PythonAnywhere的数据库连接字符串和本地差异很大,本地可能用SQLite或本地MySQL,而PythonAnywhere要求特定格式:比如MySQL需用mysql+pymysql://username:password@username.mysql.pythonanywhere-services.com/username$database_name。检查你的engine配置是否完全匹配,尤其注意数据库名带$前缀(如yourusername$nationals_db),用户名、密码有没有输错。

  • 网络访问限制
    PythonAnywhere免费账户对外访问有部分限制,部分RSS源可能屏蔽了平台IP,或者请求被目标网站识别为非浏览器请求导致拦截。可以在PythonAnywhere控制台手动执行feedparser.parse("目标RSS链接")测试能否获取内容;若失败,要么更换RSS源,要么给feedparser添加请求头模拟浏览器,比如:

    parsed_feed = feedparser.parse(feed['url'], request_headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'})
    
  • 静默异常导致流程中断
    代码中的try-except块过于宽泛,没有捕获具体异常,爬取或数据库操作出错时无法察觉。建议把except:改为except Exception as e:并打印异常信息,比如:

    try:
        summary = entry.summary[:400] or "No summary available"
    except Exception as e:
        print(f"获取摘要失败: {str(e)}")
        summary = "none"
    

    另外,可查看PythonAnywhere的Web应用日志(页面"Logs"选项卡),里面会记录运行时的报错详情。

  • Pandas版本兼容性问题
    本地与PythonAnywhere的Pandas版本可能不一致,df.append()在Pandas 2.0+已被弃用,若PythonAnywhere上的版本为2.0+,该语句会报错,导致后续数据库写入未执行。建议替换为Pandas推荐的pd.concat写法:

    df = pd.concat([df, articles], ignore_index=True)
    
  • 数据库权限不足
    检查PythonAnywhere上的数据库用户是否拥有CREATE、INSERT、REPLACE权限,权限不足会导致to_sql的if_exists='replace'执行失败。可在PythonAnywhere数据库管理页面查看用户权限,或手动执行SQL语句测试修改表的权限。

  • 请求超时中断执行
    PythonAnywhere的Web请求有超时限制(免费账户通常为30秒),爬取10多个RSS源可能未完成就被中断,导致数据未写入数据库。建议将爬取逻辑改为后台任务(使用PythonAnywhere的Tasks功能),或减少单次爬取的RSS源数量,同时优化爬取速度(比如采用异步请求)。

内容的提问来源于stack exchange,提问作者elksie5000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 05:55:04