Python Dash访问数据库时dcc.store值无法更新问题求助
问题描述
使用Python Dash开发仪表盘,需求是每5秒访问MongoDB数据库并更新图表:首次获取数据后,将DataFrame中Count列的值以列表形式存入dcc.Store;后续每5秒获取新数据后追加到列表。但实际运行中dcc.Store无法保存更新值,每次都会回到初始状态。改用random.random()生成数据测试时,代码可正常实现列表累加功能。
附相关代码:
from dash import Dash, html, dcc, Input, Output, State import plotly.express as px import dash import time import random import pandas as pd import numpy as np import json import asyncio from bson import ObjectId from pathlib import Path from datetime import datetime from concurrent.futures import ThreadPoolExecutor from dash.exceptions import PreventUpdate from get_data import init # DASH APP app = Dash(__name__) # Interval in milliseconds interval = 5000 app.layout = html.Div( children=[ html.H1(children="Time Series Reach Analysis"), html.P(children="Time Series Dataset from August 2021 to October 2022"), html.Div(id="timeseries-graph", children=[]), html.Div(id="interval-counter", children=[]), dcc.Store(id='df-storage', data=[], storage_type="memory"), dcc.Interval(id='interval-time', interval=interval, n_intervals=0) ], style={ "text-align" : "center" } ) @app.callback( # Returned Values Output("df-storage", "data"), Output("interval-counter", "children"), # Parameters Input("interval-time", "n_intervals"), State("df-storage", "data") ) def refresh_graph(n_intervals, data): storage_value = [] if n_intervals == 0: fetched_df = init() print("0 - Df = ", fetched_df) data.extend(fetched_df["_id"].tolist()) print(f"00 - Values for n_intervals-{n_intervals} and data-{data}") return data , f"Intervals : {n_intervals}" else: fetched_df = init() data.extend(fetched_df["_id"].tolist()) return data , f"Intervals : {n_intervals}" @app.callback( Output("timeseries-graph", "children"), Input("df-storage", "data"), ) def create_graph(data): df = pd.DataFrame(data=data, columns=["_id"]) fig = px.line( df, y="_id", x=df.index.tolist() ) return dcc.Graph(figure=fig) if __name__ == '__main__': app.run_server(debug=True)
init()函数负责从MongoDB获取数据并转换为含时间戳和Count的DataFrame。
核心原因
问题出在MongoDB返回的_id是bson.ObjectId类型,而dcc.Store要求存储的数据必须是可JSON序列化的类型(如字符串、数字、列表、字典等)。ObjectId属于自定义对象,无法被JSON序列化,导致回调返回更新后的数据时,dcc.Store无法正确保存,下一次回调触发时拿到的依然是初始的空列表。而用random.random()生成的是float类型,属于可序列化的数值,所以能正常累加。
另外,代码中实际操作的是_id列,但需求描述是要存储Count列,这里可能存在笔误,也需要注意。
解决思路与代码修改
1. 序列化MongoDB的ObjectId(如果仍需使用_id列)
在init()函数中,将_id字段转换为字符串,确保数据可被JSON序列化:
def init(): # 原有从MongoDB查询数据的代码 # ... # 将ObjectId转换为字符串 fetched_df["_id"] = fetched_df["_id"].astype(str) return fetched_df
2. 改用需求中的Count列(推荐,更适合图表展示)
如果需求是存储Count列,直接替换_id为Count,并确保Count列是数值类型:
def init(): # 原有从MongoDB查询数据的代码 # ... # 确保Count列是数值类型(避免字符串转数值问题) fetched_df["Count"] = pd.to_numeric(fetched_df["Count"], errors="coerce") # 过滤无效值(可选) fetched_df = fetched_df.dropna(subset=["Count"]) return fetched_df
然后修改回调中的数据处理逻辑,用Count列替代_id:
@app.callback( Output("df-storage", "data"), Output("interval-counter", "children"), Input("interval-time", "n_intervals"), State("df-storage", "data") ) def refresh_graph(n_intervals, data): fetched_df = init() # 获取Count列的数值列表 new_values = fetched_df["Count"].tolist() # 创建新列表进行累加(避免直接修改State传入的data对象) updated_data = data + new_values return updated_data, f"Intervals : {n_intervals}"
3. 修复图表生成逻辑(对应Count列)
修改create_graph回调,使用Count列生成图表:
@app.callback( Output("timeseries-graph", "children"), Input("df-storage", "data"), ) def create_graph(data): df = pd.DataFrame(data=data, columns=["Count"]) fig = px.line(df, y="Count", x=df.index.tolist(), title="Time Series Count") return dcc.Graph(figure=fig)
额外注意事项
- 确保
init()每次返回的是增量数据:如果init()每次返回的是全量数据,需要添加去重逻辑,避免重复追加数据到dcc.Store中。 - 调试时可打印数据类型:在回调中加入
print(type(data), data),确认每次返回的数据是可序列化的列表,且值正确累加。 dcc.Store的storage_type:使用"memory"是合适的,但若需要页面刷新后保留数据,可改用"local"或"session",但同样要求数据可JSON序列化。
内容的提问来源于stack exchange,提问作者Xahram
相关产品推荐
相关产品推荐

