You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Dash中上传多数据集并为每个DataFrame生成图表

多数据集上传与批量图表展示的Dash实现

问题需求

现有简易Dash应用仅支持单个数据集上传并生成前两列的散点图,需要改造为支持多数据集上传,并为每个上传的数据集单独展示对应的散点图。

修改后的完整代码

import base64
import io
import plotly.graph_objs as go
import dash
from dash.dependencies import Input, Output
from dash import dcc, html, Patch
import pandas as pd

external_stylesheets = ["https://codepen.io/chriddyp/pen/bWLwgP.css"]

app = dash.Dash(__name__, external_stylesheets=external_stylesheets)
server = app.server

colors = {"graphBackground": "#F5F5F5", "background": "#ffffff", "text": "#000000"}

app.layout = html.Div(
    [
        dcc.Upload(
            id="upload-data",
            children=html.Div(["Drag and Drop or ", html.A("Select Files")]),
            multiple=True,
            style={
                'width': '100%',
                'height': '60px',
                'lineHeight': '60px',
                'borderWidth': '1px',
                'borderStyle': 'dashed',
                'borderRadius': '5px',
                'textAlign': 'center',
                'margin': '10px'
            }
        ),
        html.Div(id="graph-container"),
    ]
)

def parse_data(contents, filename):
    content_type, content_string = contents.split(",")
    decoded = base64.b64decode(content_string)
    try:
        if "csv" in filename:
            df = pd.read_csv(io.StringIO(decoded.decode("utf-8")))
        elif "xls" in filename:
            df = pd.read_excel(io.BytesIO(decoded))
        elif "txt" in filename or "tsv" in filename:
            df = pd.read_csv(io.StringIO(decoded.decode("utf-8")), delimiter=r"\s+")
    except Exception as e:
        print(e)
        return html.Div([f"处理文件 {filename} 时出错: {str(e)}"])
    return df

@app.callback(
    Output("graph-container", "children"),
    [Input("upload-data", "contents"), Input("upload-data", "filename")],
    prevent_initial_call=True
)
def update_graphs(contents, filenames):
    graph_list = []
    if contents and filenames:
        for content, filename in zip(contents, filenames):
            data_result = parse_data(content, filename)
            # 检查解析结果是否为DataFrame
            if isinstance(data_result, pd.DataFrame):
                # 验证前两列是否为数值类型
                if pd.api.types.is_numeric_dtype(data_result.iloc[:,0]) and pd.api.types.is_numeric_dtype(data_result.iloc[:,1]):
                    fig = go.Figure(
                        data=[go.Scatter(
                            mode="markers",
                            x=data_result.iloc[:,0],
                            y=data_result.iloc[:,1],
                            showlegend=False
                        )],
                        layout=go.Layout(
                            title=f"数据集: {filename}",
                            width=500,
                            height=500,
                            plot_bgcolor=colors["graphBackground"],
                            paper_bgcolor=colors["graphBackground"]
                        )
                    )
                    graph_list.append(html.Div([
                        dcc.Graph(figure=fig),
                        html.Hr()  # 添加分隔线区分不同图表
                    ]))
                else:
                    graph_list.append(html.Div([f"文件 {filename} 的前两列不是连续数值类型,无法生成散点图。"]))
            else:
                # 解析出错的提示
                graph_list.append(data_result)
    return graph_list

if __name__ == "__main__":
    app.run_server(debug=True)

关键修改说明

  • 回调输出目标调整:将原Output的"graph1", "figure"改为"graph-container", "children",因为要生成多个图表组件,而非单个Figure对象。
  • 遍历处理多文件:不再只取第一个文件,而是通过zip(contents, filenames)遍历所有上传的文件,逐个解析并生成图表。
  • 添加类型验证:新增对解析结果是否为DataFrame的判断,以及前两列是否为数值类型的校验,避免非合规数据导致报错。
  • 用户体验优化:为每个图表添加标题显示对应文件名,添加分隔线区分不同图表;上传组件增加样式,更直观。
  • 异常处理增强:解析出错时返回带文件名的错误提示,便于用户定位问题。

内容的提问来源于stack exchange,提问作者Niam45

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:46:08