如何通过Dash上传PDF并展示从中提取的Pandas DataFrame
问题描述
我正在从PDF中提取数据并转换为Pandas DataFrame,目前用Dash教程里的代码在Dash应用上展示数据。接下来想实现上传PDF的功能,而不是在代码里预先指定PDF文件。能找到CSV的类似示例,但PDF没法按同样方式实现。
原代码如下:
pdfFileObj = open('test.pdf', 'rb') pdfReader = PyPDF2.PdfFileReader(pdfFileObj) #some operations on pdf to produce df1 and df2 using PyPDF2 app = Dash(__name__) app.layout = html.Div([ html.H4('Some title'), html.P(id='table_out'), dash_table.DataTable( id='table', columns=[{"name": i, "id": i} for i in df1.columns], data=df1.to_dict('records'), style_cell=dict(textAlign='left'), style_header=dict(backgroundColor="paleturquoise"), style_data=dict(backgroundColor="lavender") ), html.H4("Some title"), html.P(id='table_out1'), dash_table.DataTable( id='table1', columns=[{"name": i, "id": i} for i in df2.columns], data=df2.to_dict('records'), style_cell=dict(textAlign='left'), style_header=dict(backgroundColor="paleturquoise"), style_data=dict(backgroundColor="lavender") ) ]) @app.callback( Output('table_out', 'children'), Input('table', 'active_cell')) @app.callback( Output('table_out1', 'children'), Input('table1', 'active_cell')) def update_graphs(active_cell): if active_cell: cell_data = df1.iloc[active_cell['row']][active_cell['column_id']] cell_data2 = df2.iloc[active_cell['row']][active_cell['column_id']] return cell_data, cell_data2 #return f"Data: \"{cell_data}\" from table cell: {active_cell}" return "Click the table" app.run_server(debug=True)
解决方案
要实现PDF上传功能,核心是用Dash的dcc.Upload组件接收文件,在回调里处理上传的PDF二进制数据,提取数据生成DataFrame后更新表格内容。具体实现步骤如下:
1. 添加文件上传组件
在布局中加入dcc.Upload,让用户可以选择或拖拽上传PDF:
dcc.Upload( id='upload-pdf', children=html.Div([ '拖拽PDF文件到这里,或者 ', html.A('选择文件') ]), style={ 'width': '100%', 'height': '60px', 'lineHeight': '60px', 'borderWidth': '1px', 'borderStyle': 'dashed', 'borderRadius': '5px', 'textAlign': 'center', 'margin': '10px' }, multiple=False # 限制单次上传单个PDF )
2. 改造表格组件
将原来硬编码的表格列和数据改为动态配置,由回调更新:
dash_table.DataTable( id='table', columns=[], # 初始为空,回调更新 data=[], # 初始为空,回调更新 style_cell=dict(textAlign='left'), style_header=dict(backgroundColor="paleturquoise"), style_data=dict(backgroundColor="lavender") ), dash_table.DataTable( id='table1', columns=[], data=[], style_cell=dict(textAlign='left'), style_header=dict(backgroundColor="paleturquoise"), style_data=dict(backgroundColor="lavender") )
3. 编写核心回调逻辑
包含两个回调:一个处理PDF上传并更新表格,另一个处理单元格点击展示数据。同时将PDF提取逻辑封装为函数,避免全局变量依赖。
完整代码示例:
import dash from dash import Dash, html, dash_table, dcc, Input, Output, State import PyPDF2 import pandas as pd import base64 app = Dash(__name__) app.layout = html.Div([ dcc.Upload( id='upload-pdf', children=html.Div([ '拖拽PDF文件到这里,或者 ', html.A('选择文件') ]), style={ 'width': '100%', 'height': '60px', 'lineHeight': '60px', 'borderWidth': '1px', 'borderStyle': 'dashed', 'borderRadius': '5px', 'textAlign': 'center', 'margin': '10px' }, multiple=False ), html.H4('表格1'), html.P(id='table_out'), dash_table.DataTable( id='table', columns=[], data=[], style_cell=dict(textAlign='left'), style_header=dict(backgroundColor="paleturquoise"), style_data=dict(backgroundColor="lavender") ), html.H4("表格2"), html.P(id='table_out1'), dash_table.DataTable( id='table1', columns=[], data=[], style_cell=dict(textAlign='left'), style_header=dict(backgroundColor="paleturquoise"), style_data=dict(backgroundColor="lavender") ) ]) # 封装PDF提取逻辑,替换成你自己的处理代码 def parse_pdf(contents): # 解码上传的Base64格式内容 content_type, content_string = contents.split(',') decoded_pdf = base64.b64decode(content_string) # 读取PDF并处理(替换为你的原有逻辑) pdf_reader = PyPDF2.PdfReader(decoded_pdf) # 示例:生成两个测试DataFrame,实际替换为你的PDF解析代码 df1 = pd.DataFrame({'列1': [1,2,3], '列2': ['a','b','c']}) df2 = pd.DataFrame({'列A': [4,5,6], '列B': ['x','y','z']}) return df1, df2 # 上传PDF后更新表格 @app.callback( [Output('table', 'columns'), Output('table', 'data'), Output('table1', 'columns'), Output('table1', 'data')], Input('upload-pdf', 'contents'), prevent_initial_call=True ) def update_tables(contents): if contents: df1, df2 = parse_pdf(contents) columns1 = [{"name": col, "id": col} for col in df1.columns] columns2 = [{"name": col, "id": col} for col in df2.columns] return columns1, df1.to_dict('records'), columns2, df2.to_dict('records') return [], [], [], [] # 处理表格单元格点击事件 @app.callback( [Output('table_out', 'children'), Output('table_out1', 'children')], [Input('table', 'active_cell'), Input('table1', 'active_cell')], [State('table', 'data'), State('table1', 'data')] ) def show_cell_data(active_cell1, active_cell2, data1, data2): if active_cell1: row = active_cell1['row'] col = active_cell1['column_id'] cell_val1 = data1[row][col] cell_val2 = data2[row][col] if row < len(data2) else '无对应数据' return f"表格1选中数据:{cell_val1}", f"表格2对应数据:{cell_val2}" elif active_cell2: row = active_cell2['row'] col = active_cell2['column_id'] cell_val2 = data2[row][col] cell_val1 = data1[row][col] if row < len(data1) else '无对应数据' return f"表格1对应数据:{cell_val1}", f"表格2选中数据:{cell_val2}" return "点击表格查看数据", "点击表格查看数据" if __name__ == '__main__': app.run_server(debug=True)
关键说明
dcc.Upload返回的contents是Base64编码字符串,必须解码为二进制数据才能被PyPDF2读取- 将PDF解析逻辑封装为
parse_pdf函数,方便后续维护和修改 - 单元格点击回调通过
State获取当前表格数据,避免依赖全局变量,适配动态生成的DataFrame - 设置
prevent_initial_call=True防止应用启动时触发上传回调
内容的提问来源于stack exchange,提问作者user20153724
相关产品推荐
相关产品推荐

