使用Python tabula库转换PDF为CSV/XLS时如何保留两位小数
解决方案
- 你当前直接调用
tabula.convert_into()的方式没有中间处理DataFrame的步骤,无法控制输出的小数精度,需要先手动读取PDF表格内容、调整数值格式后再导出CSV,调整后代码如下:
import tabula import pandas as pd # 读取PDF时添加参数保证数值读取完整,lattice=True适用于带明确边框的表格,无框表格换成stream=True df_list = tabula.read_pdf("File1.pdf", pages='all', lattice=True, pandas_options={'dtype': str}) # 合并多页表格为单个DataFrame,各页表头不一致的话可以删除该行逐页处理 full_df = pd.concat(df_list, ignore_index=True) # 遍历所有列,将可转换为数值的列统一设置保留两位小数 for col in full_df.columns: # 尝试转换为数值,转换失败的文本类列直接跳过 full_df[col] = pd.to_numeric(full_df[col], errors='ignore') # 仅对数值类型列做格式化处理 if pd.api.types.is_numeric_dtype(full_df[col]): full_df[col] = full_df[col].apply(lambda x: f"{x:.2f}" if pd.notna(x) else x) # 导出处理后的内容到CSV,不导出pandas自带的索引列 full_df.to_csv("File1.csv", index=False, encoding='utf-8-sig')
- 批量处理目录下所有PDF的调整后代码如下:
import os import tabula import pandas as pd input_dir = "input_directory" for filename in os.listdir(input_dir): if filename.lower().endswith(".pdf"): pdf_path = os.path.join(input_dir, filename) csv_path = os.path.join(input_dir, filename.rsplit('.', 1)[0] + ".csv") df_list = tabula.read_pdf(pdf_path, pages='all', lattice=True, pandas_options={'dtype': str}) full_df = pd.concat(df_list, ignore_index=True) for col in full_df.columns: full_df[col] = pd.to_numeric(full_df[col], errors='ignore') if pd.api.types.is_numeric_dtype(full_df[col]): full_df[col] = full_df[col].apply(lambda x: f"{x:.2f}" if pd.notna(x) else x) full_df.to_csv(csv_path, index=False, encoding='utf-8-sig')
如果不需要给一位小数的数值补0(例如不希望1.2变成1.20),可以把格式化逻辑替换为`lambda x: f"{x:.2f}".rstrip('0').rstrip('.') if pd.notna(x) and '.' in f"{x:.2f}" else f"{x:.2f}"即可。
内容的提问来源于stack exchange,提问作者linux01
相关产品推荐
相关产品推荐

