UiPath集成Python报错:Series真值模糊问题求解
问题:UiPath调用Python脚本时tabulate报错“Truth value of a Series is ambiguous”
在UiPath中调用Python脚本时,执行到table = tabulate(data,tablefmt="plain")行时触发以下错误:
Truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all()
原脚本代码:
import sys from tabula import read_pdf from tabulate import tabulate import pandas as pd import io def pdf_to_csv(file_path): # 读取文件的所有页面 data = read_pdf(file_path,pages = 'all',multiple_tables = True,stream = True,area=(94.03, 65.9,496.1, 969.36)) # 将结果转换为字符串表格格式 table = tabulate(data,tablefmt="plain") # 将表格转换为DataFrame df = pd.read_fwf(io.StringIO(table)) df.to_csv('temp.csv',index=False,index_label=False,header=False)
解决方案
错误原因
当tabula.read_pdf设置multiple_tables=True时,返回的data是DataFrame列表,但如果PDF中部分区域读取到的不是标准表格(比如单行/单列数据、读取异常),会混入Series对象。tabulate在处理混合的DataFrame和Series时,会触发布尔值判断的歧义错误。
修正方案1:直接处理DataFrame(推荐,更高效)
跳过tabulate中转步骤,直接过滤并合并有效表格后导出CSV:
import sys from tabula import read_pdf import pandas as pd def pdf_to_csv(file_path): # 读取PDF表格数据 data = read_pdf(file_path,pages = 'all',multiple_tables = True,stream = True,area=(94.03, 65.9,496.1, 969.36)) # 过滤出非空的有效DataFrame valid_tables = [df for df in data if isinstance(df, pd.DataFrame) and not df.empty] if not valid_tables: raise ValueError("未读取到有效表格数据") # 合并所有表格(若需保留分表可跳过此步,单独处理每个表格) combined_df = pd.concat(valid_tables, ignore_index=True) # 直接导出CSV combined_df.to_csv('temp.csv', index=False, header=False)
修正方案2:保留tabulate中转(仅当必须使用此流程时)
确保传入tabulate的每个元素都是可处理的列表格式,同时处理可能存在的Series:
import sys from tabula import read_pdf from tabulate import tabulate import pandas as pd import io def pdf_to_csv(file_path): data = read_pdf(file_path,pages = 'all',multiple_tables = True,stream = True,area=(94.03, 65.9,496.1, 969.36)) table_rows = [] for item in data: if isinstance(item, pd.DataFrame) and not item.empty: # 将DataFrame转为含表头的列表 table_rows.append(item.columns.tolist()) table_rows.extend(item.values.tolist()) elif isinstance(item, pd.Series) and not item.empty: # 将Series转为单列表格格式 table_rows.append([item.name]) table_rows.extend([[val] for val in item.values]) # 生成plain格式表格字符串 table = tabulate(table_rows, tablefmt="plain") # 读取为DataFrame并导出 df = pd.read_fwf(io.StringIO(table)) df.to_csv('temp.csv', index=False, index_label=False, header=False)
内容的提问来源于stack exchange,提问作者Umer Shahid
相关产品推荐
相关产品推荐

