You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Camelot提取PDF表格报TypeError错误求助

解决Camelot读取PDF出现TypeError: bytes or integer address expected instead of str instance

修复步骤及优化方案

1. 删除无效代码

代码中第一行open("YP-brochure-draft_final_08_2019.pdf","rb")未赋值给变量,属于无效操作且会造成文件句柄泄漏,直接删除该行。

2. 排查依赖环境

该错误多与Camelot依赖的Ghostscript或底层库调用异常有关:

  • 确认已安装Ghostscript,并将其路径添加至系统环境变量PATH。
  • 安装兼容的稳定版本依赖:
    pip install camelot-py[cv]==0.10.1 ghostscript
    

3. 确保PDF路径正确

保证PDF文件与脚本在同一目录,或使用绝对路径(Windows系统需转义路径,如C:\\docs\\file.pdf,或用原始字符串r"C:\docs\file.pdf")。

4. 优化后的可运行代码

import camelot

# 替换为你的PDF文件路径
file_name = "YP-brochure-draft_final_08_2019.pdf"

# 提取表格,根据表格结构选择flavor:lattice(有边框)/stream(无边框)
tables = camelot.read_pdf(file_name, flavor='lattice')

print(f"Total tables extracted: {tables.n}")

# 输出并保存第一个表格
if tables.n > 0:
    print(tables[0].df)
    tables[0].to_csv("extracted_table.csv")

5. 额外调试建议

  • 若提取效果不佳,尝试切换flavor参数。
  • 验证PDF文件是否损坏,或用其他工具确认表格结构是否清晰。

内容的提问来源于stack exchange,提问作者freelance_designer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 18:10:02