You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python从PDF财务报表提取Accounts Payable关联数值遇阻

问题描述
  • 目标:从PDF财务报表中提取**Accounts Payable(应付账款)**及对应2021、2022年的数值,预期输出格式为「Accounts Payable 219,200 290,000」
  • 现状:代码运行未达预期,已尝试打印文本行排查关键词、去除$符号排除干扰,但问题仍未解决
原始错误代码
import PyPDF2
import re
# Path: \"C:\\Users\\thavo\\Downloads\\ECCREPORT.pdf\"
# \"C:\\Users\\thavo\\Downloads\\FY23_Q1_Consolidated_Financial_Statements.pdf\"

with open(r\"C:\\Users\\thavo\\Downloads\\FY23_01_Consolidated_Financial_Statements.pdf\", \"rb\") as file:
    pdf_reader = PyPDF2.PdfReader(file)
    page = pdf_reader.pages[0]
    content_of_file = page.extract_text()
    lines_of_desired_page = content_of_file.split(\"\\n\")
    print(lines_of_desired_page)
    for line in lines_of_desired_page:
        if \"net\" in line.lower():
            financial_data = line.strip().split(\" \")#[-2:]
            values = [re.sub(',', '', value) for value in financial_data]
            print(values[1])
        else:
            print(\"Cant find\") 
            break

*注:代码存在核心逻辑错误:

  1. 循环遇到第一行不含"net"的内容就直接终止,无法遍历所有行查找目标关键词
  2. 目标是查找Accounts Payable,但判断条件写成了"net",完全偏离需求
  3. 字符串转义符冗余,路径中的\"属于无效写法*
修复后的代码
import PyPDF2
import re

# 替换为你的PDF文件路径
pdf_path = r"C:\Users\thavo\Downloads\FY23_01_Consolidated_Financial_Statements.pdf"

with open(pdf_path, "rb") as file:
    pdf_reader = PyPDF2.PdfReader(file)
    # 遍历所有页面查找目标数据(若确定页码可直接指定,如pdf_reader.pages[0])
    for page in pdf_reader.pages:
        content = page.extract_text()
        if not content:
            continue
        # 正则匹配Accounts Payable及后续两个年份的数值
        pattern = re.compile(r'Accounts Payable\s+([\d,]+)\s+([\d,]+)')
        match_result = pattern.search(content)
        if match_result:
            year_2021_val = match_result.group(1)
            year_2022_val = match_result.group(2)
            print(f"Accounts Payable {year_2021_val}    {year_2022_val}")
            break
    else:
        print("未找到Accounts Payable相关数据")
修复说明
  • 修正核心逻辑:将匹配关键词改为Accounts Payable,用正则表达式精准匹配目标文本格式,避免拆分字符串的不确定性
  • 移除无效终止逻辑:遍历所有行(甚至所有页面)查找目标,而非第一行不匹配就停止
  • 优化数据提取:正则直接捕获应付账款后的两个数值,自动忽略多余空格
  • 简化代码写法:清理无效转义符,添加注释提升可读性

内容的提问来源于stack exchange,提问作者Ty_Guy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 21:07:16