You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PHPExcel/PHPSpreadsheet或Python提取Excel中嵌入的PDF

提取Excel中嵌入的PDF文件方案

PHP 实现(基于PHPSpreadsheet/原生ZIP/OLE解析)

PHPSpreadsheet本身未提供提取非图片嵌入对象的直接API,但可以通过以下两种方式处理:

1. 处理XLSX格式(推荐)

XLSX本质是ZIP压缩包,嵌入的PDF会存放在xl/embeddings/目录下,可直接解压提取:

$zip = new ZipArchive();
if ($zip->open('target_file.xlsx') === true) {
    for ($i = 0; $i < $zip->numFiles; $i++) {
        $filePath = $zip->getNameIndex($i);
        if (str_starts_with($filePath, 'xl/embeddings/')) {
            $fileContent = $zip->getFromIndex($i);
            // 通过PDF文件头验证格式
            if (str_starts_with($fileContent, '%PDF-')) {
                $outputName = 'extracted_' . basename($filePath) . '.pdf';
                file_put_contents($outputName, $fileContent);
            }
        }
    }
    $zip->close();
}

2. 处理XLS格式

针对旧版XLS文件,需借助phpoffice/phpole库(可通过Composer安装)解析OLE容器:

require 'vendor/autoload.php';

$ole = new \PhpOffice\PhpOle\Ole();
$ole->open('target_file.xls');

foreach ($ole->getRoot() as $entry) {
    if ($entry instanceof \PhpOffice\PhpOle\OleEntry) {
        $content = $entry->getData();
        if (str_starts_with($content, '%PDF-')) {
            file_put_contents('extracted_' . $entry->getName() . '.pdf', $content);
        }
    }
}
$ole->close();

Python 实现

方案1:直接解压XLSX(跨平台通用)

利用XLSX的ZIP特性,直接提取嵌入文件并验证格式:

import zipfile

with zipfile.ZipFile('target_file.xlsx', 'r') as zip_ref:
    for file_info in zip_ref.infolist():
        if file_info.filename.startswith('xl/embeddings/'):
            content = zip_ref.read(file_info)
            if content.startswith(b'%PDF-'):
                output_name = f"extracted_{file_info.filename.split('/')[-1]}.pdf"
                with open(output_name, 'wb') as f:
                    f.write(content)

方案2:结合openpyxl获取对象信息

通过openpyxl定位嵌入对象,再提取对应文件:

from openpyxl import load_workbook
import zipfile

wb = load_workbook('target_file.xlsx')
ws = wb.active

for shape in ws.shapes:
    if shape.shapeType == 'Picture':
        continue  # 跳过已能提取的图片
    if hasattr(shape, 'ole'):
        embed_path = shape.ole.target
        with zipfile.ZipFile('target_file.xlsx', 'r') as zip_ref:
            content = zip_ref.read(embed_path)
            if content.startswith(b'%PDF-'):
                output_name = f"extracted_{embed_path.split('/')[-1]}.pdf"
                with open(output_name, 'wb') as f:
                    f.write(content)

内容的提问来源于stack exchange,提问作者Vijeet Deliwala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 17:52:35