You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Amazon Textract处理PDF时遇UnsupportedDocumentException错误求助

解决Amazon Textract "Request has unsupported document format" 400错误

问题概述

使用Amazon Textract分析PDF时收到400 Bad Request错误,提示"Request has unsupported document format"。已验证文件MIME类型为PDF、符合大小限制,且已编写PDF拆分逻辑但未调用。

解决方案

1. 直接使用S3对象路径而非本地文件字节

Textract支持直接引用S3上的文件,避免本地文件读写可能导致的字节损坏问题。修改processTextract方法:

protected function processTextract($s3FilePath)
{
    $textractClient = new TextractClient([
        'region' => env('AWS_DEFAULT_REGION'),
        'version' => 'latest',
        'credentials' => [
            'key' => env('AWS_ACCESS_KEY_ID'),
            'secret' => env('AWS_SECRET_ACCESS_KEY'),
        ],
        'use_path_style_endpoint' => env('AWS_USE_PATH_STYLE_ENDPOINT'),
    ]);

    try {
        $result = $textractClient->analyzeDocument([
            'Document' => [
                'S3Object' => [
                    'Bucket' => env('AWS_BUCKET'),
                    'Name' => $s3FilePath
                ]
            ],
            'FeatureTypes' => ['TABLES', 'FORMS'],
        ]);

        return $result;
    } catch (AwsException $e) {
        return $e->getMessage();
    }
}

2. 启用PDF拆分逻辑

Textract的AnalyzeDocument接口单PDF最多支持11页,必须调用已编写的splitPdf方法拆分文件后逐个处理。修改uploadFile方法:

public function uploadFile(Request $request)
{
    $request->validate([
        'file' => 'required|mimes:pdf|max:2048',
        'validity' => 'required|date',
        'name' => 'required|string',
        'logo' => 'nullable|image|mimes:jpeg,png,jpg,gif,svg|max:2048',
    ]);

    // Upload do arquivo PDF
    $file = $request->file('file');
    $path = $file->store('pdf_tabelas_precos', 's3');

    // URL do arquivo PDF armazenado no S3
    $urlFile = Storage::disk('s3')->url($path);

    // Upload do logo, se fornecido
    $logoPath = null;
    $urlLogo = '';
    if ($request->hasFile('logo')) {
        $logo = $request->file('logo');
        $logoPath = $logo->store('logos', 's3');
        $urlLogo = Storage::disk('s3')->url($logoPath);
    }

    // 拆分PDF并处理每个子文件
    list($splitFiles, $downloadLinks) = $this->splitPdf($path);
    $allExtractedData = [];

    foreach ($splitFiles as $splitFile) {
        $extracted = $this->processTextractLocalFile($splitFile);
        $allExtractedData[] = $extracted;
        unlink($splitFile); // 清理本地拆分文件
    }

    // Salvar detalhes no banco de dados
    $operator = new Operator();
    $operator->name = $request->input('name');
    $operator->validity = $request->input('validity');
    $operator->url = $urlLogo;
    $operator->url_file = $urlFile;
    $operator->extracted_data = json_encode($allExtractedData);
    $operator->save();

    // Retornar links de download dos arquivos divididos
    return back()->with('success', 'Arquivo enviado e processado com sucesso!')
        ->with('file', $urlFile)
        ->with('downloadLinks', $downloadLinks);
}

// 新增处理本地拆分文件的方法
protected function processTextractLocalFile($localFilePath)
{
    $textractClient = new TextractClient([
        'region' => env('AWS_DEFAULT_REGION'),
        'version' => 'latest',
        'credentials' => [
            'key' => env('AWS_ACCESS_KEY_ID'),
            'secret' => env('AWS_SECRET_ACCESS_KEY'),
        ],
        'use_path_style_endpoint' => env('AWS_USE_PATH_STYLE_ENDPOINT'),
    ]);

    try {
        $result = $textractClient->analyzeDocument([
            'Document' => [
                'Bytes' => file_get_contents($localFilePath),
            ],
            'FeatureTypes' => ['TABLES', 'FORMS'],
        ]);

        return $result;
    } catch (AwsException $e) {
        return $e->getMessage();
    }
}

3. 验证PDF文件完整性

  • 检查上传的PDF是否为可搜索PDF(非扫描件图片PDF):若为扫描件,需使用Textract的DetectDocumentText接口
  • 验证本地临时文件的字节完整性,添加哈希校验:
// 在原processTextract方法中添加哈希验证
$fileContent = Storage::disk('s3')->get($s3FilePath);
$localPath = tempnam(sys_get_temp_dir(), 'pdf_') . '.pdf';
file_put_contents($localPath, $fileContent);

// 验证哈希一致性
$s3Hash = md5($fileContent);
$localHash = md5_file($localPath);
if ($s3Hash !== $localHash) {
    unlink($localPath);
    throw new \Exception('Arquivo local corrompido durante o download do S3.');
}

关键注意事项

  • Textract的AnalyzeDocument仅支持文本可搜索的PDF,扫描件需用DetectDocumentText接口处理
  • 单PDF文件页数限制为11,必须拆分后处理
  • 直接使用S3对象路径能避免本地文件读写的字节错误

内容的提问来源于stack exchange,提问作者Marcello Pato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 04:00:54