You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

react-pdf大文件字符计数返回0,小文件正常,求解决方案

解决react-pdf大文件字符计数返回0的问题

问题原因分析

你的代码在小文件场景正常但大文件返回0,大概率是以下几个问题导致:

  1. 多余的异步等待:content.items.reduce是同步操作,不需要加await,多余的await会让大文件处理逻辑出现未预期的阻塞或错误。
  2. 大文件加载异常未定位:react-pdf处理大文件时,单页加载可能因超时、内存不足触发catch分支,但原代码的错误日志未输出详细信息,无法定位具体问题。
  3. 串行处理效率低下:大文件页数多,串行循环加载页面耗时过长,可能触发浏览器异步任务超时。

修复后的代码实现

1. 修正countCharacters函数

去掉reduce前的await,改用并行处理页面提升效率,并补充详细错误日志:

export const countCharacters = async (pdf) => {
  try {
    let totalCharacters = 0;
    // 并行处理所有页面,提升大文件处理速度
    const pagePromises = Array.from({ length: pdf.numPages }, (_, i) => i + 1).map(async (pageNum) => {
      const page = await pdf.getPage(pageNum);
      const content = await page.getTextContent();
      return content.items.reduce((count, item) => count + item.str.length, 0);
    });
    
    const pageCounts = await Promise.all(pagePromises);
    totalCharacters = pageCounts.reduce((sum, count) => sum + count, 0);
    
    console.log("Total Characters:", totalCharacters);
    return totalCharacters;
  } catch (error) {
    // 打印详细错误信息,定位大文件问题
    console.error("Error while counting characters:", error.message, error.stack);
    return 0;
  }
};

2. 优化onFileLoad函数

避免async/await与.then混合使用,保持代码逻辑一致性:

const onFileLoad = async (file) => {
  try {
    const res = await countCharacters(file);
    const tokens = res / 4;
    if (tokens > 70000) {
      console.log("limit exceeds");
    } else {
      console.log("limit does not exceed");
    }
  } catch (err) {
    console.log(err);
  }
};

3. 调整react-pdf加载配置

增加超时设置并配置worker,避免大文件加载失败:

import { Document, pdfjs } from 'react-pdf';

// 配置pdfjs worker,规避默认加载问题
pdfjs.GlobalWorkerOptions.workerSrc = `//cdnjs.cloudflare.com/ajax/libs/pdf.js/${pdfjs.version}/pdf.worker.min.js`;

// 使用组件时添加超时配置
<Document 
  file={selectedFile} 
  onLoadSuccess={onFileLoad} 
  className={'hidden'}
  options={{ timeout: 60000 }} // 设为60秒超时,可根据文件大小调整
/>

其他PDF字符计数实现方法

1. 直接使用pdfjs-dist(react-pdf底层依赖)

跳过react-pdf的组件封装,直接调用底层库处理,灵活性更高:

import * as pdfjsLib from 'pdfjs-dist';
import 'pdfjs-dist/build/pdf.worker.entry';

pdfjsLib.GlobalWorkerOptions.workerSrc = 'pdfjs-dist/build/pdf.worker.entry.js';

export const countCharsWithPdfjs = async (file) => {
  const arrayBuffer = await file.arrayBuffer();
  const pdf = await pdfjsLib.getDocument({ data: arrayBuffer }).promise;
  let total = 0;
  for (let i = 1; i <= pdf.numPages; i++) {
    const page = await pdf.getPage(i);
    const content = await page.getTextContent();
    total += content.items.reduce((sum, item) => sum + item.str.length, 0);
  }
  return total;
};

2. 后端处理(推荐大文件场景)

前端处理大文件易受内存和性能限制,可将文件传到后端计算:

  • Node.js后端:使用pdf-parse库
    const pdfParse = require('pdf-parse');
    const fs = require('fs');
    
    const countCharsBackend = async (filePath) => {
      const dataBuffer = fs.readFileSync(filePath);
      const pdfData = await pdfParse(dataBuffer);
      return pdfData.text.length;
    };
    
  • Python后端:使用pdfplumber(文本提取准确性更高)
    import pdfplumber
    
    def count_chars_python(file_path):
        total = 0
        with pdfplumber.open(file_path) as pdf:
            for page in pdf.pages:
                total += len(page.extract_text())
        return total
    

内容的提问来源于stack exchange,提问作者Zahid Chandio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 19:22:37