You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用express-fileupload上传.docx后Node.js无法读取,显示乱码

解决Express上传docx后读取乱码的问题

嘿,我来帮你搞定这个docx读取乱码的问题~

你遇到的核心问题是:docx不是纯文本文件,它本质是一个包含XML结构的压缩zip包,直接用utf8编码去读取二进制的docx文件,肯定会输出乱码。

解决方案步骤:

  1. 安装专门解析docx的库
    推荐用mammoth(轻量易用,适合快速提取纯文本),先安装依赖:

    npm install mammoth
    
  2. 修改代码,以二进制方式读取文件并解析
    上传环节没问题,但读取时要放弃utf8编码,读取原始二进制Buffer,再用mammoth解析内容:

    const mammoth = require('mammoth');
    const fs = require('fs');
    const path = require('path');
    
    app.post('/upload', (req, res, next) => {
      let file = req.files.file;
      const filePath = path.join(__dirname, 'public', req.body.filename);
      
      file.mv(filePath, function(err) {
        if (err) {
          return res.status(500).send(err);
        }
        
        // 以二进制方式读取文件(不指定编码,默认返回Buffer)
        fs.readFile(filePath, (err, buffer) => {
          if (err) {
            return res.status(500).send('读取文件失败:' + err.message);
          }
          
          // 使用mammoth提取docx中的纯文本
          mammoth.extractRawText({ buffer: buffer })
            .then(result => {
              const text = result.value; // 解析出的纯文本内容
              const parseTips = result.messages; // 解析过程中的提示(比如格式警告)
              
              console.log('解析后的文本:', text);
              res.send({ content: text, tips: parseTips });
            })
            .catch(parseErr => {
              res.status(500).send('解析docx失败:' + parseErr.message);
            });
        });
      });
    });
    

为什么这样能解决问题?

  • 当你调用fs.readFile时不指定编码,Node.js会返回原始的二进制Buffer,这才是docx文件的正确读取方式。
  • mammoth库会自动解析docx的内部压缩结构,提取出其中的文本内容,而不是把二进制数据强行转成utf8文本导致乱码。

如果之后需要读取docx的样式、表格、图片等更复杂的内容,可以试试docx库,不过它的API相对复杂一些,适合需要精细操作docx的场景。

内容的提问来源于stack exchange,提问作者m-ketan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:18:40