You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Nightwatch/Node.js+Cucumber自动化验证下载PDF的内容?

验证下载PDF内容的可行方案(Nightwatch/Node.js + Cucumber)

我刚好有过类似的场景经验,给你分享几个实用的方法,完全适配你的技术栈:

一、文本内容验证(最常用)

用pdf-parse这个轻量的Node.js库就能搞定,它能直接提取PDF里的文本内容,然后和预期值做断言。

步骤:

  1. 安装依赖:
npm install pdf-parse --save-dev
  1. 在Cucumber的步骤定义里编写逻辑:
const fs = require('fs');
const pdfParse = require('pdf-parse');
const { Given, When, Then } = require('@cucumber/cucumber');

Then('the downloaded PDF should contain text {string}', async function(expectedText) {
  // 替换成你的PDF下载路径,注意Nightwatch要配置好默认下载目录
  const pdfPath = './downloads/generated.pdf';
  // 等待文件下载完成(可以自己封装一个等待文件存在的工具函数)
  await waitForFileExists(pdfPath);
  
  const pdfBuffer = fs.readFileSync(pdfPath);
  const data = await pdfParse(pdfBuffer);
  
  // 用Nightwatch内置断言或者Cucumber断言库验证
  this.assert.include(data.text, expectedText);
});

// 辅助工具:等待文件存在
const waitForFileExists = (filePath, timeout = 10000) => {
  return new Promise((resolve, reject) => {
    const checkInterval = setInterval(() => {
      if (fs.existsSync(filePath)) {
        clearInterval(checkInterval);
        resolve();
      }
      timeout -= 100;
      if (timeout <= 0) {
        clearInterval(checkInterval);
        reject(new Error(`File ${filePath} not found within timeout`));
      }
    }, 100);
  });
};

注意:如果PDF是加密的,pdf-parse支持传入解密密码的配置项,只需在调用时添加{password: 'your-password'}参数即可。

二、图片内容验证

验证PDF里的图片有两种常见场景:

1. 对比图片像素一致性

可以用pdf-poppler把PDF页面转成图片,再用pixelmatch做像素对比,适合验证图片排版、样式是否符合预期:

  • 安装依赖:
npm install pdf-poppler pixelmatch pngjs --save-dev
  • 步骤逻辑示例:
const pdfPoppler = require('pdf-poppler');
const pixelmatch = require('pixelmatch');
const { PNG } = require('pngjs');
const fs = require('fs');
const { Then } = require('@cucumber/cucumber');

Then('the first page of PDF should match the reference image', async function() {
  const pdfPath = './downloads/generated.pdf';
  const outputImgPath = './temp/pdf-page-1.png';
  const referenceImgPath = './test/references/pdf-reference.png';
  
  await waitForFileExists(pdfPath);
  
  // 把PDF第一页转成PNG格式
  await pdfPoppler.convert(pdfPath, {
    format: 'png',
    page: 1,
    out_dir: './temp'
  });
  
  // 读取两张图片做像素对比
  const generatedImg = PNG.sync.read(fs.readFileSync(outputImgPath));
  const referenceImg = PNG.sync.read(fs.readFileSync(referenceImgPath));
  const diffImg = new PNG({ width: generatedImg.width, height: generatedImg.height });
  
  // 计算像素差异,threshold设置容错率(0-1,值越大容错越高)
  const mismatchedPixels = pixelmatch(generatedImg.data, referenceImg.data, diffImg.data, generatedImg.width, generatedImg.height, { threshold: 0.1 });
  
  // 断言像素差异在允许范围内
  this.assert.ok(mismatchedPixels < 100, `PDF page has ${mismatchedPixels} mismatched pixels, exceeds allowed limit`);
});

2. 验证图片中的文本(OCR方式)

如果需要验证图片里的文字内容(比如PDF里的扫描件图片),可以用tesseract.js做OCR识别:

  • 安装依赖:
npm install tesseract.js --save-dev
  • 逻辑示例:
const Tesseract = require('tesseract.js');
const pdfPoppler = require('pdf-poppler');
const { Then } = require('@cucumber/cucumber');

Then('the image in PDF page {int} should contain text {string}', async function(pageNum, expectedText) {
  const pdfPath = './downloads/generated.pdf';
  const imgPath = `./temp/pdf-page-${pageNum}.png`;
  
  await waitForFileExists(pdfPath);
  
  // 转PDF页面为图片
  await pdfPoppler.convert(pdfPath, {
    format: 'png',
    page: pageNum,
    out_dir: './temp'
  });
  
  // OCR识别图片文本
  const { data: { text } } = await Tesseract.recognize(imgPath, 'eng');
  
  // 断言文本存在
  this.assert.include(text.trim(), expectedText);
});

三、Nightwatch配置注意事项

要确保能正确获取到下载的PDF,需要在Nightwatch配置里设置固定下载目录,避免浏览器弹窗干扰:

// nightwatch.conf.js
module.exports = {
  test_settings: {
    default: {
      desiredCapabilities: {
        browserName: 'chrome',
        'goog:chromeOptions': {
          prefs: {
            'download.default_directory': require('path').resolve(__dirname, './downloads'),
            'download.prompt_for_download': false,
            'download.directory_upgrade': true,
            'plugins.always_open_pdf_externally': true // 禁止浏览器直接打开PDF,强制下载
          }
        }
      }
    }
  }
};

内容的提问来源于stack exchange,提问作者subhajit chakraborty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:23:53