如何通过Nightwatch/Node.js+Cucumber自动化验证下载PDF的内容?
验证下载PDF内容的可行方案(Nightwatch/Node.js + Cucumber)
我刚好有过类似的场景经验,给你分享几个实用的方法,完全适配你的技术栈:
一、文本内容验证(最常用)
用pdf-parse这个轻量的Node.js库就能搞定,它能直接提取PDF里的文本内容,然后和预期值做断言。
步骤:
- 安装依赖:
npm install pdf-parse --save-dev
- 在Cucumber的步骤定义里编写逻辑:
const fs = require('fs'); const pdfParse = require('pdf-parse'); const { Given, When, Then } = require('@cucumber/cucumber'); Then('the downloaded PDF should contain text {string}', async function(expectedText) { // 替换成你的PDF下载路径,注意Nightwatch要配置好默认下载目录 const pdfPath = './downloads/generated.pdf'; // 等待文件下载完成(可以自己封装一个等待文件存在的工具函数) await waitForFileExists(pdfPath); const pdfBuffer = fs.readFileSync(pdfPath); const data = await pdfParse(pdfBuffer); // 用Nightwatch内置断言或者Cucumber断言库验证 this.assert.include(data.text, expectedText); }); // 辅助工具:等待文件存在 const waitForFileExists = (filePath, timeout = 10000) => { return new Promise((resolve, reject) => { const checkInterval = setInterval(() => { if (fs.existsSync(filePath)) { clearInterval(checkInterval); resolve(); } timeout -= 100; if (timeout <= 0) { clearInterval(checkInterval); reject(new Error(`File ${filePath} not found within timeout`)); } }, 100); }); };
注意:如果PDF是加密的,pdf-parse支持传入解密密码的配置项,只需在调用时添加{password: 'your-password'}参数即可。
二、图片内容验证
验证PDF里的图片有两种常见场景:
1. 对比图片像素一致性
可以用pdf-poppler把PDF页面转成图片,再用pixelmatch做像素对比,适合验证图片排版、样式是否符合预期:
- 安装依赖:
npm install pdf-poppler pixelmatch pngjs --save-dev
- 步骤逻辑示例:
const pdfPoppler = require('pdf-poppler'); const pixelmatch = require('pixelmatch'); const { PNG } = require('pngjs'); const fs = require('fs'); const { Then } = require('@cucumber/cucumber'); Then('the first page of PDF should match the reference image', async function() { const pdfPath = './downloads/generated.pdf'; const outputImgPath = './temp/pdf-page-1.png'; const referenceImgPath = './test/references/pdf-reference.png'; await waitForFileExists(pdfPath); // 把PDF第一页转成PNG格式 await pdfPoppler.convert(pdfPath, { format: 'png', page: 1, out_dir: './temp' }); // 读取两张图片做像素对比 const generatedImg = PNG.sync.read(fs.readFileSync(outputImgPath)); const referenceImg = PNG.sync.read(fs.readFileSync(referenceImgPath)); const diffImg = new PNG({ width: generatedImg.width, height: generatedImg.height }); // 计算像素差异,threshold设置容错率(0-1,值越大容错越高) const mismatchedPixels = pixelmatch(generatedImg.data, referenceImg.data, diffImg.data, generatedImg.width, generatedImg.height, { threshold: 0.1 }); // 断言像素差异在允许范围内 this.assert.ok(mismatchedPixels < 100, `PDF page has ${mismatchedPixels} mismatched pixels, exceeds allowed limit`); });
2. 验证图片中的文本(OCR方式)
如果需要验证图片里的文字内容(比如PDF里的扫描件图片),可以用tesseract.js做OCR识别:
- 安装依赖:
npm install tesseract.js --save-dev
- 逻辑示例:
const Tesseract = require('tesseract.js'); const pdfPoppler = require('pdf-poppler'); const { Then } = require('@cucumber/cucumber'); Then('the image in PDF page {int} should contain text {string}', async function(pageNum, expectedText) { const pdfPath = './downloads/generated.pdf'; const imgPath = `./temp/pdf-page-${pageNum}.png`; await waitForFileExists(pdfPath); // 转PDF页面为图片 await pdfPoppler.convert(pdfPath, { format: 'png', page: pageNum, out_dir: './temp' }); // OCR识别图片文本 const { data: { text } } = await Tesseract.recognize(imgPath, 'eng'); // 断言文本存在 this.assert.include(text.trim(), expectedText); });
三、Nightwatch配置注意事项
要确保能正确获取到下载的PDF,需要在Nightwatch配置里设置固定下载目录,避免浏览器弹窗干扰:
// nightwatch.conf.js module.exports = { test_settings: { default: { desiredCapabilities: { browserName: 'chrome', 'goog:chromeOptions': { prefs: { 'download.default_directory': require('path').resolve(__dirname, './downloads'), 'download.prompt_for_download': false, 'download.directory_upgrade': true, 'plugins.always_open_pdf_externally': true // 禁止浏览器直接打开PDF,强制下载 } } } } } };
内容的提问来源于stack exchange,提问作者subhajit chakraborty
相关产品推荐
相关产品推荐

