You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自动化测试:使用WebdriverIO校验PDF内容

WebdriverIO + Node.js 实现PDF与网页/API数据校验方案

核心思路

WebdriverIO本身不提供PDF解析能力,需结合Node.js生态的PDF处理工具,分三步完成:提取PDF内容、获取网页/API数据、结构化对比校验。


步骤1:安装依赖

需要用到PDF解析库、HTTP请求库和断言库:

npm install pdf-parse axios chai --save-dev

步骤2:配置WebdriverIO下载目录(针对网页生成的PDF)

如果PDF是通过网页下载获取,需在wdio.conf.js中配置浏览器下载路径,避免弹窗干扰:

exports.config = {
  // 其他配置项...
  capabilities: [{
    browserName: 'chrome',
    'goog:chromeOptions': {
      prefs: {
        'download.default_directory': require('path').resolve(__dirname, './downloads'),
        'download.prompt_for_download': false,
        'download.directory_upgrade': true,
        'plugins.always_open_pdf_externally': true // 禁止浏览器内置PDF预览
      }
    }
  }]
}

步骤3:编写测试用例

const fs = require('fs');
const path = require('path');
const pdfParse = require('pdf-parse');
const axios = require('axios');
const { expect } = require('chai');

// 辅助函数:获取下载目录中最新的PDF文件
const getLatestPdf = () => {
  const downloadDir = path.resolve(__dirname, './downloads');
  const files = fs.readdirSync(downloadDir)
    .filter(file => file.endsWith('.pdf'))
    .map(file => ({
      name: file,
      time: fs.statSync(path.join(downloadDir, file)).mtime.getTime()
    }))
    .sort((a, b) => b.time - a.time);
  return files.length ? path.join(downloadDir, files[0].name) : null;
};

describe('PDF数据一致性校验', () => {
  let pdfExtractedData;
  let webPageData;
  let apiResponseData;

  before(async () => {
    // 1. 触发PDF下载
    await browser.url('/target-page-url');
    await $('#pdf-download-button').click();
    
    // 等待下载完成(用文件存在检查替代固定延时,提升稳定性)
    await browser.waitUntil(async () => {
      return getLatestPdf() !== null;
    }, { timeout: 10000, timeoutMsg: 'PDF下载超时' });

    // 解析PDF内容
    const pdfPath = getLatestPdf();
    const pdfBuffer = fs.readFileSync(pdfPath);
    const pdfText = (await pdfParse(pdfBuffer)).text;

    // 从PDF文本提取关键数据(根据实际格式调整正则)
    pdfExtractedData = {
      orderId: pdfText.match(/订单编号:(\S+)/)[1],
      totalAmount: pdfText.match(/应付金额:(\d+\.\d+)/)[1],
      customerName: pdfText.match(/客户姓名:(\S+)/)[1]
    };

    // 2. 获取网页展示数据
    webPageData = {
      orderId: await $('#order-id').getText(),
      totalAmount: await $('#total-amount').getText().replace('¥', ''),
      customerName: await $('#customer-name').getText()
    };

    // 3. 获取API返回数据
    const apiRes = await axios.get('/api/order-info');
    apiResponseData = {
      orderId: apiRes.data.orderId,
      totalAmount: apiRes.data.totalAmount.toString(),
      customerName: apiRes.data.customerName
    };
  });

  it('PDF数据与网页数据完全匹配', () => {
    expect(pdfExtractedData.orderId).to.equal(webPageData.orderId);
    expect(pdfExtractedData.totalAmount).to.equal(webPageData.totalAmount);
    expect(pdfExtractedData.customerName).to.equal(webPageData.customerName);
  });

  it('PDF数据与API返回数据完全匹配', () => {
    expect(pdfExtractedData.orderId).to.equal(apiResponseData.orderId);
    expect(pdfExtractedData.totalAmount).to.equal(apiResponseData.totalAmount);
    expect(pdfExtractedData.customerName).to.equal(apiResponseData.customerName);
  });
});

注意事项

  • 如果PDF是扫描件(图片格式),需替换pdf-parse为OCR库(如tesseract.js)进行文本识别
  • 正则表达式需根据实际PDF文本格式调整,确保准确提取目标数据
  • 下载等待逻辑优先用文件存在检查,避免固定延时导致的不稳定

内容的提问来源于stack exchange,提问作者Jay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 18:39:42