You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CI环境编码转换测试失败但本地正常的原因排查

问题:ISO-8859-1转换函数的Jest测试在CI环境失败但本地正常

转换函数代码

用于将文件内容转换为ISO-8859-1编码的TypeScript函数:

export const convertCharsetToIso8859 = async (
    file: Blob | File,
): Promise<Blob> => {
    type CHARSET = 'iso-8859-1' | 'UTF-8' | 'UTF-16';

    const readFileAsText = (file: Blob | File, encoding: CHARSET) =>
        new Promise<string>((resolve, reject) => {
            const reader = new FileReader();
            reader.onload = () => resolve(reader.result as string);
            reader.onerror = reject;
            reader.readAsText(file, encoding);
        });

    const encoding: CHARSET = 'iso-8859-1';
    const content = await readFileAsText(file, encoding);

    return new Blob([content], { type: `text/csv;charset=${encoding}` });
};

Jest单元测试代码

对应的单元测试逻辑:

describe('convertCharsetToIso8859', () => {
    it('should handle ISO-8859-1 charset correctly', async () => {
        const encoding = 'iso-8859-1';
        const mockFileContent = 
            'test content mit Umlaute: ä, ö, ü und rumänischen Zeichen: ș, ț, ă';
        const blob = new Blob([mockFileContent], {
            type: `text/csv;charset=${encoding}`,
        });
        const expectedText = new TextDecoder(encoding).decode(
            new TextEncoder().encode(mockFileContent),
        );

        const resultBlob = await convertCharsetToIso8859(blob);

        const resultText = await new Promise((resolve, reject) => {
            const reader = new FileReader();
            reader.onload = () => resolve(reader.result);
            reader.onerror = reject;
            reader.readAsText(resultBlob);
        });

        expect(resultText).toBe(expectedText);
        expect(resultBlob.type).toBe(`text/csv;charset=${encoding}`);
    });
});

CI环境测试失败报错

使用public.ecr.aws/docker/library/node:20镜像的GitLab CI环境中,测试抛出如下断言错误:

● convertCharsetToIso8859 › should handle ISO-8859-1 charset correctly
    expect(received).toBe(expected) // Object.is equality
    Expected: "test content mit Umlaute: ä, ö, ü und rumänischen Zeichen: È, È, Ä"
    Received: "test content mit Umlaute: ä, ö, ü und rumänischen Zeichen: ș, ț, ă"
      23 |         });
      24 |
    > 25 |         expect(resultText).toBe(expectedText);
         |                            ^
      26 |         expect(resultBlob.type).toBe(`text/csv;charset=${encoding}`);
      27 |     });
      28 | });
      at Object.toBe (src/tools/utils/helpers.test.ts:25:28)

已尝试的排查方案

  • 使用iconv替代TextDecoder/TextEncoder,出现更多错误
  • 在本地相同Docker环境运行测试,结果正常
  • 设置locale为en_us.UTF-8,无效果
  • 检测UTF8 BOM,测试字符串无BOM
  • 更换为node:22镜像,无效果

可能的原因分析

  1. 非标准字符的编码 fallback 逻辑差异:ISO-8859-1编码本身不支持罗马尼亚语的ș、ț、ă字符,这些字符属于ISO-8859-2或UTF-8编码集。当FileReader.readAsText以iso-8859-1读取包含这些字符的内容时,不同环境对无法编码字符的 fallback 处理逻辑不同——本地环境可能将其替换为单字节占位符,而CI环境保留了UTF-8的多字节表示,导致最终字符串不一致。
  2. Jest环境的Blob/FileReader模拟差异:本地测试可能使用浏览器端的FileReader实现(依赖jsdom),而CI环境的Jest可能使用Node.js端的polyfill实现,两者对非标准编码字符的解析逻辑存在细微差别,导致测试结果不一致。
  3. 代码文件的编码读取差异:测试代码中的mockFileContent字符串在本地和CI环境被Node.js读取时的编码设置可能不同。尽管代码以UTF-8保存,但CI环境的Node.js可能以其他编码解析代码文件,导致特殊字符被错误解析,进而影响后续的编码转换逻辑。
  4. Docker容器Locale的深层配置问题:虽然手动设置了en_us.UTF-8,但容器内部的Locale可能并未真正生效,或者Node.js运行时未正确加载Locale配置,导致字符处理逻辑异常。

内容的提问来源于stack exchange,提问作者mles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 12:07:37