TypeScript读取XML转Base64后乱码问题排查
XML读取编码后解码乱码的排查方向(PlayWright+TypeScript+Cucumber项目)
在基于PlayWright + TypeScript + Cucumber构建的项目中,开发读取XML文件、格式化后编码为带填充Base64的功能时,已实现读取与编码逻辑,但解码后出现大量乱码。相同代码在另一仓库可正常运行,怀疑是当前仓库配置或依赖包版本问题,以下是具体排查方向:
具体实现细节
- XML文件内容
<?xml version="1.0" encoding="UTF-8" standalone="yes"?> <DataContent xmlns="http://esw.vlv.com/somecontent" xmlns:ns6="http://esw.vlv.com/somecontent" xmlns:ns5="http://esw.vlv.com/somecontent" xmlns:ns8="http://esw.vlv.com/somecontent" xmlns:ns7="http://esw.vlv.com/somecontent" xmlns:ns9="http://esw.vlv.com/somecontent" xmlns:ns11="http://somecontent.com/somecontent/1_1" xmlns:ns10="http://vcg.somecontent.com/somecontent" xmlns:ns4="http://esw.vlv.com/somecontent" schemaVersion="2.7"> <Status> <DataCompleted/> </Status> <Messages> <Message> <ReasonCode>33</ReasonCode> <ReasonDesc>Info:</ReasonDesc> <ReasonDesc>Val = 2</ReasonDesc> <ReasonDesc>BLE = 1234567890</ReasonDesc> <ReasonDesc>BLA = SOMETHING 12345</ReasonDesc> </Message> </Messages> <DataContentOutputs> <DataContentOutput> <Inputs> <Input> <InputVal schemaVersion="2.1" InputCode="BLAP2"> <ns5:Val> <ns4:Bool>true</ns4:Bool> </ns5:Val> </InputVal> </Input> </Inputs> <FailedInputs/> </DataContentOutput> </DataContentOutputs> </DataContent>
- 读取并格式化XML的代码
const stringFileContent= fs.readFileSync(filePathGoesHere, 'utf-8') const xmlFileContent = xmlFormat(stringFileContent, { indentation: ' ', collapseContent: true, lineSeparator: '\r\n' })
- Base64编码逻辑
const xmlEncoded = this.xmlEncodeBase64WithPadding(xmlFileContent) private xmlEncodeBase64WithPadding(content: any){ const encodedBytes = Buffer.from(content, 'utf-8').toString('base64'); let encodedStr = encodedBytes; // Add padding if necessary const paddingNeeded = encodedStr.length % 4; if (paddingNeeded) { encodedStr += '='.repeat(4 - paddingNeeded); } return encodedStr; }
- 乱码现象
解码后出现类似「熖煆V%Ǘ5&W7VLJ3У´FF6ö熖煦W÷'C」的不可识别字符,XML片段大量损坏。
排查方向
- 依赖与环境版本对比
- 对比两个仓库的
package.json,重点检查XML格式化库(如xml-formatter)的版本,不同版本可能在字符串处理、编码保留上有差异。 - 确认Node.js版本是否一致:Buffer的编码逻辑在不同Node版本中可能存在细微差异,尤其是UTF-8字符的处理。
- 对比两个仓库的
- 文件读取阶段验证
- 读取文件后直接打印
stringFileContent,确认原始内容是否已出现乱码,排除文件实际编码与读取指定编码(utf-8)不匹配的问题(比如文件实际是GBK编码)。 - 执行
Buffer.from(stringFileContent).toString('utf-8')二次转换后检查内容,验证读取时的编码是否被正确解析。
- 读取文件后直接打印
- XML格式化过程检查
- 打印格式化后的
xmlFileContent,对比原XML内容,确认是否引入了不可见字符、错误换行符,或修改了XML的UTF-8编码声明。 - 检查格式化库的配置参数,比如
lineSeparator: '\r\n'是否在当前环境下被错误处理为其他字符。
- 打印格式化后的
- Base64编码逻辑验证
- 移除手动添加padding的代码:
Buffer.toString('base64')会自动生成符合规范的带padding字符串,手动添加可能导致重复padding,引发解码错误。 - 复制
xmlFileContent到在线Base64工具编码后解码,若结果正常,说明当前环境下的Buffer处理存在异常。
- 移除手动添加padding的代码:
- 项目配置差异排查
- 对比两个仓库的
tsconfig.json,重点检查target、charset、module等配置,低版本target可能导致字符串处理逻辑差异。 - 检查是否存在Babel或其他转译工具,转译后的代码是否篡改了字符串或Buffer的原生处理逻辑。
- 排查项目中是否有全局polyfill对Buffer进行了重写,导致编码逻辑异常。
- 对比两个仓库的
内容的提问来源于stack exchange,提问作者Chewie
相关产品推荐
相关产品推荐

