You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从文本文件读取字节数组还原PDF时出现文件损坏问题

问题原因及解决方案

问题根源

PDF属于二进制文件,其字节序列包含大量不符合UTF-8编码规范的字节。当你用Encoding.UTF8.GetString(SourceBytes)把二进制字节转成UTF-8字符串时,所有无效的UTF-8字节都会被自动替换成�(替换字符)。后续再用Encoding.UTF8.GetBytes()转回字节数组时,这些被替换的字节已经无法还原成原始值,最终导致生成的PDF文件损坏。

正确解决方案:使用Base64编码

Base64是专门设计用来将二进制数据转换为可安全存储/传输的文本格式的编码方式,能保证二进制数据和文本之间的无损转换。修改后的代码如下:

using System.Text;

string SourceFile = @"C:\Files\originalfile.pdf";
string SerializedFile = @"C:\Files\serialized.txt";
string RevivedFile = @"C:\Files\revived.pdf";

SerializeFile(SourceFile, SerializedFile);
DeSerializeFile(SerializedFile, RevivedFile);

// 序列化:二进制转Base64文本
static void SerializeFile(string SourceFilePath, string SerializedFilePath)
{
    byte[] SourceBytes = File.ReadAllBytes(SourceFilePath);
    string base64String = Convert.ToBase64String(SourceBytes);
    File.WriteAllText(SerializedFilePath, base64String);
    Console.WriteLine("Serialized File created");
}

// 反序列化:Base64文本转二进制
static void DeSerializeFile(string SerializedFilePath, string RevivedFilePath)
{       
    string base64String = File.ReadAllText(SerializedFilePath);
    byte[] RevivedBytes = Convert.FromBase64String(base64String);
    File.WriteAllBytes(RevivedFilePath, RevivedBytes);
    Console.WriteLine("File Revived.");
}

代码说明

  • 序列化时:用Convert.ToBase64String()把PDF的字节数组转换成Base64格式的字符串,再写入文本文件。
  • 反序列化时:读取文本文件中的Base64字符串,用Convert.FromBase64String()还原成原始的字节数组,最后写入PDF文件。

这样处理就能保证原始二进制数据完全无损,还原后的PDF可以正常打开。

内容的提问来源于stack exchange,提问作者Bluemarble

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 16:20:26