如何从编码输入中无损恢复数据及相关代码问题咨询
自制压缩格式还原问题分析与修复
核心问题
- 无损恢复失败,输出文件尺寸始终大于源文件
- 解压循环无法正确终止,存在数据遗漏
- 疑问:位操作
(x & 7)在还原中的作用,误以为它干扰数据
压缩代码片段
while (top > 0) { str_xweight += bitset<4>(top).to_string(); top >>= 4; // long long unsigned int u32t = pow(2,63); uint64_t u8t = top; top = 0; while ((u8t >> 3) > 0) { str_xweight += "1"; str_xweight += bitset<3>(u8t).to_string(); u8t >>= 3; } str_xweight += "0"; str_xweight += bitset<3>(u8t).to_string(); }
注:该代码会被多次调用,输出的累计数据前会添加
-$标记,但不会以该标记结尾。
解压代码片段
bool chk = false; for (i = 0; section_str.length() > 4 ; ) { if (!chk) full_comp += section_str.substr(0,4); else return full_comp; section_str = section_str.substr(4); while (section_str.front() == '1') { full_comp += section_str.substr(1,3); section_str = section_str.substr(4); cout << "." << flush; chk = true; } if (chk && section_str.front() == '0') { cout << "?" << flush; full_comp += section_str.substr(1,3); section_str = section_str.substr(4); chk = false; } }
压缩文件数据读取代码片段
uint64_t buf_cntA = 0, buf_cntB = 2, cntAB= 0, return_len = 0; string binary_return_str = ""; ofstream out {filename_str.c_str(), std::ios_base::out | std::ios_base::trunc }; for (size_t i = 0; (full_size > buf_cntB + 8) ; i++) { buf_cntA = buf.find("-$",buf_cntB); buf_cntB = buf.find("-$",buf_cntA); string buf_keep = buf.substr(buf_cntA+2,buf_cntB); buf_cntB += 2; if (buf_keep.length() == 0) { out << " " << flush; // binary_return_str.clear(); continue; } string return_str = retrieve(buf_keep); binary_return_str = bitsToBytes(return_str, return_len); out << binary_return_str << flush; binary_return_str.clear(); cout << "\rLeft: " << (full_size - buf_cntB) << " \r " << flush; } }
问题分析与修复方案
1. 解压循环终止错误
- 解压代码逻辑缺陷:
for循环条件section_str.length() > 4会直接丢弃长度≤4的剩余数据;且当chk = true时直接return full_comp,导致单个块未处理完就提前返回,丢失后续数据。- 修复:移除
else return full_comp分支,改为处理完当前块后继续循环;将循环条件改为!section_str.empty(),并在内部判断剩余长度是否足够处理下一段位数据。
- 修复:移除
- 数据读取循环终止条件过严:
full_size > buf_cntB +8的条件会忽略最后一段无-$标记的数据,因为最后一次buf_cntB会返回string::npos,导致这段数据完全没被处理。- 修复:单独处理最后一段:当
buf_cntB == string::npos时,截取从buf_cntA+2到缓冲区结尾的内容进行解压处理。
- 修复:单独处理最后一段:当
2. 输出文件过大问题
- 解压代码中
chk的逻辑错误,可能导致重复写入数据;另外bitsToBytes函数如果在补位时错误添加了多余的字节(比如未正确处理剩余不足8位的位流),也会导致文件体积膨胀。- 修复:严格匹配压缩与解压的位流结构:压缩时每个块的结构是
[4位初始值] + ([1前缀+3位数据]*) + [0前缀+3位数据],解压时需完全遵循该结构读取,避免重复或错误读取位段;同时检查bitsToBytes函数,确保仅在必要时补0位,且补位方向(高位补0/低位补0)与压缩时的字节转换逻辑一致。
- 修复:严格匹配压缩与解压的位流结构:压缩时每个块的结构是
3. (x & 7)的作用解析
7的二进制是0b111,x & 7的作用是精准提取x的低3位二进制数据,这和压缩代码中bitset<3>(u8t)的功能完全等价——都是用来拆分固定长度的位段。这个操作是位压缩的标准操作,不会干扰原有数据,是你的压缩逻辑中拆分u8t的核心步骤,完全正常。
内容的提问来源于stack exchange,提问作者Anthony Pulse
相关产品推荐
相关产品推荐

