C++反向写入文件时换行符数量翻倍的原因排查
问题:反向文件后换行符数量翻倍的原因及解决方法
我实现了一个仅处理ASCII字符的C++反向写入文件函数,使用MSVC 19.41.34123 x86编译器时,发现反向后的文件换行符数量翻倍:原文件有10个\n,反向后变为20个,两次反向后达40个。调试时发现原文件的每个\n在反向读取时会被识别为两个\n。
代码实现
void reverse_file(const std::filesystem::path& input_file, const std::filesystem::path& out_file) { std::ifstream ifs{ input_file }; ifs.exceptions(std::ifstream::badbit); std::ofstream ofs{out_file}; ofs.exceptions(std::ofstream::badbit | std::ofstream::failbit); // moves one to the left such that get does not already read at end ifs.seekg(-1, std::ios::end); auto end_point = ifs.tellg() ; for (; end_point >= 0; end_point -= 1) { ifs.seekg(end_point, std::ios::beg); char c = static_cast<char>(ifs.get()); if (ifs.eof()) return; else if (ifs.fail()) throw std::ifstream::failure("Invalid character "); ofs << c; } }
原测试文件
The quick brown fox jumps over the lazy dog. C++ is a powerful programming language for system-level development. Lorem ipsum dolor sit amet, consectetur adipiscing elit. 1234567890 - Numbers can be part of a text file too! How many lines can your program read in a single pass? Testing edge cases is essential for robust software. This line intentionally left blank. Whitespace matters: spaces, tabs, and newlines are characters too. A single character: X. Did your program handle this line correctly? End
反向结果
dnE ?yltcerroc enil siht eldnah margorp ruoy diD .X :retcarahc elgnis A .oot sretcarahc era senilwen dna ,sbat ,secaps :srettam ecapsetihW .knalb tfel yllanoitnetni enil sihT .erawtfos tsubor rof laitnesse si sesac egde gnitseT ?ssap elgnis a ni daer margorp ruoy nac senil ynam woH !oot elif txet a fo trap eb nac srebmuN - 0987654321 .tile gnicsipida rutetcesnoc ,tema tis rolod muspi meroL .tnempoleved level-metsys rof egaugnal gnimmargorp lufrewop a si ++C .god yzal eht revo spmuj xof nworb kciuq ehT
补充信息
使用count_newlines函数统计换行符,结果为:original:10、reversed_once:20、reversed_twice:40,调试时原文件开头读取的字符为'd'、'n'、'E'、'\n'、'\n'、'?'。
原因分析及解决方法
问题根源
问题出在文本模式下的文件读写逻辑:
- Windows系统中文本文件的换行是
\r\n(CRLF)组合,但在文本模式打开文件时,C++标准库会自动将磁盘上的\r\n转换为单个\n(LF)供程序读取; - 写入时,标准库又会将程序中的
\n转换回\r\n写入磁盘。
你的代码使用seekg直接操作文件指针,跳过了文本模式的转换逻辑。原文件中每个换行在磁盘上是两个字节(\r和\n),你逐字节读取时会把这两个字节当成独立字符,写入时每个\n又会被转换为\r\n,最终导致换行符数量翻倍。
解决方法
打开文件时使用二进制模式,禁用标准库的换行符自动转换,直接读写磁盘上的原始字节:
// 修改文件打开方式,添加std::ios::binary std::ifstream ifs{ input_file, std::ios::binary }; std::ofstream ofs{ out_file, std::ios::binary };
这样就能准确读取每个原始字节,反向操作后不会出现换行符数量翻倍的问题。
内容的提问来源于stack exchange,提问作者a_floating_point
相关产品推荐
相关产品推荐

