You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++反向写入文件时换行符数量翻倍的原因排查

问题:反向文件后换行符数量翻倍的原因及解决方法

我实现了一个仅处理ASCII字符的C++反向写入文件函数,使用MSVC 19.41.34123 x86编译器时,发现反向后的文件换行符数量翻倍:原文件有10个\n,反向后变为20个,两次反向后达40个。调试时发现原文件的每个\n在反向读取时会被识别为两个\n。

代码实现

void reverse_file(const std::filesystem::path& input_file, const std::filesystem::path& out_file)
{
    std::ifstream ifs{ input_file };
    ifs.exceptions(std::ifstream::badbit);
    std::ofstream ofs{out_file};
    ofs.exceptions(std::ofstream::badbit | std::ofstream::failbit);
    
    // moves one to the left such that get does not already read at end
    ifs.seekg(-1, std::ios::end); 
    auto end_point = ifs.tellg() ;

    for (; end_point >= 0; end_point -= 1)
    {
        ifs.seekg(end_point, std::ios::beg);
        char c = static_cast<char>(ifs.get());
        if (ifs.eof()) return;
        else if (ifs.fail()) throw std::ifstream::failure("Invalid character ");
        ofs << c;
    }
}

原测试文件

The quick brown fox jumps over the lazy dog.
C++ is a powerful programming language for system-level development.
Lorem ipsum dolor sit amet, consectetur adipiscing elit.
1234567890 - Numbers can be part of a text file too!
How many lines can your program read in a single pass?
Testing edge cases is essential for robust software.
This line intentionally left blank.
Whitespace matters: spaces, tabs, and newlines are characters too.
A single character: X.
Did your program handle this line correctly?
End

反向结果

dnE

?yltcerroc enil siht eldnah margorp ruoy diD

.X :retcarahc elgnis A

.oot sretcarahc era senilwen dna ,sbat ,secaps :srettam ecapsetihW

.knalb tfel yllanoitnetni enil sihT

.erawtfos tsubor rof laitnesse si sesac egde gnitseT

?ssap elgnis a ni daer margorp ruoy nac senil ynam woH

!oot elif txet a fo trap eb nac srebmuN - 0987654321

.tile gnicsipida rutetcesnoc ,tema tis rolod muspi meroL

.tnempoleved level-metsys rof egaugnal gnimmargorp lufrewop a si ++C

.god yzal eht revo spmuj xof nworb kciuq ehT

补充信息

使用count_newlines函数统计换行符,结果为:original:10、reversed_once:20、reversed_twice:40,调试时原文件开头读取的字符为'd'、'n'、'E'、'\n'、'\n'、'?'。


原因分析及解决方法

问题根源

问题出在文本模式下的文件读写逻辑:

  • Windows系统中文本文件的换行是\r\n(CRLF)组合,但在文本模式打开文件时,C++标准库会自动将磁盘上的\r\n转换为单个\n(LF)供程序读取;
  • 写入时,标准库又会将程序中的\n转换回\r\n写入磁盘。

你的代码使用seekg直接操作文件指针,跳过了文本模式的转换逻辑。原文件中每个换行在磁盘上是两个字节(\r和\n),你逐字节读取时会把这两个字节当成独立字符,写入时每个\n又会被转换为\r\n,最终导致换行符数量翻倍。

解决方法

打开文件时使用二进制模式,禁用标准库的换行符自动转换,直接读写磁盘上的原始字节:

// 修改文件打开方式,添加std::ios::binary
std::ifstream ifs{ input_file, std::ios::binary };
std::ofstream ofs{ out_file, std::ios::binary };

这样就能准确读取每个原始字节,反向操作后不会出现换行符数量翻倍的问题。


内容的提问来源于stack exchange,提问作者a_floating_point

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 05:14:54