Linux按行数分割Windows文本文件时保留换行符问题求助
Hey there! I see you're running into line break problems when splitting Windows-formatted TXT files using bash scripts—let's break down what's happening and fix it.
The Root Cause
Windows uses \r\n (CRLF) as the line ending, while Linux systems default to \n (LF). When you use tools like split or awk directly on a Windows TXT file, the \r (carriage return) character stays attached to the end of each line. When you view these split files, the \r can cause unexpected behavior: some tools interpret it as "jump back to the start of the line," making subsequent lines overwrite the previous ones, which looks like all lines are merged together (line1line2line3).
Solutions
Here are a few straightforward fixes, depending on whether you want to modify the original file or process it on-the-fly:
1. Convert the File First (Recommended)
First, convert the Windows line endings to Linux format to avoid future issues. If you have dos2unix installed (common on most Linux systems), run:
dos2unix LargeFile.txt
If dos2unix isn't available, use sed or awk to strip the \r characters:
# Using sed (modifies the file in-place) sed -i 's/\r$//' LargeFile.txt # Using awk (writes to a new file, leaves original intact) awk '{sub(/\r$/,""); print}' LargeFile.txt > FixedLargeFile.txt
Once converted, your original split or awk commands will work correctly:
# Split into 5000-line chunks, then add .txt extensions split -l 5000 LargeFile.txt SmallFile for file in SmallFile*; do mv "$file" "$file.txt"; done # Or use awk (with file handle cleanup) awk '{filename = "wrd." int((NR-1)/5000) ".txt"; print >> filename; close(filename)}' LargeFile.txt
2. Split Without Modifying the Original File
If you can't edit the original file, process it on-the-fly to strip \r before splitting:
- For
split:tr -d '\r' < LargeFile.txt | split -l 5000 - SmallFile # Add .txt extensions afterward for file in SmallFile*; do mv "$file" "$file.txt"; done - For
awk:awk '{ sub(/\r$/,""); # Strip the carriage return from the end of the line filename = "wrd." int((NR-1)/5000) ".txt"; print >> filename; close(filename); # Close the file to avoid too many open handles }' LargeFile.txt
Verify the Fix
To check if the line endings are correct, use cat -A to visualize invisible characters:
cat -A SmallFileaa.txt
You should see $ at the end of each line (indicating \n), instead of ^M$ (which is \r\n).
内容的提问来源于stack exchange,提问作者iamcj

