如何实现CSV文件每行固定为982个非逗号分隔字符?
Solution to Enforce 982 Non-Comma Characters Per CSV Line
Got it, let's break down how to solve this problem. The key goal is to process your CSV so every line has exactly 982 non-comma characters, pulling from subsequent lines if a line is too short, and continuing until all content is handled. Here's a step-by-step approach with code:
Core Approach
- Read all content first: Instead of processing line-by-line in isolation, we'll read the entire file into a single stream of characters. This makes it easy to pull characters from "the next line" seamlessly, without getting stuck on original line breaks.
- Track non-comma count: As we build each output line, we'll count how many non-comma characters we've added—this is the metric we care about hitting.
- Build lines incrementally: Once we hit 982 non-comma characters, we finalize that line and start building the next one. Any remaining characters (including commas) carry over to the next line's buffer.
Python Implementation
def process_csv(input_file_path, output_file_path, target_non_comma=982): # Read the entire input file, stripping out newlines to create a continuous character stream with open(input_file_path, 'r', encoding='utf-8') as f: content = f.read().replace('\n', '').replace('\r', '') current_line = [] non_comma_count = 0 with open(output_file_path, 'w', encoding='utf-8') as out_f: for char in content: current_line.append(char) # Only increment count if the character isn't a comma if char != ',': non_comma_count += 1 # Check if we've hit our target non-comma character count if non_comma_count == target_non_comma: out_f.write(''.join(current_line) + '\n') # Reset buffer and count for the next line current_line = [] non_comma_count = 0 # Handle any leftover content that doesn't make a full 982 non-comma characters if current_line: out_f.write(''.join(current_line) + '\n') # Example usage if __name__ == "__main__": # Replace with your actual input/output file paths process_csv("dna.csv", "processed_dna.csv")
How This Works
- Continuous content stream: By stripping newlines, we treat the entire file as one long string—so pulling characters from "the next line" is just grabbing the next character in the stream, no extra logic needed.
- Line building: We add every character to the current line buffer, but only count non-comma characters toward our 982 target.
- Finalizing lines: As soon as we hit 982 non-comma characters, we write the line to the output and reset our buffer.
- Leftover content: If the total number of non-comma characters isn't a perfect multiple of 982, the remaining content gets written as a final (short) line. If you need to loop back to the start of the file to fill this final line to 982, just let me know and I can adjust the code!
内容的提问来源于stack exchange,提问作者xion
相关产品推荐
相关产品推荐

