You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现CSV文件每行固定为982个非逗号分隔字符?

Solution to Enforce 982 Non-Comma Characters Per CSV Line

Got it, let's break down how to solve this problem. The key goal is to process your CSV so every line has exactly 982 non-comma characters, pulling from subsequent lines if a line is too short, and continuing until all content is handled. Here's a step-by-step approach with code:

Core Approach

  • Read all content first: Instead of processing line-by-line in isolation, we'll read the entire file into a single stream of characters. This makes it easy to pull characters from "the next line" seamlessly, without getting stuck on original line breaks.
  • Track non-comma count: As we build each output line, we'll count how many non-comma characters we've added—this is the metric we care about hitting.
  • Build lines incrementally: Once we hit 982 non-comma characters, we finalize that line and start building the next one. Any remaining characters (including commas) carry over to the next line's buffer.

Python Implementation

def process_csv(input_file_path, output_file_path, target_non_comma=982):
    # Read the entire input file, stripping out newlines to create a continuous character stream
    with open(input_file_path, 'r', encoding='utf-8') as f:
        content = f.read().replace('\n', '').replace('\r', '')
    
    current_line = []
    non_comma_count = 0
    
    with open(output_file_path, 'w', encoding='utf-8') as out_f:
        for char in content:
            current_line.append(char)
            # Only increment count if the character isn't a comma
            if char != ',':
                non_comma_count += 1
            
            # Check if we've hit our target non-comma character count
            if non_comma_count == target_non_comma:
                out_f.write(''.join(current_line) + '\n')
                # Reset buffer and count for the next line
                current_line = []
                non_comma_count = 0
        
        # Handle any leftover content that doesn't make a full 982 non-comma characters
        if current_line:
            out_f.write(''.join(current_line) + '\n')

# Example usage
if __name__ == "__main__":
    # Replace with your actual input/output file paths
    process_csv("dna.csv", "processed_dna.csv")

How This Works

  1. Continuous content stream: By stripping newlines, we treat the entire file as one long string—so pulling characters from "the next line" is just grabbing the next character in the stream, no extra logic needed.
  2. Line building: We add every character to the current line buffer, but only count non-comma characters toward our 982 target.
  3. Finalizing lines: As soon as we hit 982 non-comma characters, we write the line to the output and reset our buffer.
  4. Leftover content: If the total number of non-comma characters isn't a perfect multiple of 982, the remaining content gets written as a final (short) line. If you need to loop back to the start of the file to fill this final line to 982, just let me know and I can adjust the code!

内容的提问来源于stack exchange,提问作者xion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:54:43