You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python按指定格式将数据集写入.txt文件?

Efficiently Generate Formatted TXT File for Document-Word Counts

It sounds like you're dealing with a large document-word matrix and need a fast, reliable way to output it in that specific text format. The key issues with np.savetxt you're facing are likely either inefficient handling of huge arrays or incorrect format specification. Let's fix that with two optimized approaches depending on whether your docword data is a NumPy array or a regular Python list.


If docword is a NumPy Array

NumPy's savetxt can be efficient if you process the data in chunks (to avoid overwhelming memory) and explicitly set the format to match your requirements:

import numpy as np

# Your predefined values
D = 68601
W = 1500
NNZ = 205806000
docword = np.array([[0, 170, 4], [0, 1856, 4], ..., [68601, 0, 0]])  # Your large array

# Open the output file
with open('docword_output.txt', 'w') as f:
    # Write the metadata lines first
    f.write(f"{D}\n")
    f.write(f"{W}\n")
    f.write(f"{NNZ}\n")
    
    # Process data in chunks to minimize memory usage and speed up writes
    chunk_size = 1_000_000  # Adjust based on your available memory
    for start_idx in range(0, len(docword), chunk_size):
        end_idx = min(start_idx + chunk_size, len(docword))
        chunk = docword[start_idx:end_idx]
        # Write chunk with integer format, space-separated
        np.savetxt(f, chunk, fmt='%d %d %d')

Why this works:

  • Chunked processing: Avoids loading the entire 200M+ row array into memory at once, which reduces overhead and prevents memory bottlenecks.
  • Explicit format: The fmt='%d %d %d' ensures each line has three integers separated by spaces, matching your exact requirement.
  • Reduced I/O operations: Writing in batches minimizes the number of file system calls, which is one of the biggest slowdowns in large file writes.

If docword is a Regular Python List

If you're working with a standard list of lists, you can still write efficiently by generating lines in chunks and using writelines (which is far faster than calling write in a loop):

# Your predefined values
D = 68601
W = 1500
NNZ = 205806000
docword = [[0, 170, 4], [0, 1856, 4], ..., [68601, 0, 0]]  # Your large list

with open('docword_output.txt', 'w') as f:
    # Write metadata first
    f.write(f"{D}\n{W}\n{NNZ}\n")
    
    # Write data in chunks
    chunk_size = 1_000_000
    for start_idx in range(0, len(docword), chunk_size):
        end_idx = min(start_idx + chunk_size, len(docword))
        chunk = docword[start_idx:end_idx]
        # Convert chunk to a list of formatted strings
        lines = [' '.join(map(str, row)) + '\n' for row in chunk]
        # Write all lines in the chunk at once
        f.writelines(lines)

Why this works:

  • Batch string generation: Converting a chunk of rows to strings in one go is more efficient than processing each row individually.
  • writelines: This method writes multiple lines in a single I/O operation, which drastically speeds up the process compared to looping with write().

Both approaches will produce a text file that meets your requirements:

  1. First three lines are D, W, NNZ each on their own line.
  2. Every subsequent line contains three space-separated integers: docID, wordID, count.
  3. The file is fully human-readable and doesn't use any binary formats like pickle.

内容的提问来源于stack exchange,提问作者nightrain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:01:18