如何用Python按指定格式将数据集写入.txt文件?
It sounds like you're dealing with a large document-word matrix and need a fast, reliable way to output it in that specific text format. The key issues with np.savetxt you're facing are likely either inefficient handling of huge arrays or incorrect format specification. Let's fix that with two optimized approaches depending on whether your docword data is a NumPy array or a regular Python list.
If docword is a NumPy Array
NumPy's savetxt can be efficient if you process the data in chunks (to avoid overwhelming memory) and explicitly set the format to match your requirements:
import numpy as np # Your predefined values D = 68601 W = 1500 NNZ = 205806000 docword = np.array([[0, 170, 4], [0, 1856, 4], ..., [68601, 0, 0]]) # Your large array # Open the output file with open('docword_output.txt', 'w') as f: # Write the metadata lines first f.write(f"{D}\n") f.write(f"{W}\n") f.write(f"{NNZ}\n") # Process data in chunks to minimize memory usage and speed up writes chunk_size = 1_000_000 # Adjust based on your available memory for start_idx in range(0, len(docword), chunk_size): end_idx = min(start_idx + chunk_size, len(docword)) chunk = docword[start_idx:end_idx] # Write chunk with integer format, space-separated np.savetxt(f, chunk, fmt='%d %d %d')
Why this works:
- Chunked processing: Avoids loading the entire 200M+ row array into memory at once, which reduces overhead and prevents memory bottlenecks.
- Explicit format: The
fmt='%d %d %d'ensures each line has three integers separated by spaces, matching your exact requirement. - Reduced I/O operations: Writing in batches minimizes the number of file system calls, which is one of the biggest slowdowns in large file writes.
If docword is a Regular Python List
If you're working with a standard list of lists, you can still write efficiently by generating lines in chunks and using writelines (which is far faster than calling write in a loop):
# Your predefined values D = 68601 W = 1500 NNZ = 205806000 docword = [[0, 170, 4], [0, 1856, 4], ..., [68601, 0, 0]] # Your large list with open('docword_output.txt', 'w') as f: # Write metadata first f.write(f"{D}\n{W}\n{NNZ}\n") # Write data in chunks chunk_size = 1_000_000 for start_idx in range(0, len(docword), chunk_size): end_idx = min(start_idx + chunk_size, len(docword)) chunk = docword[start_idx:end_idx] # Convert chunk to a list of formatted strings lines = [' '.join(map(str, row)) + '\n' for row in chunk] # Write all lines in the chunk at once f.writelines(lines)
Why this works:
- Batch string generation: Converting a chunk of rows to strings in one go is more efficient than processing each row individually.
writelines: This method writes multiple lines in a single I/O operation, which drastically speeds up the process compared to looping withwrite().
Both approaches will produce a text file that meets your requirements:
- First three lines are
D,W,NNZeach on their own line. - Every subsequent line contains three space-separated integers:
docID,wordID,count. - The file is fully human-readable and doesn't use any binary formats like pickle.
内容的提问来源于stack exchange,提问作者nightrain

