内存受限下迭代将字典写入CSV的技术实现问询
Got it, let's fix this iterative CSV writing for memory-constrained scenarios. The key here is to write each processed row immediately to the file instead of buffering everything in memory, and handle the header correctly when the file doesn't exist yet. Here's the polished implementation:
import csv import os import numpy as np # Assuming these are defined elsewhere: no_phons, get_phonema def save_phons_2_csv(pandas_dataset, csv_name): # Check if we need to write the header (only for new files) write_header = not os.path.isfile(csv_name) # Open file in append mode to avoid overwriting existing data with open(csv_name, 'a', newline='', encoding='utf-8') as csv_file: csv_writer = csv.writer(csv_file) for index_r, row in pandas_dataset.iterrows(): # Iterate through each phonetic entry in the row's 'phons' field for index, phon_dict in enumerate(row['phons']): if phon_dict['phon'] not in no_phons: # Get processed phonema data as a dictionary dicc = get_phonema(row, index) # Convert dict values to a list for CSV writing csv_row = list(dicc.values()) # Write header once if the file is new if write_header: csv_writer.writerow(dicc.keys()) write_header = False # Disable header writing after first use # Write the single row immediately to avoid memory buildup csv_writer.writerow(csv_row)
Key improvements for memory efficiency:
- Immediate row flushing: Instead of collecting all rows in a large array/list, we write each processed phonetic entry right after generating it. This keeps memory usage constant (only holding one row at a time) instead of growing with the dataset.
- Safe CSV handling: The
csv.writermodule automatically handles edge cases like commas in values, quotes, and newline characters—something manual string writing would easily break. - Simplified file setup: We skip creating an empty file upfront; using append mode (
'a') handles both new and existing files seamlessly. - Header logic: We only write the header once when the file is first created, ensuring consistent column structure.
Optional speed optimization:
If your pandas dataset is very large, iterrows() can be slow. Swap it for itertuples() which returns lightweight tuples instead of Series objects, boosting iteration speed:
for row in pandas_dataset.itertuples(index=False): for index, phon_dict in enumerate(row.phons): # Same processing logic, but access fields with row.field_name instead of row['field_name'] if phon_dict['phon'] not in no_phons: dicc = get_phonema(row, index) # ... rest of the writing code
Notes:
- Ensure
get_phonemareturns a dictionary with consistent keys every time—this guarantees the header matches all subsequent rows. - The
newline=''parameter inopen()prevents extra blank lines from appearing in CSV files on Windows. - If you need to include
dicc['label']in the output, just make sure it's part of the dictionary returned byget_phonema.
内容的提问来源于stack exchange,提问作者Isaac
相关产品推荐
相关产品推荐

