如何在Python脚本中实现CSV对比后将新数据追加至历史文件?
Got it, let's tweak your script to handle appending the new entries to allhistory.csv seamlessly. I'll show you two versions—one that stays close to your original code, and a more robust one using the csv module properly (better for edge cases like weird formatting or headers).
First version: Minimal changes to your existing code
This keeps your original comparison logic but adds the append step without re-reading the update.csv file (more efficient):
import csv # Compare files and collect new entries first with open('allhistory.csv', 'r') as t1, open('filewithnewdata.csv', 'r') as t2: fileone = t1.readlines() filetwo = t2.readlines() new_entries = [line for line in filetwo if line not in fileone] # Write to update.csv (your original output) with open('update.csv', 'w') as outFile: outFile.writelines(new_entries) # Append new entries to allhistory.csv if new_entries: # Only run if there's actual new data to add with open('allhistory.csv', 'a', newline='') as history_file: history_file.writelines(new_entries)
Key tweaks here:
- Stored the new entries in a variable first so we don't have to read
update.csvagain later - Used
a(append) mode forallhistory.csv—this adds content to the end instead of overwriting it - Added
newline=''to avoid extra blank lines showing up on Windows (a common CSV gotcha) - Added a check for empty
new_entriesto skip unnecessary file operations
Second version: More robust (using csv module properly)
If your CSV files have headers, quoted fields, or inconsistent line endings, using the csv module's reader/writer is safer. This avoids false positives for "duplicate" lines caused by formatting differences:
import csv def get_csv_rows(file_path): """Helper function to read all rows from a CSV file""" with open(file_path, 'r', newline='') as f: return list(csv.reader(f)) # Read rows from both files history_rows = get_csv_rows('allhistory.csv') new_data_rows = get_csv_rows('filewithnewdata.csv') # Find unique new rows (convert to tuples since lists can't be stored in sets) history_row_set = set(tuple(row) for row in history_rows) new_entries = [row for row in new_data_rows if tuple(row) not in history_row_set] # Write new entries to update.csv with open('update.csv', 'w', newline='') as outfile: writer = csv.writer(outfile) writer.writerows(new_entries) # Append to allhistory.csv if there's new data if new_entries: with open('allhistory.csv', 'a', newline='') as history_file: writer = csv.writer(history_file) writer.writerows(new_entries)
Why this is better:
- Handles CSV-specific formatting (like fields with commas inside quotes) correctly
- Doesn't get confused by different line endings (
\nvs\r\n) - Makes it easier to handle headers if needed (e.g., you could skip the first row of
filewithnewdata.csvif you don't want to re-add headers)
Just pick whichever version fits your use case—both will get the job done of appending your new entries to the history file!
内容的提问来源于stack exchange,提问作者user1959064

