使用Python字典在CSV中实现多字符串批量替换(修订版)
Hey there! Sounds like you’re off to a solid start by building that replacement dictionary—great job taking the first step as a new programmer. Let’s walk through how to put that dictionary to work to fix up your target CSV exactly like your example shows.
First, let’s confirm you’ve got your dictionary mapped correctly from misspelled terms to their fixes. If you’re loading this from a correction guide CSV (which is the cleanest way), here’s how you can build that dict using Python’s built-in csv module:
import csv # Initialize an empty dictionary to hold our corrections correction_dict = {} # Load the correction guide CSV with open('correction_guide.csv', mode='r') as guide_file: # Assuming your guide has columns named "misspelled" and "correct" reader = csv.DictReader(guide_file) for row in reader: correction_dict[row['misspelled']] = row['correct']
Just adjust the column names ("misspelled"/"correct") to match whatever your correction CSV uses—no fancy libraries needed here!
Now let’s take that dictionary and use it to update your target CSV. We’ll read each row, check if the first column (since your example uses the first column for misspellings) needs a fix, then write the corrected row to a new file (always safer to write to a new file instead of overwriting the original!):
# Define input and output file paths input_csv = 'target.csv' output_csv = 'corrected_target.csv' with open(input_csv, mode='r') as infile, open(output_csv, mode='w', newline='') as outfile: reader = csv.reader(infile) writer = csv.writer(outfile) # Write the header row first (if your CSV has one) header = next(reader) writer.writerow(header) # Loop through each row in the target CSV for row in reader: # Grab the value from the first column (adjust index if your misspellings are elsewhere) original_term = row[0] # Use .get() to safely replace—if the term isn't in the dict, keep it as-is corrected_term = correction_dict.get(original_term, original_term) # Update the row with the corrected term row[0] = corrected_term # Write the fixed row to the output file writer.writerow(row)
Suppose your correction guide CSV looks like this:
misspelled,correct chikn,chicken bcon,Bacon egs,eggs
And your target CSV is:
item,id,code chikn,1,a bcon,2,b egs,3,c
After running the code, your corrected_target.csv will be exactly what you want:
item,id,code chicken,1,a Bacon,2,b eggs,3,c
- Backup First: Always make a copy of your original target CSV before running scripts—accidents happen, and you don’t want to lose your raw data!
- Adjust Column Index: If your misspelled terms are in the second column instead of the first, just change
row[0]torow[1](remember Python uses 0-indexing!). - Case Insensitivity: If you need to fix terms regardless of capitalization (like
ChiknorBCON), tweak the dictionary and replacement step to use lowercase:# When building the dict correction_dict[row['misspelled'].lower()] = row['correct'] # When replacing corrected_term = correction_dict.get(original_term.lower(), original_term)
内容的提问来源于stack exchange,提问作者lostnthought247

