在difflib.get_close_matches中使用导入的CSV列作为字符串
How to Match CSV Columns with difflib.get_close_matches
Hey there! Let's walk through how to get this done step by step—since you're new to Python, I'll keep things clear and include fully working code you can adapt to your files.
First, we'll use pandas (a super popular library for data handling) to read/write CSVs, and of course difflib for the matching logic. If you don't have pandas installed yet, run this in your terminal first:
pip install pandas
Full Working Code
Here's a complete example with comments explaining each part—just replace the placeholder names with your actual file paths and column names:
import pandas as pd from difflib import get_close_matches # Step 1: Load your CSV files # Replace 'csv1.csv' and 'csv2.csv' with your actual file paths csv1 = pd.read_csv('csv1.csv') csv2 = pd.read_csv('csv2.csv') # Step 2: Extract the Master column from CSV2 into a list # Replace 'Master' with the exact column name in your CSV2 master_list = csv2['Master'].tolist() # Step 3: Define a function to find matches for each string def find_close_match(input_string): # Handle empty values in CSV1's column (so we don't throw errors) if pd.isna(input_string): return "" # Use get_close_matches: # - input_string: the value from CSV1 we want to match # - master_list: all values from CSV2's Master column # - n=1: return up to 1 best match (adjust if you want more) # - cutoff=0.6: only return matches with similarity > 0.6 matches = get_close_matches(input_string, master_list, n=1, cutoff=0.6) # Return the first match if there is one, else return an empty string return matches[0] if matches else "" # Step 4: Apply the function to your target column in CSV1 # Replace 'Your_Target_Column' with the column name from CSV1 you want to compare # 'Matched_Result' is the name of the new column we'll add to CSV1 csv1['Matched_Result'] = csv1['Your_Target_Column'].apply(find_close_match) # Step 5: Save the updated CSV1 with the new column csv1.to_csv('updated_csv1.csv', index=False)
Key Explanations
- Why pandas? It makes reading/writing CSVs and manipulating columns way easier than the built-in
csvmodule, especially for beginners. Theapply()method lets us run our matching function on every row in the target column with one line of code. - get_close_matches parameters:
n: How many top matches you want to return (I set it to 1 for simplicity, but you can change it to 2/3 if you want multiple matches).cutoff: The similarity threshold you specified (0.6)—only matches with a ratio higher than this will be kept.
- Handling empty values: We added a check for
pd.isna(input_string)to avoid errors if your CSV1 column has blank cells.
Troubleshooting Tips
- Double-check that your column names in the code exactly match the ones in your CSVs (capitalization matters!).
- If you get a "file not found" error, make sure your CSV files are in the same folder as your Python script, or use the full file path (like
C:/Users/YourName/Documents/csv1.csv). - If you want to return multiple matches instead of just one, change
n=1ton=3(for example) and adjust the return line to something like", ".join(matches) if matches else ""to join them into a single string.
内容的提问来源于stack exchange,提问作者PythonNooob2018
相关产品推荐
相关产品推荐

