You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在difflib.get_close_matches中使用导入的CSV列作为字符串

How to Match CSV Columns with difflib.get_close_matches

Hey there! Let's walk through how to get this done step by step—since you're new to Python, I'll keep things clear and include fully working code you can adapt to your files.

First, we'll use pandas (a super popular library for data handling) to read/write CSVs, and of course difflib for the matching logic. If you don't have pandas installed yet, run this in your terminal first:

pip install pandas

Full Working Code

Here's a complete example with comments explaining each part—just replace the placeholder names with your actual file paths and column names:

import pandas as pd
from difflib import get_close_matches

# Step 1: Load your CSV files
# Replace 'csv1.csv' and 'csv2.csv' with your actual file paths
csv1 = pd.read_csv('csv1.csv')
csv2 = pd.read_csv('csv2.csv')

# Step 2: Extract the Master column from CSV2 into a list
# Replace 'Master' with the exact column name in your CSV2
master_list = csv2['Master'].tolist()

# Step 3: Define a function to find matches for each string
def find_close_match(input_string):
    # Handle empty values in CSV1's column (so we don't throw errors)
    if pd.isna(input_string):
        return ""
    
    # Use get_close_matches:
    # - input_string: the value from CSV1 we want to match
    # - master_list: all values from CSV2's Master column
    # - n=1: return up to 1 best match (adjust if you want more)
    # - cutoff=0.6: only return matches with similarity > 0.6
    matches = get_close_matches(input_string, master_list, n=1, cutoff=0.6)
    
    # Return the first match if there is one, else return an empty string
    return matches[0] if matches else ""

# Step 4: Apply the function to your target column in CSV1
# Replace 'Your_Target_Column' with the column name from CSV1 you want to compare
# 'Matched_Result' is the name of the new column we'll add to CSV1
csv1['Matched_Result'] = csv1['Your_Target_Column'].apply(find_close_match)

# Step 5: Save the updated CSV1 with the new column
csv1.to_csv('updated_csv1.csv', index=False)

Key Explanations

  • Why pandas? It makes reading/writing CSVs and manipulating columns way easier than the built-in csv module, especially for beginners. The apply() method lets us run our matching function on every row in the target column with one line of code.
  • get_close_matches parameters:
    • n: How many top matches you want to return (I set it to 1 for simplicity, but you can change it to 2/3 if you want multiple matches).
    • cutoff: The similarity threshold you specified (0.6)—only matches with a ratio higher than this will be kept.
  • Handling empty values: We added a check for pd.isna(input_string) to avoid errors if your CSV1 column has blank cells.

Troubleshooting Tips

  • Double-check that your column names in the code exactly match the ones in your CSVs (capitalization matters!).
  • If you get a "file not found" error, make sure your CSV files are in the same folder as your Python script, or use the full file path (like C:/Users/YourName/Documents/csv1.csv).
  • If you want to return multiple matches instead of just one, change n=1 to n=3 (for example) and adjust the return line to something like ", ".join(matches) if matches else "" to join them into a single string.

内容的提问来源于stack exchange,提问作者PythonNooob2018

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:27:14