3.3M脏城市名匹配2.7K干净城市名的高效方案咨询
高效城市名近似匹配优化方案
针对3.3M条脏城市名与2.7K条干净城市名的匹配场景,核心问题是原方案的O(n*m)全量比较复杂度导致计算量爆炸(约8.9亿次匹配)。以下是从「减少计算量」和「提升单匹配速度」两个维度出发的优化方案:
1. 先做文本预处理(必做前置步骤)
统一文本格式,减少无效匹配和重复计算:
import pandas as pd # 预处理函数:转小写、去首尾空格、移除非字母字符 def preprocess(text): text = str(text).strip().lower() return ''.join([c for c in text if c.isalpha()]) # 对两个DataFrame的城市名做标准化处理 df1['cleaned_city'] = df1['city_name'].apply(preprocess) df2['cleaned_city'] = df2['city_name'].apply(preprocess) # 对脏数据去重,先处理唯一值再映射回原表,大幅减少重复计算 df1_unique = df1[['cleaned_city']].drop_duplicates().reset_index(drop=True) clean_city_list = df2['cleaned_city'].unique().tolist()
2. 用RapidFuzz替代FuzzyWuzzy(速度提升50-100倍)
RapidFuzz是FuzzyWuzzy的C重写版本,API兼容但性能碾压纯Python实现,同时通过score_cutoff过滤低匹配度结果:
from rapidfuzz import process, fuzz def get_top_matches(text, candidates, limit=10, score_cutoff=70): # WRatio适配大小写、部分匹配等场景,适合城市名匹配 matches = process.extract(text, candidates, limit=limit, scorer=fuzz.WRatio, score_cutoff=score_cutoff) return [(match[0], match[1]) for match in matches] # 对去重后的脏数据批量匹配 df1_unique['matches'] = df1_unique['cleaned_city'].apply(lambda x: get_top_matches(x, clean_city_list)) # 将匹配结果映射回原3.3M条数据 df1 = df1.merge(df1_unique[['cleaned_city', 'matches']], on='cleaned_city', how='left')
3. N-Gram倒排索引(进一步减少比较次数)
通过3-Gram索引快速筛选候选城市,把O(nm)复杂度降到O(nk)(k为每个脏城市的候选数,远小于2.7K):
from collections import defaultdict from rapidfuzz import process, fuzz # 构建3-Gram倒排索引:key是3-Gram片段,value是对应的干净城市集合 def build_ngram_index(candidates, n=3): ngram_index = defaultdict(set) for city in candidates: if len(city) < n: ngrams = {city} else: ngrams = {city[i:i+n] for i in range(len(city)-n+1)} for gram in ngrams: ngram_index[gram].add(city) return ngram_index # 生成索引 ngram_index = build_ngram_index(clean_city_list) # 先通过3-Gram筛选候选,再做模糊匹配 def get_top_matches_with_candidates(text, ngram_index, all_candidates, limit=10, score_cutoff=70): # 提取当前脏城市的3-Gram片段 if len(text) < 3: ngrams = {text} else: ngrams = {text[i:i+3] for i in range(len(text)-2)} # 找到有重叠3-Gram的候选城市 candidates = set() for gram in ngrams: candidates.update(ngram_index.get(gram, set())) # 无候选时 fallback 到全量匹配 candidates = list(candidates) if candidates else all_candidates # 用RapidFuzz做精准匹配 matches = process.extract(text, candidates, limit=limit, scorer=fuzz.WRatio, score_cutoff=score_cutoff) return [(match[0], match[1]) for match in matches] # 批量处理去重后的脏数据 df1_unique['matches'] = df1_unique['cleaned_city'].apply(lambda x: get_top_matches_with_candidates(x, ngram_index, clean_city_list)) # 映射回原表 df1 = df1.merge(df1_unique[['cleaned_city', 'matches']], on='cleaned_city', how='left')
4. 并行处理(可选,多核机器提速)
用swifter自动实现并行化处理,无需手动编写多进程代码:
import swifter df1_unique['matches'] = df1_unique['cleaned_city'].swifter.apply(lambda x: get_top_matches_with_candidates(x, ngram_index, clean_city_list))
内容的提问来源于stack exchange,提问作者Anisha A
相关产品推荐
相关产品推荐

