You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

3.3M脏城市名匹配2.7K干净城市名的高效方案咨询

高效城市名近似匹配优化方案

针对3.3M条脏城市名与2.7K条干净城市名的匹配场景,核心问题是原方案的O(n*m)全量比较复杂度导致计算量爆炸(约8.9亿次匹配)。以下是从「减少计算量」和「提升单匹配速度」两个维度出发的优化方案:

1. 先做文本预处理(必做前置步骤)

统一文本格式,减少无效匹配和重复计算:

import pandas as pd

# 预处理函数:转小写、去首尾空格、移除非字母字符
def preprocess(text):
    text = str(text).strip().lower()
    return ''.join([c for c in text if c.isalpha()])

# 对两个DataFrame的城市名做标准化处理
df1['cleaned_city'] = df1['city_name'].apply(preprocess)
df2['cleaned_city'] = df2['city_name'].apply(preprocess)

# 对脏数据去重,先处理唯一值再映射回原表,大幅减少重复计算
df1_unique = df1[['cleaned_city']].drop_duplicates().reset_index(drop=True)
clean_city_list = df2['cleaned_city'].unique().tolist()

2. 用RapidFuzz替代FuzzyWuzzy(速度提升50-100倍)

RapidFuzz是FuzzyWuzzy的C重写版本,API兼容但性能碾压纯Python实现,同时通过score_cutoff过滤低匹配度结果:

from rapidfuzz import process, fuzz

def get_top_matches(text, candidates, limit=10, score_cutoff=70):
    # WRatio适配大小写、部分匹配等场景,适合城市名匹配
    matches = process.extract(text, candidates, limit=limit, scorer=fuzz.WRatio, score_cutoff=score_cutoff)
    return [(match[0], match[1]) for match in matches]

# 对去重后的脏数据批量匹配
df1_unique['matches'] = df1_unique['cleaned_city'].apply(lambda x: get_top_matches(x, clean_city_list))

# 将匹配结果映射回原3.3M条数据
df1 = df1.merge(df1_unique[['cleaned_city', 'matches']], on='cleaned_city', how='left')

3. N-Gram倒排索引(进一步减少比较次数)

通过3-Gram索引快速筛选候选城市,把O(nm)复杂度降到O(nk)(k为每个脏城市的候选数,远小于2.7K):

from collections import defaultdict
from rapidfuzz import process, fuzz

# 构建3-Gram倒排索引:key是3-Gram片段,value是对应的干净城市集合
def build_ngram_index(candidates, n=3):
    ngram_index = defaultdict(set)
    for city in candidates:
        if len(city) < n:
            ngrams = {city}
        else:
            ngrams = {city[i:i+n] for i in range(len(city)-n+1)}
        for gram in ngrams:
            ngram_index[gram].add(city)
    return ngram_index

# 生成索引
ngram_index = build_ngram_index(clean_city_list)

# 先通过3-Gram筛选候选,再做模糊匹配
def get_top_matches_with_candidates(text, ngram_index, all_candidates, limit=10, score_cutoff=70):
    # 提取当前脏城市的3-Gram片段
    if len(text) < 3:
        ngrams = {text}
    else:
        ngrams = {text[i:i+3] for i in range(len(text)-2)}
    # 找到有重叠3-Gram的候选城市
    candidates = set()
    for gram in ngrams:
        candidates.update(ngram_index.get(gram, set()))
    # 无候选时 fallback 到全量匹配
    candidates = list(candidates) if candidates else all_candidates
    # 用RapidFuzz做精准匹配
    matches = process.extract(text, candidates, limit=limit, scorer=fuzz.WRatio, score_cutoff=score_cutoff)
    return [(match[0], match[1]) for match in matches]

# 批量处理去重后的脏数据
df1_unique['matches'] = df1_unique['cleaned_city'].apply(lambda x: get_top_matches_with_candidates(x, ngram_index, clean_city_list))

# 映射回原表
df1 = df1.merge(df1_unique[['cleaned_city', 'matches']], on='cleaned_city', how='left')

4. 并行处理(可选,多核机器提速)

用swifter自动实现并行化处理,无需手动编写多进程代码:

import swifter

df1_unique['matches'] = df1_unique['cleaned_city'].swifter.apply(lambda x: get_top_matches_with_candidates(x, ngram_index, clean_city_list))

内容的提问来源于stack exchange,提问作者Anisha A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 08:08:30