如何使用C#匹配存在差异的字符串并关联对应ID?
企业名称模糊匹配方案(C#实现)
针对CRM系统企业简称与Web应用完整注册名的匹配需求,可通过字符串标准化+相似度计算的组合方案实现精准匹配,具体实现如下:
核心思路
- 字符串标准化:统一格式、移除干扰字符、替换常见缩写,消除简称与全称的格式差异
- 相似度计算:通过编辑距离(Levenshtein Distance)量化两个标准化后字符串的相似程度,设定阈值判断是否匹配
代码实现
1. 字符串标准化方法
将不同格式的企业名称转换为统一的标准格式,消除简称、特殊符号、大小写带来的差异:
private static string NormalizeCompanyName(string name) { if (string.IsNullOrWhiteSpace(name)) return string.Empty; // 统一转为小写 var normalized = name.ToLowerInvariant(); // 移除所有非字母数字和空格的字符 normalized = System.Text.RegularExpressions.Regex.Replace(normalized, @"[^a-z0-9\s]", ""); // 替换行业常见缩写为全称,可根据业务场景扩展 var abbreviationMap = new Dictionary<string, string> { {"ltd", "limited"}, {"pty", "proprietary"}, {"inc", "incorporated"}, {"corp", "corporation"}, {"co", "company"} }; foreach (var kvp in abbreviationMap) { // 匹配独立单词,避免误替换 normalized = System.Text.RegularExpressions.Regex.Replace(normalized, @"\b" + kvp.Key + @"\b", kvp.Value); } // 合并多余空格并去除首尾空格 normalized = System.Text.RegularExpressions.Regex.Replace(normalized, @"\s+", " ").Trim(); return normalized; }
2. 编辑距离计算方法
计算两个字符串之间的编辑距离(即最少需要多少次增删改操作才能将一个字符串转为另一个),用于后续相似度计算:
private static int CalculateLevenshteinDistance(string s, string t) { int n = s.Length; int m = t.Length; int[,] d = new int[n + 1, m + 1]; if (n == 0) return m; if (m == 0) return n; for (int i = 0; i <= n; d[i, 0] = i++) { } for (int j = 0; j <= m; d[0, j] = j++) { } for (int i = 1; i <= n; i++) { for (int j = 1; j <= m; j++) { int cost = (t[j - 1] == s[i - 1]) ? 0 : 1; d[i, j] = Math.Min(Math.Min(d[i - 1, j] + 1, d[i, j - 1] + 1), d[i - 1, j - 1] + cost); } } return d[n, m]; }
3. 匹配判断与结果输出
结合标准化和相似度计算,判断名称是否匹配,并生成所需的结果列表:
// 定义输出结果类(与需求一致) public class AccountPairResult { public string AccountNameA { get; set; } public int AccountNameB_ID { get; set; } } // 匹配判断方法 public static bool AreCompanyNamesMatching(string nameA, string nameB, double similarityThreshold = 0.9) { var normalizedA = NormalizeCompanyName(nameA); var normalizedB = NormalizeCompanyName(nameB); if (string.IsNullOrWhiteSpace(normalizedA) || string.IsNullOrWhiteSpace(normalizedB)) return false; // 优先判断完全匹配 if (normalizedA == normalizedB) return true; // 计算相似度:1 - 编辑距离/最长字符串长度 int maxLength = Math.Max(normalizedA.Length, normalizedB.Length); if (maxLength == 0) return false; int distance = CalculateLevenshteinDistance(normalizedA, normalizedB); double similarity = 1.0 - (double)distance / maxLength; // 达到阈值则判定为匹配 return similarity >= similarityThreshold; } // 批量处理AccountPair,生成匹配结果 public static List<AccountPairResult> MatchAccounts(List<AccountPair> accountPairs) { var results = new List<AccountPairResult>(); foreach (var pair in accountPairs) { if (AreCompanyNamesMatching(pair.AccountNameA, pair.AccountNameB)) { results.Add(new AccountPairResult { AccountNameA = pair.AccountNameA, AccountNameB_ID = pair.AccountNameB_ID }); } } return results; } // 原始输入的数据类 public class AccountPair { public string AccountNameA { get; set; } public string AccountNameB { get; set; } public int AccountNameB_ID { get; set; } }
优化建议
- 扩展缩写映射表:根据业务涉及的行业,添加更多专属缩写与全称的对应关系,提升标准化精度
- 调整相似度阈值:如果简称与全称差异较大(比如包含地区后缀),可适当降低阈值(如0.8)
- 分组预处理:针对大量数据,可先提取公司核心名称(如"John Doe")进行分组,减少需要计算相似度的配对数量,提升性能
- 特殊规则补充:针对特定场景(如带地区后缀、特殊行业名称),添加额外的规则处理(如移除"(UK)"、"(PTE)"等后缀)
内容的提问来源于stack exchange,提问作者Hennericho
相关产品推荐
相关产品推荐

