You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用C#匹配存在差异的字符串并关联对应ID?

企业名称模糊匹配方案(C#实现)

针对CRM系统企业简称与Web应用完整注册名的匹配需求,可通过字符串标准化+相似度计算的组合方案实现精准匹配,具体实现如下:

核心思路

  1. 字符串标准化:统一格式、移除干扰字符、替换常见缩写,消除简称与全称的格式差异
  2. 相似度计算:通过编辑距离(Levenshtein Distance)量化两个标准化后字符串的相似程度,设定阈值判断是否匹配

代码实现

1. 字符串标准化方法

将不同格式的企业名称转换为统一的标准格式,消除简称、特殊符号、大小写带来的差异:

private static string NormalizeCompanyName(string name)
{
    if (string.IsNullOrWhiteSpace(name))
        return string.Empty;

    // 统一转为小写
    var normalized = name.ToLowerInvariant();

    // 移除所有非字母数字和空格的字符
    normalized = System.Text.RegularExpressions.Regex.Replace(normalized, @"[^a-z0-9\s]", "");

    // 替换行业常见缩写为全称,可根据业务场景扩展
    var abbreviationMap = new Dictionary<string, string>
    {
        {"ltd", "limited"},
        {"pty", "proprietary"},
        {"inc", "incorporated"},
        {"corp", "corporation"},
        {"co", "company"}
    };

    foreach (var kvp in abbreviationMap)
    {
        // 匹配独立单词,避免误替换
        normalized = System.Text.RegularExpressions.Regex.Replace(normalized, @"\b" + kvp.Key + @"\b", kvp.Value);
    }

    // 合并多余空格并去除首尾空格
    normalized = System.Text.RegularExpressions.Regex.Replace(normalized, @"\s+", " ").Trim();

    return normalized;
}

2. 编辑距离计算方法

计算两个字符串之间的编辑距离(即最少需要多少次增删改操作才能将一个字符串转为另一个),用于后续相似度计算:

private static int CalculateLevenshteinDistance(string s, string t)
{
    int n = s.Length;
    int m = t.Length;
    int[,] d = new int[n + 1, m + 1];

    if (n == 0) return m;
    if (m == 0) return n;

    for (int i = 0; i <= n; d[i, 0] = i++) { }
    for (int j = 0; j <= m; d[0, j] = j++) { }

    for (int i = 1; i <= n; i++)
    {
        for (int j = 1; j <= m; j++)
        {
            int cost = (t[j - 1] == s[i - 1]) ? 0 : 1;
            d[i, j] = Math.Min(Math.Min(d[i - 1, j] + 1, d[i, j - 1] + 1), d[i - 1, j - 1] + cost);
        }
    }

    return d[n, m];
}

3. 匹配判断与结果输出

结合标准化和相似度计算,判断名称是否匹配,并生成所需的结果列表:

// 定义输出结果类(与需求一致)
public class AccountPairResult
{
    public string AccountNameA { get; set; }
    public int AccountNameB_ID { get; set; }
}

// 匹配判断方法
public static bool AreCompanyNamesMatching(string nameA, string nameB, double similarityThreshold = 0.9)
{
    var normalizedA = NormalizeCompanyName(nameA);
    var normalizedB = NormalizeCompanyName(nameB);

    if (string.IsNullOrWhiteSpace(normalizedA) || string.IsNullOrWhiteSpace(normalizedB))
        return false;

    // 优先判断完全匹配
    if (normalizedA == normalizedB)
        return true;

    // 计算相似度:1 - 编辑距离/最长字符串长度
    int maxLength = Math.Max(normalizedA.Length, normalizedB.Length);
    if (maxLength == 0)
        return false;

    int distance = CalculateLevenshteinDistance(normalizedA, normalizedB);
    double similarity = 1.0 - (double)distance / maxLength;

    // 达到阈值则判定为匹配
    return similarity >= similarityThreshold;
}

// 批量处理AccountPair,生成匹配结果
public static List<AccountPairResult> MatchAccounts(List<AccountPair> accountPairs)
{
    var results = new List<AccountPairResult>();

    foreach (var pair in accountPairs)
    {
        if (AreCompanyNamesMatching(pair.AccountNameA, pair.AccountNameB))
        {
            results.Add(new AccountPairResult
            {
                AccountNameA = pair.AccountNameA,
                AccountNameB_ID = pair.AccountNameB_ID
            });
        }
    }

    return results;
}

// 原始输入的数据类
public class AccountPair
{
    public string AccountNameA { get; set; }
    public string AccountNameB { get; set; }
    public int AccountNameB_ID { get; set; }
}

优化建议

  • 扩展缩写映射表:根据业务涉及的行业,添加更多专属缩写与全称的对应关系,提升标准化精度
  • 调整相似度阈值:如果简称与全称差异较大(比如包含地区后缀),可适当降低阈值(如0.8)
  • 分组预处理:针对大量数据,可先提取公司核心名称(如"John Doe")进行分组,减少需要计算相似度的配对数量,提升性能
  • 特殊规则补充:针对特定场景(如带地区后缀、特殊行业名称),添加额外的规则处理(如移除"(UK)"、"(PTE)"等后缀)

内容的提问来源于stack exchange,提问作者Hennericho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 15:29:54