C#实现Excel指定列文件名匹配(80%相似度)及文件重命名问询
嘿,这个需求我之前帮不少开发者实现过!要从完全匹配改成80%匹配度的模糊查找,核心是得有个靠谱的字符串相似度计算方法,最常用的就是编辑距离(Levenshtein Distance)——它能衡量两个文件名的差异程度,再转换成匹配度百分比就搞定了。下面给你一步步拆解实现思路和代码:
核心思路拆解
- 计算字符串相似度:用编辑距离算法算出目标文件名和目录中每个文件名的差异,再转换为0-100%的匹配度
- 遍历筛选文件:遍历指定目录下的所有文件,逐个计算与目标文件名的相似度,筛选出≥80%的文件
- 安全重命名:对匹配的文件执行重命名,同时处理文件名重复、权限异常等问题
相似度计算工具类
先写一个静态工具类,实现Levenshtein编辑距离的计算,再封装成相似度百分比的方法:
public static class StringSimilarityHelper { // 计算两个字符串的编辑距离(插入/删除/替换的最少操作次数) private static int CalculateLevenshteinDistance(string s, string t) { int n = s.Length; int m = t.Length; int[,] distanceMatrix = new int[n + 1, m + 1]; if (n == 0) return m; if (m == 0) return n; // 初始化矩阵边界 for (int i = 0; i <= n; distanceMatrix[i, 0] = i++) { } for (int j = 0; j <= m; distanceMatrix[0, j] = j++) { } // 填充矩阵计算距离 for (int i = 1; i <= n; i++) { for (int j = 1; j <= m; j++) { int cost = (t[j - 1] == s[i - 1]) ? 0 : 1; distanceMatrix[i, j] = Math.Min( Math.Min(distanceMatrix[i - 1, j] + 1, distanceMatrix[i, j - 1] + 1), distanceMatrix[i - 1, j - 1] + cost); } } return distanceMatrix[n, m]; } // 计算相似度百分比(返回0-100之间的数值) public static double CalculateSimilarity(string str1, string str2) { if (string.IsNullOrEmpty(str1) && string.IsNullOrEmpty(str2)) return 100; if (string.IsNullOrEmpty(str1) || string.IsNullOrEmpty(str2)) return 0; // 忽略大小写(可选,根据需求注释掉) str1 = str1.ToLowerInvariant(); str2 = str2.ToLowerInvariant(); int editDistance = CalculateLevenshteinDistance(str1, str2); int maxLength = Math.Max(str1.Length, str2.Length); return maxLength == 0 ? 100 : (1 - (double)editDistance / maxLength) * 100; } }
完整业务实现代码
假设你已经通过EPPlus/NPOI等库读取到Excel指定列的invoiceName值,下面是遍历目录、筛选匹配文件并重命名的完整逻辑:
using System.IO; using System.Linq; public class InvoiceFileRenamer { public static void ProcessFileRename(string targetDirectory, string invoiceName, double requiredSimilarity = 80) { if (!Directory.Exists(targetDirectory)) { throw new DirectoryNotFoundException("指定的目录不存在,请检查路径"); } // 获取目录下所有文件(可添加扩展名过滤,比如只找.pdf/.xlsx:Directory.GetFiles(targetDirectory, "*.pdf")) var allFiles = Directory.GetFiles(targetDirectory); foreach (var filePath in allFiles) { // 获取文件名(可选择是否去掉扩展名:Path.GetFileNameWithoutExtension(filePath)) string currentFileName = Path.GetFileName(filePath); double similarity = StringSimilarityHelper.CalculateSimilarity(currentFileName, invoiceName); if (similarity >= requiredSimilarity) { Console.WriteLine($"找到匹配文件:{currentFileName} | 相似度:{similarity:F2}%"); // 构造新文件路径,避免重名覆盖 string newFilePath = Path.Combine(targetDirectory, invoiceName); int duplicateCounter = 1; while (File.Exists(newFilePath)) { string nameWithoutExt = Path.GetFileNameWithoutExtension(invoiceName); string ext = Path.GetExtension(invoiceName); newFilePath = Path.Combine(targetDirectory, $"{nameWithoutExt}({duplicateCounter}){ext}"); duplicateCounter++; } // 执行重命名并处理异常 try { File.Move(filePath, newFilePath); Console.WriteLine($"重命名成功:{currentFileName} -> {Path.GetFileName(newFilePath)}"); } catch (IOException ex) { Console.WriteLine($"重命名失败(文件被占用或无权限):{ex.Message}"); } catch (Exception ex) { Console.WriteLine($"未知错误:{ex.Message}"); } } } } }
额外注意事项
- 算法调整:如果Levenshtein距离不符合你的匹配场景,也可以换成Jaccard相似度(基于字符交集),不过编辑距离对文件名的错字、漏字匹配更友好
- 多匹配处理:如果有多个文件相似度≥80%,上面的代码会逐个重命名,你可以修改逻辑,比如只保留相似度最高的那个,或者提示用户手动选择
- 性能优化:如果目录下文件极多,可以先按文件名长度过滤(比如只找长度和目标文件名相差±20%的文件),再计算相似度,提升效率
内容的提问来源于stack exchange,提问作者Richeek Dey
相关产品推荐
相关产品推荐

