如何在.NET中查找并读取名称存在编码异常的文件?
我太懂这种头疼的情况了——在Linux下创建了一个带无效UTF-8编码的文件名,.NET明明能枚举到它,可真要打开的时候就报错找不到文件,简直离谱。先把你的问题场景再理清楚:
你遇到的具体操作与问题
首先是创建这个畸形文件名的命令:
printf '\xC3\x28malformed_name.txt' | xargs touch
用ls验证文件确实存在:
ls ''$'\303''(malformed_name.txt'
然后用F#代码枚举目录,.NET确实能看到这个文件:
open System.IO Directory.EnumerateFiles "." |> Seq.toList // 输出 ["./�(malformed_name.txt"]
但尝试打开时直接报错,这段代码:
open System.IO let path = Directory.EnumerateFiles(".") |> Seq.head printfn "%s" path let fileStream = File.OpenRead(path) let reader = new StreamReader(fileStream) let content = reader.ReadToEnd() printfn "%s" content
运行后得到错误:
dotnet fsi scratch.fsx
./�(malformed_name.txt
System.IO.FileNotFoundException: Could not find file '/demo/�(malformed_name.txt'.
File name: '/demo/�(malformed_name.txt'
at Interop.ThrowExceptionForIoErrno(ErrorInfo errorInfo, String path, Boolean isDirError)
at Microsoft.Win32.SafeHandles.SafeFileHandle.Open(String path, OpenFlags flags, Int32 mode, Boolean failForSymlink, Boolean& wasSymlink, Func`4 createOpenException)
...
Stopped due to error
问题根源
Linux系统的文件名本质是原始字节序列,并不要求必须是有效的UTF-8编码。但.NET默认会把这些字节序列尝试解码成UTF-8字符串,遇到无效编码的字节时,会自动替换成通用替换字符�。这就导致你拿到的路径字符串和文件实际的原始字节路径不匹配,调用系统API打开时自然找不到文件。
解决方案
针对这种情况,.NET从Core 3.0开始提供了直接操作原始字节路径的API,绕过UTF-8解码的坑,具体操作如下:
1. 使用FileSystemInfo.FullNameAsBytes获取原始字节路径
你可以通过DirectoryInfo枚举文件,然后用FullNameAsBytes属性获取文件路径的原始字节序列,再用这个字节数组打开文件:
open System.IO let dirInfo = DirectoryInfo(".") // 获取目标文件的FileInfo对象 let targetFile = dirInfo.EnumerateFiles() |> Seq.find (fun fi -> fi.Name.Contains("malformed_name")) // 获取原始字节形式的完整路径 let rawPathBytes = targetFile.FullNameAsBytes // 用字节数组直接打开文件 use fileStream = File.OpenRead(rawPathBytes) use reader = new StreamReader(fileStream) let content = reader.ReadToEnd() printfn "文件内容:%s" content
2. 其他注意事项
- 这种问题只出现在类Unix系统(Linux/macOS),Windows的文件名基于UTF-16,不允许存在无效编码的文件名。
- 处理这类文件时,尽量避免将路径转换为字符串操作,保持字节序列的原始性,才能准确匹配系统中的文件。
内容来源于stack exchange

