如何修改C#代码读取含多个JSON文件的GZ压缩包
处理GZ压缩包中多个JSON文件的C#代码修改方案
你的核心问题是:原代码将GZ解压后的内容当成纯文本处理,但实际上这是一个TAR归档文件(从示例里的ustar标识可看出),里面包含多个独立的JSON文件。直接查找第一个[会把TAR头和其他文件内容混在一起,导致多文件场景失效。
原代码(仅支持单个JSON文件)
string responseBodyUrl = getResponseObject.response_body_url; // get the tar.gr file and expand WebClient webClient = new WebClient(); Stream stream = webClient.OpenRead(responseBodyUrl); MemoryStream memoryStream = new MemoryStream(); GZipStream gzipStream = new GZipStream(stream, CompressionMode.Decompress); gzipStream.CopyTo(memoryStream); gzipStream.Close(); stream.Close(); memoryStream.Position = 0; StreamReader reader = new StreamReader(memoryStream); string memstreamjson = reader.ReadToEnd(); reader.Close(); memoryStream.Close(); // find the index of the first '[' character int index = memstreamjson.IndexOf('['); System.IO.File.AppendAllText(@"memstreamjson.log",memstreamjson.ToString().TrimEnd() + Environment.NewLine); // if found if (index != -1) { // get the substring from that index to the end string indexedmemstreamjson = memstreamjson.Substring(index); // parse the JSON string as an array JArray arr = JArray.Parse(indexedmemstreamjson.ToString()); // loop through each element of the array foreach (JObject obj in arr) { // get the status_code value of the JObject string status_code = (string)obj["status_code"]; // ... 后续逻辑 } }
优化方案:解析TAR归档提取每个JSON文件
从.NET 6开始,官方提供System.Formats.Tar库可直接解析TAR文件,无需第三方依赖。下面是修改后的代码,会逐个提取TAR中的JSON文件并处理:
修改后代码
using System.Formats.Tar; using System.IO.Compression; using Newtonsoft.Json.Linq; string responseBodyUrl = getResponseObject.response_body_url; // 下载并解压GZ流,直接处理Tar内容(无需全部写入内存) using (WebClient webClient = new WebClient()) using (Stream gzStream = webClient.OpenRead(responseBodyUrl)) using (GZipStream decompressStream = new GZipStream(gzStream, CompressionMode.Decompress)) using (TarReader tarReader = new TarReader(decompressStream)) { TarEntry entry; while ((entry = tarReader.GetNextEntry()) != null) { // 跳过目录和非JSON文件 if (entry.EntryType != TarEntryType.RegularFile || !entry.Name.EndsWith(".json", StringComparison.OrdinalIgnoreCase)) { continue; } Console.WriteLine($"正在处理文件: {entry.Name}"); // 读取当前JSON文件的内容 using (Stream entryStream = entry.OpenEntryStream()) using (StreamReader reader = new StreamReader(entryStream)) { string jsonContent = reader.ReadToEnd(); // 解析JSON数组(假设每个JSON文件内容都是数组) if (JArray.TryParse(jsonContent, out JArray jsonArr)) { foreach (JObject obj in jsonArr) { string statusCode = obj["status_code"]?.ToString(); Console.WriteLine($" Status Code: {statusCode}"); // 这里可以添加你的后续业务逻辑 } } else { Console.WriteLine($" 警告:{entry.Name} 不是有效的JSON数组"); } } } }
关键改进点
- 直接流式处理:无需把整个解压内容写入内存,节省内存占用,适配大文件场景
- 正确解析TAR归档:通过
TarReader遍历每个文件条目,自动跳过目录和非JSON文件 - 独立处理每个JSON:对每个JSON文件单独读取、解析,避免内容混杂导致的解析失败
- 容错性提升:用
JArray.TryParse替代直接解析,单个文件出错不会中断整个流程
注意事项
- 若项目基于.NET 5或更早版本,需使用第三方库(如
SharpCompress)解析TAR文件,用法逻辑一致 - 若存在单个JSON对象的文件(而非数组),需将解析逻辑调整为
JObject.Parse
内容的提问来源于stack exchange,提问作者Gerald Oakham
相关产品推荐
相关产品推荐

