Zig语言:while循环中ArrayList追加值被最后条目覆盖问题
问题描述
作为Zig新手,我编写了一个函数,接收已清理特殊字符、每行仅含单个单词的文件路径,返回排序后的文件路径。但运行时发现,ArrayList最终全被最后一个单词的片段覆盖。
主函数代码:
const std = @import("std"); const GPA = std.heap.GeneralPurposeAllocator; const log = std.log; const File = std.fs.File; pub fn sortFile(file_path: []u8, buffer: []u8) ![]u8 { // 创建通用分配器 var general_purpose_allocator = GPA(.{}){}; const gpa = general_purpose_allocator.allocator(); // 根据OS适配路径 var encoded_path_buffer = gpa.alloc(u8, file_path.len) catch unreachable; const encoded_file_sub_path = try encodePathForOs(file_path, encoded_path_buffer); // 创建文件读取器 var file_read: File = try cwd().openFile(encoded_file_sub_path, .{}); defer file_read.close(); const file_reader = file_read.reader(); gpa.free(encoded_new_file_sub_path); // 存储单词的ArrayList var lines = std.ArrayList([]u8).init(gpa); // 逐行读取并插入 while (try file_reader.readUntilDelimiterOrEof(buffer, '\n')) |line| { log.warn("line: {s}\n", .{line}); lines.append(line) catch |err| { return err; }; log.warn("lines: {s}\n", .{lines.items}); } const items = lines.items; log.warn("items: {s}\n", .{items}); /// 以下为非核心代码,仅用于复现问题: defer lines.deinit(); return file_path; }
路径编码工具函数:
const builtin = @import("builtin"); const os_tag = builtin.os.tag; const unicode = std.unicode; const std = @import("std"); fn encodePathForOs(path: []u8, encoded_path_buffer: []u8) ![]u8 { if (os_tag == .windows) { var i: usize = 0; while (i < path.len) : (i += 1) { const codepoint = try unicode.utf8Decode(path[i .. i + 1]); _ = try unicode.wtf8Encode(codepoint, encoded_path_buffer[i..]); } return encoded_path_buffer; } else { return path; } }
输入文件内容:
hello this is a test this is a test this is a test this is a test
迭代日志显示,每次追加后ArrayList的内容逐渐被后续行覆盖,最终所有元素都变成最后一个单词的片段。尝试过克隆行、扩容ArrayList、使用insert等方法,均未解决问题。
问题原因
核心问题在于file_reader.readUntilDelimiterOrEof(buffer, '\n')返回的line是传入的buffer的切片,并非新分配的内存。每次循环时,新读取的内容会覆盖buffer里的旧数据,而ArrayList中存储的只是指向buffer的指针。当循环结束时,buffer中保留的是最后一行的内容,因此ArrayList的所有元素看起来都被最后一个单词覆盖了。
另外代码中还有两处小错误:
cwd()未指定命名空间,应改为std.fs.cwd();- 释放内存时误用了未定义的变量
encoded_new_file_sub_path,实际应为encoded_path_buffer。
解决方案
需要为每次读取到的line分配新内存,复制内容后再存入ArrayList,避免所有元素指向同一块缓冲区。具体修改如下:
- 在循环中,使用分配器的
dupe方法复制line的内容,将复制后的切片加入ArrayList:
while (try file_reader.readUntilDelimiterOrEof(buffer, '\n')) |line| { log.warn("line: {s}\n", .{line}); // 复制line到新分配的内存 const copied_line = try gpa.dupe(u8, line); lines.append(copied_line) catch |err| { gpa.free(copied_line); // 分配成功但append失败时,释放已分配的内存 return err; }; log.warn("lines: {s}\n", .{lines.items}); }
- 在
defer lines.deinit()之前,手动释放ArrayList中每个元素的内存(ArrayList的deinit只会释放自身存储元素指针的数组,不会释放每个元素指向的内存):
defer { for (lines.items) |line| { gpa.free(line); } lines.deinit(); }
- 修正代码中的两处小错误(补充
std.fs前缀、释放正确的变量)。
修改后的完整主函数代码:
const std = @import("std"); const GPA = std.heap.GeneralPurposeAllocator; const log = std.log; const File = std.fs.File; pub fn sortFile(file_path: []u8, buffer: []u8) ![]u8 { var general_purpose_allocator = GPA(.{}){}; const gpa = general_purpose_allocator.allocator(); defer _ = general_purpose_allocator.deinit(); // 补充:释放通用分配器 var encoded_path_buffer = gpa.alloc(u8, file_path.len) catch unreachable; const encoded_file_sub_path = try encodePathForOs(file_path, encoded_path_buffer); var file_read: File = try std.fs.cwd().openFile(encoded_file_sub_path, .{}); defer file_read.close(); const file_reader = file_read.reader(); gpa.free(encoded_path_buffer); var lines = std.ArrayList([]u8).init(gpa); defer { for (lines.items) |line| { gpa.free(line); } lines.deinit(); } while (try file_reader.readUntilDelimiterOrEof(buffer, '\n')) |line| { log.warn("line: {s}\n", .{line}); const copied_line = try gpa.dupe(u8, line); lines.append(copied_line) catch |err| { gpa.free(copied_line); return err; }; log.warn("lines: {s}\n", .{lines.items}); } const items = lines.items; log.warn("items: {s}\n", .{items}); return file_path; }
这样修改后,ArrayList中的每个元素都指向独立的内存块,不会再被后续读取的内容覆盖。
内容的提问来源于stack exchange,提问作者avifen
相关产品推荐
相关产品推荐

