Matlab如何提取长文本文件指定字段及指定迭代轮次数据
Matlab提取超长观测文本指定字段与指定迭代记录
针对重复出现的固定结构观测记录块,采用逐行扫描+正则匹配的方式实现大文件低内存占用提取,仅保留需要的坐标、真太阳俯角字段,以及第1次、第4次、最后一次迭代的记录。
单条观测记录参考格式如下:
Target Name: Tlil@z=0070km Observation Time: 2022-Dec-23 at 09:35:48.500 UT ---> Lon: 4.990 deg +/- .030 deg ---> Lat: 55.5 deg +/- 0.054 deg ---> Alt: 84.962 km +/- 35.577 meters True Solar Depression Angle: 4.1 deg Slant Range: 590.253 km
实现逻辑
- 遍历指定目录下所有观测文本文件,采用逐行读取模式,避免一次性加载全量超大文件导致内存溢出
- 以
Observation Time字段作为单条观测迭代记录的起始标记,先扫描全文件统计总迭代次数,定位需要提取的第1、4、最后一条记录的行位置 - 对目标记录行,通过正则匹配提取经度、纬度、高度、真太阳俯角4个字段的有效数值,自动跳过误差标注、斜距字段及其他无关冗余内容
- 提取结果自动结构化存储,支持直接导出为表格文件
实现代码
%% 初始化配置 % 将路径替换为实际存储观测文件的文件夹路径 fileFolder = fullfile(pwd, 'observation_files'); filePattern = fullfile(fileFolder, '*.txt'); txtFiles = dir(filePattern); % 初始化结果存储表 result = table([], {}, [], [], [], [], ... 'VariableNames', {'FileID','ObsTime','Lon_deg','Lat_deg','Alt_km','SolarDepAngle_deg'}); %% 逐文件处理 for fileIdx = 1:length(txtFiles) filePath = fullfile(txtFiles(fileIdx).folder, txtFiles(fileIdx).name); fid = fopen(filePath, 'r'); if fid == -1 warning('文件 %s 打开失败,已跳过', txtFiles(fileIdx).name); continue end % 第一轮扫描:定位所有观测记录的起始行,统计总记录数 recordStartLine = []; lineIdx = 0; tline = fgetl(fid); while ischar(tline) lineIdx = lineIdx + 1; if contains(tline, 'Observation Time:') recordStartLine = [recordStartLine, lineIdx]; end tline = fgetl(fid); end totalRecord = length(recordStartLine); if totalRecord < 1 fclose(fid); warning('文件 %s 未检测到有效观测记录,已跳过', txtFiles(fileIdx).name); continue end % 生成需要提取的记录索引,自动过滤超出总记录数的索引 targetIdx = unique([1, min(4, totalRecord), totalRecord]); % 第二轮扫描:回到文件开头,仅提取目标记录的指定字段 frewind(fid); lineIdx = 0; currentRecordIdx = 0; inTargetRecord = false; tempStruct = struct(); tline = fgetl(fid); while ischar(tline) lineIdx = lineIdx + 1; % 检测是否进入目标记录块 if ismember(lineIdx, recordStartLine) currentRecordIdx = currentRecordIdx + 1; if ismember(currentRecordIdx, targetIdx) inTargetRecord = true; tempStruct = struct(); % 提取观测时间 timeMatch = regexp(tline, 'Observation Time:\s*(.*?)\s*UT', 'tokens', 'once'); tempStruct.ObsTime = timeMatch{1}; else inTargetRecord = false; end end % 仅在目标记录块内提取所需字段 if inTargetRecord % 提取经度 if contains(tline, '---> Lon:') lonMatch = regexp(tline, 'Lon:\s*([\d\.]+)\s*deg', 'tokens', 'once'); tempStruct.Lon_deg = str2double(lonMatch{1}); end % 提取纬度 if contains(tline, '---> Lat:') latMatch = regexp(tline, 'Lat:\s*([\d\.]+)\s*deg', 'tokens', 'once'); tempStruct.Lat_deg = str2double(latMatch{1}); end % 提取高度 if contains(tline, '---> Alt:') altMatch = regexp(tline, 'Alt:\s*([\d\.]+)\s*km', 'tokens', 'once'); tempStruct.Alt_km = str2double(altMatch{1}); end % 提取真太阳俯角,提取完成后存入结果表 if contains(tline, 'True Solar Depression Angle:') sdaMatch = regexp(tline, 'True Solar Depression Angle:\s*([\d\.]+)\s*deg', 'tokens', 'once'); tempStruct.SolarDepAngle_deg = str2double(sdaMatch{1}); newRow = {fileIdx, tempStruct.ObsTime, tempStruct.Lon_deg, ... tempStruct.Lat_deg, tempStruct.Alt_km, tempStruct.SolarDepAngle_deg}; result = [result; newRow]; inTargetRecord = false; end end tline = fgetl(fid); end fclose(fid); end %% 结果处理 % 命令行打印提取结果 disp('指定记录提取完成:'); disp(result); % 取消注释可将结果导出为CSV文件 % writetable(result, 'extracted_target_data.csv');
使用说明
- 代码默认匹配
.txt后缀的文本文件,若观测文件为其他后缀,可修改filePattern中的后缀参数 - 正则匹配规则自动忽略字段前后空格、
+/-误差标注等无关内容,直接提取有效数值,无需额外清洗 - 当单文件内观测记录总数不足4条时,代码会自动适配,不会抛出错误,仅提取存在的对应序号记录
- 逐行读取的内存占用稳定在1MB以内,可适配单文件GB级的超长文本处理
内容的提问来源于stack exchange,提问作者refinedbread
相关产品推荐
相关产品推荐

