You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Matlab如何提取长文本文件指定字段及指定迭代轮次数据

Matlab提取超长观测文本指定字段与指定迭代记录

针对重复出现的固定结构观测记录块,采用逐行扫描+正则匹配的方式实现大文件低内存占用提取,仅保留需要的坐标、真太阳俯角字段,以及第1次、第4次、最后一次迭代的记录。

单条观测记录参考格式如下:

Target Name: Tlil@z=0070km
 Observation Time: 2022-Dec-23 at 09:35:48.500 UT
---> Lon: 4.990 deg +/- .030 deg
---> Lat: 55.5 deg +/- 0.054 deg
---> Alt: 84.962 km +/- 35.577 meters
True Solar Depression Angle:  4.1 deg
Slant Range: 590.253 km

实现逻辑

  • 遍历指定目录下所有观测文本文件,采用逐行读取模式,避免一次性加载全量超大文件导致内存溢出
  • 以Observation Time字段作为单条观测迭代记录的起始标记,先扫描全文件统计总迭代次数,定位需要提取的第1、4、最后一条记录的行位置
  • 对目标记录行,通过正则匹配提取经度、纬度、高度、真太阳俯角4个字段的有效数值,自动跳过误差标注、斜距字段及其他无关冗余内容
  • 提取结果自动结构化存储,支持直接导出为表格文件

实现代码

%% 初始化配置
% 将路径替换为实际存储观测文件的文件夹路径
fileFolder = fullfile(pwd, 'observation_files');
filePattern = fullfile(fileFolder, '*.txt');
txtFiles = dir(filePattern);
% 初始化结果存储表
result = table([], {}, [], [], [], [], ...
    'VariableNames', {'FileID','ObsTime','Lon_deg','Lat_deg','Alt_km','SolarDepAngle_deg'});

%% 逐文件处理
for fileIdx = 1:length(txtFiles)
    filePath = fullfile(txtFiles(fileIdx).folder, txtFiles(fileIdx).name);
    fid = fopen(filePath, 'r');
    if fid == -1
        warning('文件 %s 打开失败,已跳过', txtFiles(fileIdx).name);
        continue
    end
    
    % 第一轮扫描:定位所有观测记录的起始行,统计总记录数
    recordStartLine = [];
    lineIdx = 0;
    tline = fgetl(fid);
    while ischar(tline)
        lineIdx = lineIdx + 1;
        if contains(tline, 'Observation Time:')
            recordStartLine = [recordStartLine, lineIdx];
        end
        tline = fgetl(fid);
    end
    totalRecord = length(recordStartLine);
    if totalRecord < 1
        fclose(fid);
        warning('文件 %s 未检测到有效观测记录,已跳过', txtFiles(fileIdx).name);
        continue
    end
    % 生成需要提取的记录索引,自动过滤超出总记录数的索引
    targetIdx = unique([1, min(4, totalRecord), totalRecord]);
    
    % 第二轮扫描:回到文件开头,仅提取目标记录的指定字段
    frewind(fid);
    lineIdx = 0;
    currentRecordIdx = 0;
    inTargetRecord = false;
    tempStruct = struct();
    
    tline = fgetl(fid);
    while ischar(tline)
        lineIdx = lineIdx + 1;
        % 检测是否进入目标记录块
        if ismember(lineIdx, recordStartLine)
            currentRecordIdx = currentRecordIdx + 1;
            if ismember(currentRecordIdx, targetIdx)
                inTargetRecord = true;
                tempStruct = struct();
                % 提取观测时间
                timeMatch = regexp(tline, 'Observation Time:\s*(.*?)\s*UT', 'tokens', 'once');
                tempStruct.ObsTime = timeMatch{1};
            else
                inTargetRecord = false;
            end
        end
        
        % 仅在目标记录块内提取所需字段
        if inTargetRecord
            % 提取经度
            if contains(tline, '---> Lon:')
                lonMatch = regexp(tline, 'Lon:\s*([\d\.]+)\s*deg', 'tokens', 'once');
                tempStruct.Lon_deg = str2double(lonMatch{1});
            end
            % 提取纬度
            if contains(tline, '---> Lat:')
                latMatch = regexp(tline, 'Lat:\s*([\d\.]+)\s*deg', 'tokens', 'once');
                tempStruct.Lat_deg = str2double(latMatch{1});
            end
            % 提取高度
            if contains(tline, '---> Alt:')
                altMatch = regexp(tline, 'Alt:\s*([\d\.]+)\s*km', 'tokens', 'once');
                tempStruct.Alt_km = str2double(altMatch{1});
            end
            % 提取真太阳俯角,提取完成后存入结果表
            if contains(tline, 'True Solar Depression Angle:')
                sdaMatch = regexp(tline, 'True Solar Depression Angle:\s*([\d\.]+)\s*deg', 'tokens', 'once');
                tempStruct.SolarDepAngle_deg = str2double(sdaMatch{1});
                
                newRow = {fileIdx, tempStruct.ObsTime, tempStruct.Lon_deg, ...
                    tempStruct.Lat_deg, tempStruct.Alt_km, tempStruct.SolarDepAngle_deg};
                result = [result; newRow];
                inTargetRecord = false;
            end
        end
        tline = fgetl(fid);
    end
    fclose(fid);
end

%% 结果处理
% 命令行打印提取结果
disp('指定记录提取完成:');
disp(result);
% 取消注释可将结果导出为CSV文件
% writetable(result, 'extracted_target_data.csv');

使用说明

  • 代码默认匹配.txt后缀的文本文件,若观测文件为其他后缀,可修改filePattern中的后缀参数
  • 正则匹配规则自动忽略字段前后空格、+/-误差标注等无关内容,直接提取有效数值,无需额外清洗
  • 当单文件内观测记录总数不足4条时,代码会自动适配,不会抛出错误,仅提取存在的对应序号记录
  • 逐行读取的内存占用稳定在1MB以内,可适配单文件GB级的超长文本处理

内容的提问来源于stack exchange,提问作者refinedbread

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 00:31:08