在R中读取AGS格式文件并将内部各表转为独立数据框的技术咨询
解决思路
- 读取AGS文件全部内容,按连续空行切割得到独立的表格块,过滤无效空内容
- 每个有效表格块遵循固定结构:首行是GROUP标记及表名,第二行为列名,第三行为对应列单位,第四行及之后为观测数据
- 逐行拆分处理字段(自动识别双引号包裹的内容,避免内容内逗号干扰拆分),将每个表格块转换为独立DataFrame,所有表格按表名存入字典方便调用
Python实现代码
首先安装依赖库:pip install pandas
完整实现代码:
import pandas as pd import csv from io import StringIO from itertools import groupby def parse_ags_row(row_str): """用csv模块解析单行AGS内容,避免内容中包含逗号导致拆分错误""" reader = csv.reader(StringIO(row_str), delimiter=',', quotechar='"') return next(reader) def read_ags(file_path, encoding='utf-8'): # 读取文件所有行 with open(file_path, 'r', encoding=encoding) as f: lines = [line.rstrip('\n') for line in f] # 按连续空行拆分表格块 table_blocks = [] for is_empty, group in groupby(lines, key=lambda x: x.strip() == ''): if not is_empty: table_blocks.append(list(group)) ags_tables = {} for block in table_blocks: # 跳过非GROUP开头的无效块 if not (block and parse_ags_row(block[0])[0] == 'GROUP'): continue # 提取表名 table_name = parse_ags_row(block[0])[1] # 提取列名 columns = parse_ags_row(block[1]) # 提取列单位,存入DataFrame属性 units = parse_ags_row(block[2]) # 处理数据行 data = [] for row in block[3:]: data.append(parse_ags_row(row)) # 生成DataFrame df = pd.DataFrame(data, columns=columns) df.attrs['units'] = dict(zip(columns, units)) ags_tables[table_name] = df return ags_tables
使用方式:
# 替换为你的本地AGS文件路径 ags_result = read_ags("./example.ags", encoding='utf-8') # 按表名获取对应DataFrame,例如获取LOCA表 loca_df = ags_result.get("LOCA") # 查看某列的单位 print(loca_df.attrs['units'].get("LOCA_ID"))
注意事项
- 若读取时出现编码错误,可尝试将encoding参数替换为
gbk、latin-1等适配文件编码 - 所有表的单位都存储在对应DataFrame的attrs属性中,可根据需要提取使用
- 如果需要将DataFrame导出为csv、excel等格式,直接调用pandas自带的
to_csv、to_excel方法即可
内容的提问来源于stack exchange,提问作者erdi
相关产品推荐
相关产品推荐

