从txt提取指定位置字符串写入DataFrame列的问题排查
问题解决:从文本文件提取数据到Pandas DataFrame
问题描述
需要从output.txt每行的第13位到行尾提取特定字符串,存入预先创建的包含name、altitude、latitude、longitude四列的空DataFrame。现有代码执行后DataFrame仍为空:
创建空DataFrame的代码:
import pandas as pd cameras = pd.DataFrame(columns=['name', 'altitude', 'latitude', 'longitude'])
提取数据的代码:
with open('output.txt','r') as f: for line in f.readlines(): if line.startswith('name'): cameras['name'] = line[13:-1] if line.startswith('NN'): cameras['altitude'] = line[13:-1] if line.startswith('lat'): cameras['latitude'] = line[13:-1] if line.startswith('lon'): cameras['longitude'] = line[13:-1]
问题原因
- 空DataFrame无任何行数据,直接给列赋值
cameras['name'] = ...仅会设置列的单个值,但不会新增行,最终DataFrame仍显示为空。 - 原代码未考虑
name/NN/lat/lon是一组对应DataFrame一行的字段,无法正确聚合单条完整记录。
解决方法
方法1:先收集所有记录的字典列表,再生成DataFrame(推荐,效率更高)
import pandas as pd records = [] current_record = {} with open('output.txt','r') as f: for line in f: line = line.strip() # 去除换行符与首尾空白 if not line: continue # 跳过空行 # 提取第13位到末尾的内容,注意Python索引从0开始,第13位对应索引12 content = line[12:] if len(line) >=13 else '' if line.startswith('name'): current_record['name'] = content elif line.startswith('NN'): current_record['altitude'] = content elif line.startswith('lat'): current_record['latitude'] = content elif line.startswith('lon'): current_record['longitude'] = content # 凑齐四个字段后,添加到记录列表并重置当前记录 records.append(current_record) current_record = {} # 生成目标DataFrame cameras = pd.DataFrame(records, columns=['name', 'altitude', 'latitude', 'longitude'])
方法2:逐行追加记录到DataFrame
import pandas as pd cameras = pd.DataFrame(columns=['name', 'altitude', 'latitude', 'longitude']) current_record = {} with open('output.txt','r') as f: for line in f: line = line.strip() if not line: continue content = line[12:] if len(line) >=13 else '' if line.startswith('name'): current_record['name'] = content elif line.startswith('NN'): current_record['altitude'] = content elif line.startswith('lat'): current_record['latitude'] = content elif line.startswith('lon'): current_record['longitude'] = content # 将当前完整记录追加到DataFrame cameras = pd.concat([cameras, pd.DataFrame([current_record])], ignore_index=True) current_record = {}
注意事项
- 确认切片索引:Python字符串索引从0开始,第13位字符对应索引12,原代码的
line[13:-1]是从第14位开始提取,需根据output.txt的实际格式调整。 - 异常处理:添加空行判断避免无效数据,同时判断字符串长度防止索引越界。
内容的提问来源于stack exchange,提问作者user13686580
相关产品推荐
相关产品推荐

