You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从txt提取指定位置字符串写入DataFrame列的问题排查

问题解决:从文本文件提取数据到Pandas DataFrame

问题描述

需要从output.txt每行的第13位到行尾提取特定字符串,存入预先创建的包含name、altitude、latitude、longitude四列的空DataFrame。现有代码执行后DataFrame仍为空:

创建空DataFrame的代码:

import pandas as pd
cameras = pd.DataFrame(columns=['name', 'altitude', 'latitude', 'longitude'])

提取数据的代码:

with open('output.txt','r') as f:
    for line in f.readlines():
        if line.startswith('name'):
            cameras['name'] = line[13:-1]
        if line.startswith('NN'):
            cameras['altitude'] = line[13:-1]
        if line.startswith('lat'):
            cameras['latitude'] = line[13:-1]
        if line.startswith('lon'):
            cameras['longitude'] = line[13:-1]

问题原因

  1. 空DataFrame无任何行数据,直接给列赋值cameras['name'] = ...仅会设置列的单个值,但不会新增行,最终DataFrame仍显示为空。
  2. 原代码未考虑name/NN/lat/lon是一组对应DataFrame一行的字段,无法正确聚合单条完整记录。

解决方法

方法1:先收集所有记录的字典列表,再生成DataFrame(推荐,效率更高)

import pandas as pd

records = []
current_record = {}

with open('output.txt','r') as f:
    for line in f:
        line = line.strip()  # 去除换行符与首尾空白
        if not line:
            continue  # 跳过空行
        # 提取第13位到末尾的内容,注意Python索引从0开始,第13位对应索引12
        content = line[12:] if len(line) >=13 else ''
        if line.startswith('name'):
            current_record['name'] = content
        elif line.startswith('NN'):
            current_record['altitude'] = content
        elif line.startswith('lat'):
            current_record['latitude'] = content
        elif line.startswith('lon'):
            current_record['longitude'] = content
            # 凑齐四个字段后,添加到记录列表并重置当前记录
            records.append(current_record)
            current_record = {}

# 生成目标DataFrame
cameras = pd.DataFrame(records, columns=['name', 'altitude', 'latitude', 'longitude'])

方法2:逐行追加记录到DataFrame

import pandas as pd

cameras = pd.DataFrame(columns=['name', 'altitude', 'latitude', 'longitude'])
current_record = {}

with open('output.txt','r') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        content = line[12:] if len(line) >=13 else ''
        if line.startswith('name'):
            current_record['name'] = content
        elif line.startswith('NN'):
            current_record['altitude'] = content
        elif line.startswith('lat'):
            current_record['latitude'] = content
        elif line.startswith('lon'):
            current_record['longitude'] = content
            # 将当前完整记录追加到DataFrame
            cameras = pd.concat([cameras, pd.DataFrame([current_record])], ignore_index=True)
            current_record = {}

注意事项

  • 确认切片索引:Python字符串索引从0开始,第13位字符对应索引12,原代码的line[13:-1]是从第14位开始提取,需根据output.txt的实际格式调整。
  • 异常处理:添加空行判断避免无效数据,同时判断字符串长度防止索引越界。

内容的提问来源于stack exchange,提问作者user13686580

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 23:06:09