You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本转CSV遇List index out of range错误求解决

求助:Python文本转CSV时出现索引越界错误

我尝试用Python把特定格式的文本文件转换成CSV,但运行代码时一直报“List index out of range”错误,实在搞不定,求帮忙解决。

文本文件内容

DC Number: V70909
Name: A, SASHWIN
Race: ALL OTHERS/UNKNOWN
Sex: MALE
Birth Date: 04/27/1996
Custody: CLOSE
Release Date: 09/23/2021
Aliases:SASHWIN A, SASHWIN ASHOK, SASHWIN ASOKAN, A SASHWIN, ASOKAN SASHWINDC Number: 522180
Name: AALIM, MIKAIL N
Race: BLACK
Sex: MALE
Birth Date: 10/12/1950
Custody: COMMUNITY
Release Date: 08/05/2005
Aliases:MIKAIL AALIM, MIKAIL N AALIM, MIKAIL NAJI AALIM, LORENZO ANDERSON, LORENZO KENNETH ANDERSON, LEROY WILLIAMS ANGUS, ANGUS WILLIAMSDC Number: Y11193
Name: AALTO, MARK K
Race: WHITE
Sex: MALE
Birth Date: 06/29/1968
Custody: MEDIUM
Release Date: 05/31/2013
Aliases:MARK AALTO, MARK K AALTO, MARK KENNETH AALTODC Number: K87086
Name: AAMIR, OMAR T
Race: BLACK
Sex: MALE
Birth Date: 06/30/1992
Custody: MINIMUM
Release Date: 08/01/2019
Aliases:OMAR T AAMIR, OMAR TERRANCE AAMIR, KADEEM THOMPSONDC Number: 138198
Name: AANENSEN, JOHN A

期望的CSV格式

CSV需包含以下列:DC Number、Name、Race、Sex、Birth Date、Custody、Release Date、Aliases,每行对应一个人员的完整信息。

我的实现代码

import csv
file_name = "readme.txt"
list_csv = []

# add header to csvlist_csv
header = ["DC Number", "Name", "Race", "Sex", "Birth Date", "Custody", "Release Date", "Aliases"]

# Write to csv file
# opening the csv file in 'a+' mode
csv_file_name = "my_csv.csv"
file = open(csv_file_name, 'w+', newline='\n')
write = csv.writer(file)
write.writerow(header)
# write.writerow(temp_list)

# headers
with open(file_name) as file_reader:
    lines = file_reader.readlines()

    for line in range(0, len(lines), 8):
        temp_list = []
        DC_Number = lines[line].split(":")[1].strip()
        name = lines[line + 1].split(":")[1].strip()
        race = lines[line + 2].split(":")[1].strip()
        sex = lines[line + 3].split(":")[1].strip()
        birth_date = lines[line + 4].split(":")[1].strip()
        custody = lines[line + 5].split(":")[1].strip()
        release_date = lines[line + 6].split(":")[1].strip()
        aliases = lines[line + 7].split(":")[1].strip()
        temp_list.append(DC_Number)
        temp_list.append(name)
        temp_list.append(race)
        temp_list.append(sex)
        temp_list.append(birth_date)
        temp_list.append(custody)
        temp_list.append(release_date)
        temp_list.append(aliases)
        # write content to csv

        write.writerow(temp_list)

file.close()  # close the file

问题原因及解决方法

错误原因

  1. 文本格式问题:原文本中Aliases行的结尾和下一个DC Number连在一起,导致readlines()读取的行并不是每8行对应一个完整人员信息,循环按步长8取索引时必然会越界。
  2. 不完整数据:最后一条人员信息只有DC Number和Name,缺少后续字段,按固定索引取值会报错。

修正后的代码

import csv

file_name = "readme.txt"
csv_file_name = "my_csv.csv"
header = ["DC Number", "Name", "Race", "Sex", "Birth Date", "Custody", "Release Date", "Aliases"]

# 读取整个文本内容
with open(file_name, 'r') as f:
    content = f.read()

# 按"DC Number:"分割,获取每个人员的信息块
person_blocks = content.split("DC Number:")[1:]  # 去掉第一个空元素

with open(csv_file_name, 'w', newline='') as csv_file:
    writer = csv.writer(csv_file)
    writer.writerow(header)

    for block in person_blocks:
        # 初始化字段默认值
        person_data = {key: "" for key in header}
        # 先给DC Number赋值
        person_data["DC Number"] = block.split("\n")[0].strip()
        # 处理剩余行
        lines = block.split("\n")[1:]
        for line in lines:
            line = line.strip()
            if not line:
                continue
            # 分割字段名和值,只分割一次
            if ":" in line:
                key_part, value_part = line.split(":", 1)
                key = key_part.strip()
                value = value_part.strip()
                # 如果是Aliases,可能包含下一个DC Number的开头,需要截断
                if key == "Aliases" and "DC Number:" in value:
                    value = value.split("DC Number:")[0].strip()
                if key in person_data:
                    person_data[key] = value
        # 按header顺序提取值,写入CSV
        writer.writerow([person_data[key] for key in header])

代码说明

  1. 按人员块分割:不再依赖固定行数,而是用DC Number:作为分隔符,直接拆分出每个人员的信息块,避免行对齐问题。
  2. 字段映射:用字典存储每个字段的值,缺失字段自动填充为空字符串,兼容不完整数据。
  3. 处理Aliases的粘连问题:检查Aliases值中是否包含下一个DC Number:,如果有则截断,提取正确的别名内容。

内容的提问来源于stack exchange,提问作者Joe janf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 06:35:28