You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取文件至多列表失败:x2/y2/x3/y3为空问题求助

问题排查与解决方案:多组列表读取失败仅第一组有数据

我来帮你梳理下代码里的问题,以及对应的解决办法:

核心问题原因

  • 文件指针偏移导致后续读取为空
    当你第一次调用f.readlines()[0:20]时,已经把文件指针移动到了第20行的位置。后续再调用f.readlines()时,会从当前指针位置开始读取,而此时已经没有剩余内容(或者剩余内容不在你指定的切片范围内),所以lines2和lines3都是空列表,自然x2、y2等也没有数据。

  • 切片索引错误
    假设每组数据是20行,你写的[21:40]其实是从第22行开始读取(Python切片是左闭右开,且索引从0开始),正确的第二组应该是[20:40],第三组是[40:60]——不过这个问题在文件指针的问题解决后才会显现。

  • 数据处理不严谨
    用line.split()[1][:4]来保留两位小数非常不可靠,比如如果数值是0.099,取前4位会得到0.09,丢失精度;而且读取的x、y都是字符串类型,绘图时可能会出现排序异常或者显示问题。

解决方案

方案1:修正硬编码行数的读取逻辑

先一次性读取所有行,再对读取到的列表进行切片,避免文件指针的问题,同时修正索引和数据类型:

import numpy as np
import matplotlib.pyplot as plt

with open('kmt_oa.txt') as f:
    # 一次性读取所有行,并过滤掉空行
    all_lines = [line.strip() for line in f if line.strip()]
    
    # 第一组:前20行
    lines1 = all_lines[0:20]
    # 转换为float类型,保留两位小数
    x1 = [round(float(line.split()[1]), 2) for line in lines1]
    y1 = [int(line.split()[2]) for line in lines1]
    
    # 第二组:第20-40行
    lines2 = all_lines[20:40]
    x2 = [round(float(line.split()[1]), 2) for line in lines2]
    y2 = [int(line.split()[2]) for line in lines2]
    
    # 第三组:第40-60行
    lines3 = all_lines[40:60]
    x3 = [round(float(line.split()[1]), 2) for line in lines3]
    y3 = [int(line.split()[2]) for line in lines3]

ax1 = plt.subplot()
ax1.scatter(x1, y1, color='r', label='Table size of 20')
ax1.scatter(x2, y2, color='g', label='Table size of 21')
ax1.scatter(x3, y3, color='b', label='Table size of 19')
ax1.set_xlabel('Load Factor')
ax1.set_ylabel('Number of Collisions')
ax1.set_title("Key Mod Tablesize & Open Addressing")
plt.legend(loc='upper left')
plt.show()

方案2:根据行号自动分组(更符合你的原始需求)

如果你希望当第一列行号重置为1时自动生成新列表,而不是硬编码行数,这个方案更灵活,适合每组行数不固定的情况:

import numpy as np
import matplotlib.pyplot as plt

x_groups = []
y_groups = []
current_x = []
current_y = []

with open('kmt_oa.txt') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        parts = line.split()
        row_num = int(parts[0])
        x_val = round(float(parts[1]), 2)
        y_val = int(parts[2])
        
        # 当行号为1时,说明是新的一组,先保存当前组(如果有的话)
        if row_num == 1 and current_x:
            x_groups.append(current_x)
            y_groups.append(current_y)
            current_x = []
            current_y = []
        current_x.append(x_val)
        current_y.append(y_val)
    # 保存最后一组数据
    if current_x:
        x_groups.append(current_x)
        y_groups.append(current_y)

# 绘图(这里假设你有3组,对应三个标签)
ax1 = plt.subplot()
colors = ['r', 'g', 'b']
labels = ['Table size of 20', 'Table size of 21', 'Table size of 19']
for i in range(len(x_groups)):
    ax1.scatter(x_groups[i], y_groups[i], color=colors[i], label=labels[i])

ax1.set_xlabel('Load Factor')
ax1.set_ylabel('Number of Collisions')
ax1.set_title("Key Mod Tablesize & Open Addressing")
plt.legend(loc='upper left')
plt.show()

额外说明

  • 用round(float(val), 2)可以准确保留两位小数,避免字符串切片的不确定性。
  • 将y值转换为int类型,因为碰撞次数是整数,绘图时更合理。
  • 自动分组的方案不需要提前知道每组的行数,完全符合你“行号重置为1时生成新列表”的需求,扩展性更好。

内容的提问来源于stack exchange,提问作者zmjackson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:28:32