You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取CSV数据:筛选至Asia行并保留South对应后代行

CSV指定片段提取实现方案

原有脚本问题

你原来的脚本只能拿到前14行的核心原因很简单:遍历到第一列值为Asia的行时直接执行break终止了整个循环,Asia节点之后的所有内容根本没进入处理流程,自然提取不到需要的South后代行。另外提前调用os.remove删结果文件完全没必要,Python用w模式写文件时会自动覆盖旧内容。

核心逻辑

整个提取流程靠状态位控制即可,不需要复杂操作:

  • 初始状态:逐行写入内容,直到碰到第一列为Asia的行,切换状态,这行本身不写入结果
  • 查找South块状态:继续往后遍历,跳过无关行,直到定位到第一列值为South的行,写入这行后切换到提取South后代的状态
  • 提取South后代状态:持续写入后续行,直到碰到第一列非空、且值不属于South子项的同级/上级节点,直接终止遍历即可。这类层级结构的CSV里,子项行的第一列一般为空,靠第二列存内容,用这个特征判断后代边界准确率很高。

修正后可直接运行的代码

import csv

source_path = "C:/Users/Documents/Python Scripts/countries_source.csv"
result_path = "C:/Users/Documents/Python Scripts/mycountries.csv"

# 一次性读取所有源文件行
with open(source_path, "r", encoding="utf-8") as source_file:
    all_rows = list(csv.reader(source_file))

with open(result_path, "w", newline="", encoding="utf-8") as result_file:
    writer = csv.writer(result_file)
    # 状态标记
    has_passed_asia = False
    has_found_south = False

    for row in all_rows:
        # 跳过空行避免索引报错
        if not row:
            continue
        first_col_val = row[0].strip()

        # 状态1:还没到Asia行,正常写入
        if not has_passed_asia:
            if first_col_val == "Asia":
                has_passed_asia = True
                continue
            writer.writerow(row)
            continue

        # 状态2:过了Asia,还没找到South行
        if not has_found_south:
            if first_col_val == "South":
                has_found_south = True
                writer.writerow(row)
            continue

        # 状态3:正在提取South后代,碰到第一列有值的非子项行就停止
        if first_col_val:
            break
        writer.writerow(row)

注意点

  • 代码统一指定了utf-8编码,避免不同环境打开CSV时出现乱码
  • 加了空行判断,源文件里如果有空行不会触发索引越界报错
  • 如果你的South块终止边界不是“第一列非空”,只需要调整状态3里的终止判断条件即可,整体逻辑不用改。

内容的提问来源于stack exchange,提问作者Vikas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 18:45:42