You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中异常致for循环中断,使用pass仍无效

问题:遍历Word文档时循环中断,无效文件导致部分有效文档未写入CSV

我编写了一段Python代码,用于遍历目录及子目录,提取Word文档(.docx/.doc)的内容并写入CSV文件。遇到无效文件(如伪装成Word的XML格式文件)时需要跳过,但目前存在循环中断问题:某路径下有6个Word文档,其中1个为无效文件,预期5个有效文档内容写入CSV,但实际仅4个内容被保存。我在异常处理块中使用了pass,但循环仍未继续,想知道问题出在哪。

我的代码

import pandas as pd
import numpy as np
import docx2txt
import csv
import re
import os

def read_files():
    
    
    list_of_lists = []

    full_path =""
    rootdir = '/Users/xxx/Documents/github/xxxx/xxxxx/Financial Statements/'
    
    for subdir, dirs, files in os.walk(rootdir):
            try:
                for file in files:
                    print(os.path.join(subdir, file))

                    content = [" "] * 22
                    
                    if file.endswith(".docx") or file.endswith(".doc"):
                        full_path = os.path.join(subdir, file)
                        print(full_path)
                        text = docx2txt.process(full_path)
                        s = " "
                        for line in text.splitlines():
                        #This will ignore empty/blank lines. 
                            if line != '':
                                s = s+ " " + line
                                s = s.replace(",", "")
                                s = " ".join(s.split())
                                s = s.strip()
                        content.insert(6, s)
                        content.insert(0, subdir)
                        content.insert(1, file)
                        list_of_lists.append(content)
                    else:
                        print("AAAA")
                #print(list_of_lists)

                
            except Exception as error:
                
                print(full_path)
                f = open("exception.txt", "a")
                f.write(full_path + "\n")
                f.write(str(error) + "\n")
                f.close()
                print("An exception occurred", error)
                pass
                
            finally:
                df = pd.DataFrame(list_of_lists)
                #print(df)
                df.to_csv('final.csv', encoding='utf-8')


read_files()

问题原因及解决方法

核心问题:异常捕获范围错误

你把整个内层文件循环都包裹在了try块中,一旦某个文件处理时抛出异常,整个内层for file in files循环会直接终止,跳到except块,该目录下剩下的文件都不会被处理。这就是为什么6个文件里1个出错,只处理了4个(出错的是第2个的话,后面3个都没处理)。pass只是让异常块不做额外操作,但无法让已经中断的循环继续。

修复方案

  1. 缩小异常捕获范围:把try-except移到单个Word文件的处理逻辑内部,这样单个文件出错只会跳过该文件,不影响其他文件的循环执行。
  2. 优化写入CSV的时机:finally块在每个子目录处理完后都会执行一次写入,效率低且没必要,应该把写入CSV的逻辑放在所有文件处理完成后,只执行一次。
  3. 优化字符串处理逻辑:避免在逐行循环里重复执行replace、split等操作,提升处理效率。

修改后的代码

import pandas as pd
import docx2txt
import os

def read_files():
    list_of_lists = []
    rootdir = '/Users/xxx/Documents/github/xxxx/xxxxx/Financial Statements/'
    
    for subdir, dirs, files in os.walk(rootdir):
        for file in files:
            print(os.path.join(subdir, file))
            content = [" "] * 22
            
            if file.endswith(".docx") or file.endswith(".doc"):
                full_path = os.path.join(subdir, file)
                print(full_path)
                
                try:
                    text = docx2txt.process(full_path)
                    # 优化字符串处理:先过滤空行,再合并
                    non_empty_lines = [line.strip() for line in text.splitlines() if line.strip()]
                    s = " ".join(non_empty_lines).replace(",", "")
                    s = " ".join(s.split())  # 合并多余空格
                    
                    content.insert(6, s)
                    content.insert(0, subdir)
                    content.insert(1, file)
                    list_of_lists.append(content)
                except Exception as error:
                    # 捕获单个文件的异常,记录后继续处理下一个文件
                    print(f"处理文件失败: {full_path}, 错误: {error}")
                    with open("exception.txt", "a") as f:
                        f.write(f"{full_path}\n{str(error)}\n\n")
            else:
                print("非Word文件,跳过")
    
    # 所有文件处理完成后,一次性写入CSV
    if list_of_lists:
        df = pd.DataFrame(list_of_lists)
        df.to_csv('final.csv', encoding='utf-8')
    else:
        print("没有找到有效Word文档")

read_files()

关键改动说明

  • 将try-except嵌套到单个Word文件的处理代码中,确保单个文件出错不影响整个目录的文件循环。
  • 使用with语句自动管理文件句柄,避免手动关闭文件的遗漏。
  • 优化字符串处理流程,减少重复操作,提升代码效率。
  • 把CSV写入逻辑移到函数末尾,所有文件处理完成后只执行一次,避免多次写入的冗余操作。

内容的提问来源于stack exchange,提问作者Adam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 11:27:09