You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本提取目录及子目录文本文件邮箱为空无报错排查求助

问题描述

需要从包含大量子目录的目录下所有文本文件中提取全部邮箱地址,手动操作工作量极大,因此编写了以下Python脚本实现自动化处理。但执行脚本后输出的是空数组,且未显示任何错误,恳请帮忙指出问题所在。

# Import Module
import os
import re
  
# Folder Path
path = "pat to the root directory"
  
# Change the directory
os.chdir(path)

#create list and index to add the emails
new_list = []
idx = 0

# I create a method to add all email address from within the subdirectories to add 
  them to an array
def read_text_file(file_path):
    with open(file_path, 'r') as f:
        emails = re.findall(r"[a-z0-9\.\-+_]+@[a-z0-9\.\-+_]+\.[a-z]+", str(f))
        new_list.insert(idx, emails)
        idx + 1
     
# iterate through all file and call the method from above
for file in os.listdir():
    # Check whether file is in text format or not
    if file.endswith(".txt"):
        p = f"{path}\{file}"
        # call read text file function
        read_text_file(p)

#print the array    
print (new_list) 
问题分析与修复方案

你的脚本存在几个关键问题,直接导致输出空数组:

  • 未递归遍历子目录:os.listdir()仅读取当前目录下的文件和子目录名称,不会深入子目录内部处理文件,所有子目录里的文本文件完全没被读取。
  • 错误读取文件内容:str(f)是把文件对象本身转为字符串(类似<_io.TextIOWrapper name='xxx.txt' mode='r' encoding='utf-8'>),而不是读取文件里的实际内容,正则自然匹配不到任何邮箱。
  • 全局变量idx未正确修改:函数内的idx + 1只做了加法运算,没有重新赋值给idx;同时修改全局变量需要声明global idx,否则函数内的idx会被视为局部变量。另外完全没必要用insert和索引,直接用append或extend更简洁。
  • 正则表达式不匹配大写邮箱:原正则只匹配小写字母的邮箱,若文本中有大写字母开头的邮箱(比如John@Example.com)会被漏掉,需要添加忽略大小写的标志。
  • 路径拼接存在风险:用f"{path}\{file}"拼接路径可能因转义字符导致错误,建议使用os.path.join()保证跨平台兼容性。

修正后的完整脚本

import os
import re

# 根目录路径,替换为你的实际路径
root_path = "path to the root directory"

new_list = []

def extract_emails_from_file(file_path):
    try:
        with open(file_path, 'r', encoding='utf-8') as f:
            content = f.read()
            # 添加re.IGNORECASE匹配大小写邮箱,优化正则规则
            emails = re.findall(r"[a-zA-Z0-9\.\-+_]+@[a-zA-Z0-9\.\-+_]+\.[a-zA-Z]+", content, flags=re.IGNORECASE)
            if emails:
                new_list.extend(emails)
    except Exception as e:
        print(f"读取文件出错: {file_path}, 错误信息: {e}")

# 递归遍历所有子目录和文件
for dirpath, _, filenames in os.walk(root_path):
    for filename in filenames:
        if filename.endswith(".txt"):
            file_path = os.path.join(dirpath, filename)
            extract_emails_from_file(file_path)

# 去重并打印结果(可选,根据需求决定是否去重)
unique_emails = list(set(new_list))
print(unique_emails)

内容的提问来源于stack exchange,提问作者snorkel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:45:47