You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从含未知文件夹的.tar.gz提取.txt文件遇EOFError求助

Fixing EOFError & Correctly Extracting .txt Files from Tar.gz

Hey there! Let's tackle your two main issues: the EOFError popping up, and getting those .txt files extracted directly into your my_path directory without keeping the unknown folder structure.

First: Resolve the EOFError

That error typically stems from one of two scenarios:

  • Your tar.gz file is corrupted or incomplete: Maybe it didn’t finish downloading, or got damaged during creation. Verify this by running this command in your terminal:
    tar -tvzf /path/to/your/my_path/your_archive.tar.gz
    
    If this command throws an error about the archive being corrupted, you’ll need to re-download or re-generate the tar.gz file—broken archives can’t be fixed programmatically.
  • Unnecessary repeated file openings: Your original code calls tarfile.open() twice for the same file, which can cause resource conflicts. We’ll fix this in the revised code below.

Second: Correct Code to Extract .txt Files to my_path

Your initial code only printed .txt file names but didn’t extract them, and it didn’t strip out the unknown parent folder. Here’s a robust, working version:

import tarfile
import os

# Replace these with your actual paths
my_path = "/your/actual/my_path"
archive_filename = "your_archive.tar.gz"
full_archive_path = os.path.join(my_path, archive_filename)

if full_archive_path.endswith('.tar.gz'):
    try:
        # Use 'with' to auto-manage tar file resources
        with tarfile.open(full_archive_path, "r:gz") as tar_archive:
            for file_info in tar_archive:
                # Only process regular .txt files (skip directories)
                if file_info.isfile() and file_info.name.endswith('.txt'):
                    # Strip off the unknown folder path, keep just the filename
                    target_filename = os.path.basename(file_info.name)
                    target_file_path = os.path.join(my_path, target_filename)
                    
                    # Extract content and write directly to my_path
                    with tar_archive.extractfile(file_info) as source_file, open(target_file_path, 'wb') as dest_file:
                        dest_file.write(source_file.read())
                    
                    print(f"Successfully extracted: {target_file_path}")
    except tarfile.ReadError as e:
        print(f"Failed to read the tar archive: {str(e)}")
        print("Double-check that your tar.gz file isn't corrupted (use the tar command mentioned earlier)!")

Key improvements in this code:

  • Uses with statements to ensure files are properly closed after use, avoiding resource leaks.
  • Skips directories with file_info.isfile() to prevent empty folders in my_path.
  • Strips the unknown parent folder using os.path.basename() so .txt files go straight to your target directory.
  • Includes error handling to catch archive issues and give clear, actionable feedback.

内容的提问来源于stack exchange,提问作者Marta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:35:38