You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将目录、文件、MP4及空目录写入CSV?重复输出问题求助

Fixing Duplicate Entries in Your CSV Output

Hey there! The duplicate entries are happening because of redundant nested os.walk loops and unnecessary repeated data collection. Let's break down what's wrong and fix it:

What's Causing the Duplicates?

  • You're running os.walk three separate times (outer loop, inner loop, and inside the empty directory collection), which means you're iterating over the same directory structure multiple times.
  • For every file in the outer loop, you loop through all MP4 files and all empty directories, writing a row for each combination—this creates way more rows than needed.
  • You're re-collecting the full list of empty directories every time you find an MP4 file, which is inefficient and leads to redundant data.

Corrected Code

Here's a revised version that collects all necessary data once, then writes clean, non-duplicate rows to your CSV:

import csv
import os
import sys

def main():
    # Ensure the user provides a root directory
    if len(sys.argv) != 2:
        print("Usage: python script.py <root_directory>")
        sys.exit(1)
    root_dir = sys.argv[1]
    
    # Pre-collect all MP4 files in one pass
    mp4_file_list = []
    for root, _, files in os.walk(root_dir):
        for file in files:
            if file.lower().endswith(".mp4"):  # Lowercase to catch .MP4 too
                mp4_file_list.append(os.path.join(root, file))
    
    # Pre-collect all empty directories in one pass
    empty_dir_list = []
    for dirpath, dirnames, filenames in os.walk(root_dir):
        # A directory is empty if it has no subdirectories and no files
        if not dirnames and not filenames:
            empty_dir_list.append(dirpath)
    
    # Write to CSV with clean data
    with open('Files.csv', 'w', newline='', encoding='utf-8') as f:
        fieldnames = ['Foldername', 'Filename', 'mp4filelist', 'Empty folders']
        writer = csv.DictWriter(f, fieldnames=fieldnames)
        writer.writeheader()
        
        # Loop through each file once to write its details
        for dirpath, _, files in os.walk(root_dir):
            for filename in files:
                # Convert lists to readable comma-separated strings
                mp4_str = ', '.join(mp4_file_list) if mp4_file_list else 'None'
                empty_dir_str = ', '.join(empty_dir_list) if empty_dir_list else 'None'
                
                writer.writerow({
                    'Foldername': dirpath,
                    'Filename': filename,
                    'mp4filelist': mp4_str,
                    'Empty folders': empty_dir_str
                })

if __name__ == "__main__":
    main()

Key Improvements

  • Single-pass data collection: We gather all MP4 files and empty directories once at the start, avoiding redundant work.
  • No nested loops: We only iterate over the directory structure once per task, eliminating duplicate row generation.
  • Readable CSV output: Lists are converted to comma-separated strings instead of raw Python list syntax, making the CSV easier to read.
  • Error handling: Added a check for the command-line argument to guide the user if they forget to provide the root directory.
  • Case insensitivity: Checks for .mp4 in lowercase to catch files with .MP4 extensions too.

If you intended to have separate rows for MP4 files or empty directories instead of including their full lists in every row, let me know and we can adjust the structure further!

内容的提问来源于stack exchange,提问作者Haider

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 17:02:38