You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

开发查重软件时批量转换多格式文件为TXT需保留原文件名的技术问询

Solution to Keep Original Filenames When Converting Files to .txt

Got it, let's sort this out—your current code is renaming source files to UUIDs first, which is why you're losing the original filenames. We can fix this by ditching that unnecessary rename step and using the original filename directly for your output .txt files.

What's Causing the Issue?

Your first loop renames every file in the source directory to a UUID, overwriting the original filenames entirely. That's why when you generate the .txt files later, you're stuck using those UUIDs instead of the original names. We don't need to modify the source files at all to extract their text!

Fixed Code

import os
import textract

# Define your directories (removed redundant os.path.join with os.getcwd() since you're using absolute paths)
source_directory = "C:/Users/syedm/Desktop/Study/FOUNDplag/Plagiarism-checker-Python/mainfolder"
training_directory = "C:/Users/syedm/Desktop/Study/FOUNDplag/Plagiarism-checker-Python/trainingdata"

# Loop through each original file in the source directory
for filename in os.listdir(source_directory):
    # Skip directories to avoid processing folders by mistake
    if os.path.isdir(os.path.join(source_directory, filename)):
        continue
    
    # Split original filename from its extension (e.g., "report" from "report.pdf")
    original_name, extension = os.path.splitext(filename)
    
    # Create target .txt filename using the original file's base name
    dest_file_path = f"{original_name}.txt"
    dest_full_path = os.path.join(training_directory, dest_file_path)
    
    try:
        # Extract text directly from the original file (no renaming needed!)
        content = textract.process(os.path.join(source_directory, filename))
        
        # Use a context manager to handle file writing (safer than manual close)
        with open(dest_full_path, "wb") as write_text_file:
            write_text_file.write(content)
        
        print(f"Successfully converted {filename} to {dest_file_path}")
    except Exception as e:
        print(f"Failed to process {filename}: {str(e)}")

Key Changes Explained:

  • Removed the UUID rename loop: We leave your source files untouched—their original names stay intact.
  • Directly use original filenames for output: We take the base name of the original file (without its extension) and append .txt to create the target file name.
  • Added error handling: Wrapped processing in a try/except block to catch issues like unsupported file types or permission errors, with clear feedback.
  • Skipped directories: Added a check to avoid trying to process folders in the source directory.
  • Used context manager for file writing: This ensures the file is properly closed even if an error occurs, which is more reliable than manual close() calls.

Quick Optional Improvement:

If you have multiple files with the same base name (e.g., notes.pdf and notes.docx), the second conversion will overwrite the first .txt file. To prevent this, you could include the original extension in the target name:

dest_file_path = f"{original_name}_{extension.lstrip('.')}.txt"
# This would create "notes_pdf.txt" and "notes_docx.txt"

内容的提问来源于stack exchange,提问作者Ahsan Ul Haq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 08:33:14