开发查重软件时批量转换多格式文件为TXT需保留原文件名的技术问询
Got it, let's sort this out—your current code is renaming source files to UUIDs first, which is why you're losing the original filenames. We can fix this by ditching that unnecessary rename step and using the original filename directly for your output .txt files.
What's Causing the Issue?
Your first loop renames every file in the source directory to a UUID, overwriting the original filenames entirely. That's why when you generate the .txt files later, you're stuck using those UUIDs instead of the original names. We don't need to modify the source files at all to extract their text!
Fixed Code
import os import textract # Define your directories (removed redundant os.path.join with os.getcwd() since you're using absolute paths) source_directory = "C:/Users/syedm/Desktop/Study/FOUNDplag/Plagiarism-checker-Python/mainfolder" training_directory = "C:/Users/syedm/Desktop/Study/FOUNDplag/Plagiarism-checker-Python/trainingdata" # Loop through each original file in the source directory for filename in os.listdir(source_directory): # Skip directories to avoid processing folders by mistake if os.path.isdir(os.path.join(source_directory, filename)): continue # Split original filename from its extension (e.g., "report" from "report.pdf") original_name, extension = os.path.splitext(filename) # Create target .txt filename using the original file's base name dest_file_path = f"{original_name}.txt" dest_full_path = os.path.join(training_directory, dest_file_path) try: # Extract text directly from the original file (no renaming needed!) content = textract.process(os.path.join(source_directory, filename)) # Use a context manager to handle file writing (safer than manual close) with open(dest_full_path, "wb") as write_text_file: write_text_file.write(content) print(f"Successfully converted {filename} to {dest_file_path}") except Exception as e: print(f"Failed to process {filename}: {str(e)}")
Key Changes Explained:
- Removed the UUID rename loop: We leave your source files untouched—their original names stay intact.
- Directly use original filenames for output: We take the base name of the original file (without its extension) and append
.txtto create the target file name. - Added error handling: Wrapped processing in a try/except block to catch issues like unsupported file types or permission errors, with clear feedback.
- Skipped directories: Added a check to avoid trying to process folders in the source directory.
- Used context manager for file writing: This ensures the file is properly closed even if an error occurs, which is more reliable than manual
close()calls.
Quick Optional Improvement:
If you have multiple files with the same base name (e.g., notes.pdf and notes.docx), the second conversion will overwrite the first .txt file. To prevent this, you could include the original extension in the target name:
dest_file_path = f"{original_name}_{extension.lstrip('.')}.txt" # This would create "notes_pdf.txt" and "notes_docx.txt"
内容的提问来源于stack exchange,提问作者Ahsan Ul Haq

