在Django Web应用中实现PDF/DOCX转TXT并前置保存时遭遇FileNotFoundError的技术求助
Fixing File Conversion Issue in Django Model's save() Method
Let's tackle this problem step by step—your error comes from a few key mistakes in how you're accessing the uploaded file and handling the conversion logic. Here's what's going wrong and how to fix it:
Key Issues in Your Current Code
- Incorrect File Path:
self.file.nameonly gives the relative filename, not the full path to the uploaded file. Before callingsuper().save(), the file hasn't been saved to yourupload_todirectory yet, so trying to open it via filename fails. Instead, use Django's built-inself.file.open()to access the file object directly. - Page Index Out of Bounds:
pdfreader.numPagesreturns the total number of pages, which are 0-indexed. UsinggetPage(x + 1)will always throw an error because you're trying to access a page that doesn't exist. - Invalid Filename Syntax:
self.file.name.txtis not valid Python syntax. You need to properly replace the file extension to.txtusingos.path.splitext(). - Missing DOCX Handling: Your code only handles PDFs, but you mentioned needing to convert
.docxfiles too.
Fixed Code with PDF + DOCX Support
First, install the required dependencies if you haven't already:
pip install PyPDF2 python-docx
Then update your models.py:
from django.db import models from django.contrib.auth.models import User from django.urls import reverse from django.conf import settings import PyPDF2 from docx import Document import os class FileUpload(models.Model): user = models.ForeignKey(User, on_delete=models.CASCADE, blank=True, null=True) file = models.FileField(upload_to='files') # Optional: Add a field to store the extracted text if you need it in the database extracted_text = models.TextField(blank=True, null=True) def save(self, *args, **kwargs): extracted_content = "" file_ext = os.path.splitext(self.file.name)[1].lower() if file_ext == '.pdf': # Open the uploaded PDF directly using Django's file object with self.file.open('rb') as pdffileobj: pdfreader = PyPDF2.PdfReader(pdffileobj) # Iterate through all pages to extract text for page in pdfreader.pages: extracted_content += page.extract_text() or "" elif file_ext == '.docx': # Handle DOCX files using python-docx with self.file.open('rb') as docxfileobj: doc = Document(docxfileobj) for para in doc.paragraphs: extracted_content += para.text + "\n" # Save the extracted text to the model field (if you added it) self.extracted_text = extracted_content # Optionally: Save the text as a .txt file in the same directory as the uploaded file if extracted_content: txt_filename = os.path.splitext(self.file.name)[0] + '.txt' txt_file_path = os.path.join(settings.MEDIA_ROOT, txt_filename) with open(txt_file_path, 'w', encoding='utf-8') as txt_file: txt_file.write(extracted_content) # Always call the parent save method last to ensure the file is saved to disk super(FileUpload, self).save(*args, **kwargs) def get_absolute_url(self): return reverse('home')
What Changed?
- Direct File Access: Used
self.file.open()to read the uploaded file without needing a filesystem path—this works even before the file is saved to yourupload_todirectory. - Valid Page Iteration: For PDFs, we loop through all pages using
pdfreader.pages(the modern PyPDF2 API—notePdfFileReaderis deprecated in newer versions). - DOCX Handling: Added support for
.docxfiles using thepython-docxlibrary, extracting text from all paragraphs. - Proper TXT Filename: Used
os.path.splitext()to safely replace the original file extension with.txt, and saved the file to yourMEDIA_ROOTdirectory. - Optional Database Storage: Added an
extracted_textfield to store the converted text directly in the database, which is useful if you need to display or search the content later. - Parent Save Last: We call
super().save()after processing to ensure the original file is saved to disk first (if you need to reference its final path).
Notes
- Make sure your Django
MEDIA_ROOTandMEDIA_URLare properly configured insettings.pyso the uploaded files and generated.txtfiles are accessible. - Handle edge cases (e.g., password-protected PDFs, corrupted files) by adding try-except blocks if needed.
内容的提问来源于stack exchange,提问作者Huma Qureshi
相关产品推荐
相关产品推荐

