You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Django Web应用中实现PDF/DOCX转TXT并前置保存时遭遇FileNotFoundError的技术求助

Fixing File Conversion Issue in Django Model's save() Method

Let's tackle this problem step by step—your error comes from a few key mistakes in how you're accessing the uploaded file and handling the conversion logic. Here's what's going wrong and how to fix it:

Key Issues in Your Current Code

  • Incorrect File Path: self.file.name only gives the relative filename, not the full path to the uploaded file. Before calling super().save(), the file hasn't been saved to your upload_to directory yet, so trying to open it via filename fails. Instead, use Django's built-in self.file.open() to access the file object directly.
  • Page Index Out of Bounds: pdfreader.numPages returns the total number of pages, which are 0-indexed. Using getPage(x + 1) will always throw an error because you're trying to access a page that doesn't exist.
  • Invalid Filename Syntax: self.file.name.txt is not valid Python syntax. You need to properly replace the file extension to .txt using os.path.splitext().
  • Missing DOCX Handling: Your code only handles PDFs, but you mentioned needing to convert .docx files too.

Fixed Code with PDF + DOCX Support

First, install the required dependencies if you haven't already:

pip install PyPDF2 python-docx

Then update your models.py:

from django.db import models
from django.contrib.auth.models import User
from django.urls import reverse
from django.conf import settings
import PyPDF2
from docx import Document
import os

class FileUpload(models.Model):
    user = models.ForeignKey(User, on_delete=models.CASCADE, blank=True, null=True)
    file = models.FileField(upload_to='files')
    # Optional: Add a field to store the extracted text if you need it in the database
    extracted_text = models.TextField(blank=True, null=True)

    def save(self, *args, **kwargs):
        extracted_content = ""
        file_ext = os.path.splitext(self.file.name)[1].lower()

        if file_ext == '.pdf':
            # Open the uploaded PDF directly using Django's file object
            with self.file.open('rb') as pdffileobj:
                pdfreader = PyPDF2.PdfReader(pdffileobj)
                # Iterate through all pages to extract text
                for page in pdfreader.pages:
                    extracted_content += page.extract_text() or ""

        elif file_ext == '.docx':
            # Handle DOCX files using python-docx
            with self.file.open('rb') as docxfileobj:
                doc = Document(docxfileobj)
                for para in doc.paragraphs:
                    extracted_content += para.text + "\n"

        # Save the extracted text to the model field (if you added it)
        self.extracted_text = extracted_content

        # Optionally: Save the text as a .txt file in the same directory as the uploaded file
        if extracted_content:
            txt_filename = os.path.splitext(self.file.name)[0] + '.txt'
            txt_file_path = os.path.join(settings.MEDIA_ROOT, txt_filename)
            with open(txt_file_path, 'w', encoding='utf-8') as txt_file:
                txt_file.write(extracted_content)

        # Always call the parent save method last to ensure the file is saved to disk
        super(FileUpload, self).save(*args, **kwargs)

    def get_absolute_url(self):
        return reverse('home')

What Changed?

  1. Direct File Access: Used self.file.open() to read the uploaded file without needing a filesystem path—this works even before the file is saved to your upload_to directory.
  2. Valid Page Iteration: For PDFs, we loop through all pages using pdfreader.pages (the modern PyPDF2 API—note PdfFileReader is deprecated in newer versions).
  3. DOCX Handling: Added support for .docx files using the python-docx library, extracting text from all paragraphs.
  4. Proper TXT Filename: Used os.path.splitext() to safely replace the original file extension with .txt, and saved the file to your MEDIA_ROOT directory.
  5. Optional Database Storage: Added an extracted_text field to store the converted text directly in the database, which is useful if you need to display or search the content later.
  6. Parent Save Last: We call super().save() after processing to ensure the original file is saved to disk first (if you need to reference its final path).

Notes

  • Make sure your Django MEDIA_ROOT and MEDIA_URL are properly configured in settings.py so the uploaded files and generated .txt files are accessible.
  • Handle edge cases (e.g., password-protected PDFs, corrupted files) by adding try-except blocks if needed.

内容的提问来源于stack exchange,提问作者Huma Qureshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 04:32:47