You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遵循SOLID原则:将Django视图中PDF读取代码提取为独立方法

提取PDF读取逻辑至独立方法的实现方案

下面是重构后的代码,将PDF读取逻辑提取为独立类方法,符合单一职责原则(SOLID中的S):

import os
import io
from PIL import Image
import pytesseract
from wand.image import Image as wi
from django.views import View
from django.conf import settings
from .forms import ProfileForm
from .models import UploadFile

class ReadingFile(View):
    def get(self, request):
        form = ProfileForm()
        return render(request, "main/create_profile.html", {
            "form": form
        })
    
    def extractingtextfromimage(self):  # 修复缺少self参数的问题
        pass

    # 新增私有方法专门处理PDF文本提取
    def _extract_text_from_pdf(self, pdf_path):
        pdfFile = wi(filename=pdf_path, resolution=300)
        text_factuur_verdi = []

        image = pdfFile.convert('jpeg')
        imageBlobs = []

        for img in image.sequence:
            imgPage = wi(image=img)
            imageBlobs.append(imgPage.make_blob('jpeg'))

        for imgBlob in imageBlobs:
            image = Image.open(io.BytesIO(imgBlob))
            text = pytesseract.image_to_string(image, lang='eng')
            text_factuur_verdi.append(text)
        
        return text_factuur_verdi

    def post(self, request):
        submitted_form = ProfileForm(request.POST, request.FILES)
        content = ''

        if submitted_form.is_valid():
            uploadfile = UploadFile(image=request.FILES["upload_file"])
            name_of_file = str(request.FILES['upload_file'])
            uploadfile.save()
            
            print('path of the file is:::', uploadfile.image.name)            
            print("Now its type is ", type(name_of_file))
            print(uploadfile.image.path)

            # 分支处理不同类型文件
            if name_of_file.endswith('.pdf'):
                content = self._extract_text_from_pdf(uploadfile.image.path)
                print(content)
            else:
                # 非PDF文件直接读取内容
                with open(os.path.join(settings.MEDIA_ROOT, f"{uploadfile.image}"), 'r') as f:
                    content = f.read()
                    print(content)

            return render(request, "main/create_profile.html", {
                'form': ProfileForm(),
                "content": content
            })

        return render(request, "main/create_profile.html", {
            "form": submitted_form,
        })

关键修改说明:

  • 新增_extract_text_from_pdf私有方法:把原post方法中PDF转图片提取文本的逻辑完整迁移过来,仅接收PDF路径参数并返回文本列表,让每个方法只负责单一职责
  • 简化post方法逻辑:移除原PDF处理代码块,替换为调用新方法,降低post方法的复杂度,更易维护
  • 修复extractingtextfromimage方法的参数问题:类实例方法必须包含self参数
  • 调整文件读取结构:将非PDF文件的with open操作移到对应分支下,避免PDF处理时不必要的文件打开操作

内容的提问来源于stack exchange,提问作者mightycode Newton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 17:15:40