You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在查词Django API中正确实现网页抓取及逻辑架构优化?

Hey there! Let's tackle cleaning up your Django API logic and fixing those 404 response issues you're facing. Your current approach cramming all logic into get_queryset is definitely going to cause structural headaches—let's refactor this properly to separate concerns and handle all your requirements smoothly.

First, Let's Fix the Core Issue: Misusing get_queryset

The get_queryset method in Django is meant only to return a queryset for the view to work with. It shouldn't handle web scraping, database saves, or return non-queryset values. That's why you're getting weird errors when trying to return JSON on 404—you're fighting against Django's view flow.

Instead, we'll split your code into two parts:

  1. A service layer to handle all business logic (database checks, scraping, saving data)
  2. A clean API view that handles requests/responses using Django REST Framework's APIView

Step 1: Create a Service Class for Business Logic

Create a new file services.py in your app to encapsulate all word-related logic. This keeps your views clean and makes the code easier to test and maintain:

# services.py
import requests
from bs4 import BeautifulSoup
from .models import Word

class WordService:
    @staticmethod
    def fetch_word_data(word):
        # First check if the word exists in our database (case-insensitive)
        existing_word = Word.objects.filter(word__iexact=word).first()
        if existing_word:
            return {
                "status": "success",
                "data": {"word": existing_word.word, "meaning": existing_word.meaning}
            }
        
        # If not in DB, scrape dictionary.com
        headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"}
        url = f"https://www.dictionary.com/browse/{word}?s=t"
        
        try:
            page = requests.get(url, headers=headers)
            page.raise_for_status()  # Raise error for HTTP status codes like 404
            soup = BeautifulSoup(page.content, 'html.parser')
            
            # Try to extract the meaning
            meaning_element = soup.find("div", {"value": "1"})
            if meaning_element:
                meaning = meaning_element.get_text(strip=True)
                # Save to database
                new_word = Word.objects.create(word=word.strip(), meaning=meaning)
                return {
                    "status": "success",
                    "data": {"word": new_word.word, "meaning": new_word.meaning}
                }
            else:
                # No meaning found—extract suggested words
                suggestions = []
                suggestion_links = soup.select(".spell-suggestions a")
                for link in suggestion_links:
                    suggestions.append(link.get_text(strip=True))
                
                if suggestions:
                    return {
                        "status": "suggestions",
                        "data": {"suggestions": suggestions}
                    }
                else:
                    return {
                        "status": "not_found",
                        "data": {"message": "Word not found and no suggestions available"}
                    }
        except requests.exceptions.RequestException:
            return {
                "status": "error",
                "data": {"message": "Failed to connect to dictionary service"}
            }

Step 2: Refactor Your View to Use the Service

Now update your view to use APIView instead of relying on get_queryset. This gives you full control over the response:

# views.py
from rest_framework.views import APIView
from rest_framework.response import Response
from rest_framework import status
from .services import WordService

class WordMeaningView(APIView):
    def get(self, request, word):
        result = WordService.fetch_word_data(word)
        
        if result["status"] == "success":
            return Response(result["data"], status=status.HTTP_200_OK)
        elif result["status"] == "suggestions":
            return Response(
                {"message": "Word not found. Did you mean?", **result["data"]},
                status=status.HTTP_404_NOT_FOUND
            )
        elif result["status"] == "not_found":
            return Response(result["data"], status=status.HTTP_404_NOT_FOUND)
        else:
            return Response(result["data"], status=status.HTTP_500_INTERNAL_SERVER_ERROR)

Key Improvements Here:

  • Separation of Concerns: All business logic lives in the service class—your view only handles request/response formatting.
  • Proper Error Handling: We catch HTTP request errors, handle missing meanings gracefully, and return clear JSON responses for every scenario.
  • Case Insensitivity: Using __iexact ensures users get results regardless of uppercase/lowercase input.
  • Clear Status Codes: Return 200 for found words, 404 for missing words with/without suggestions, and 500 for service errors.

Bonus Tips for Production:

  • Add caching (e.g., using Django's cache framework) for scraped results to avoid hitting dictionary.com on every repeated request.
  • Add rate limiting to prevent abuse of your API and avoid getting blocked by dictionary.com.
  • Add indexes to your Word model's word field (db_index=True) to speed up database lookups.
  • Consider adding retry logic for failed scrapes (using a library like tenacity) to handle temporary network issues.

内容的提问来源于stack exchange,提问作者LDomain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:37:50