You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Django Rest Framework构建HTML内容爬取API的项目结构咨询

嘿,我来帮你梳理下这个 Django+DRF 爬虫 API 项目的规划思路,结合你侧重内容提取而非网页本身的需求,咱们一步步拆解:

项目结构与实现方案

我会推荐一个解耦清晰的结构,把爬取逻辑、数据存储、API 服务分开,方便后续维护和扩展。

1. 基础项目结构(Django + DRF)

先搭一个标准的 Django 项目,再拆分出专门的模块负责爬取和 API:

my_content_api/
├── my_content_api/          # Django 核心配置目录
│   ├── __init__.py
│   ├── settings.py
│   ├── urls.py
│   └── wsgi.py
├── content_store/           # 负责数据模型与 ORM 操作
│   ├── __init__.py
│   ├── admin.py
│   ├── models.py
│   ├── serializers.py       # DRF 序列化器
│   ├── views.py             # DRF 视图
│   └── urls.py              # API 子路由
├── content_scraper/         # 独立的爬取逻辑模块
│   ├── __init__.py
│   ├── parsers/             # 不同网站的内容解析器
│   │   ├── __init__.py
│   │   ├── example_site_parser.py
│   │   └── another_site_parser.py
│   ├── fetchers.py          # 请求工具函数
│   └── tasks.py             # 爬取任务脚本(可用于定时执行)
├── manage.py
└── requirements.txt
  • content_store:专注于数据持久化和 API 输出,和爬取逻辑完全解耦
  • content_scraper:只负责从网页提取核心内容,把数据交给content_store存储

2. Django ORM 模型设计

根据你要爬取的内容类型(比如文章、产品信息等)设计模型,核心是存储提取后的内容而非原始 HTML:

# content_store/models.py
from django.db import models

class ScrapedContent(models.Model):
    title = models.CharField(max_length=255, verbose_name="内容标题")
    core_content = models.TextField(verbose_name="核心内容")
    source_url = models.URLField(unique=True, verbose_name="来源URL")
    source_domain = models.CharField(max_length=100, verbose_name="来源域名")
    scraped_at = models.DateTimeField(auto_now_add=True, verbose_name="首次爬取时间")
    updated_at = models.DateTimeField(auto_now=True, verbose_name="最后更新时间")

    class Meta:
        ordering = ['-scraped_at']
        verbose_name_plural = "爬取内容库"

    def __str__(self):
        return f"{self.title} | {self.source_domain}"
  • source_url设为unique,避免重复爬取同一页面
  • source_domain方便后续按网站分类筛选内容

3. 爬取逻辑与 Django 整合

你觉得 Scrapy 侧重网页爬取,但其实它也能做内容提取;如果需求简单,用requests+BeautifulSoup更轻量,这里两种方案都给你参考:

方案A:轻量版(Requests + BeautifulSoup)

适合单页/少量页面的内容提取,直接在content_scraper里写工具函数:

# content_scraper/fetchers.py
import requests
from bs4 import BeautifulSoup

def fetch_and_extract(url):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    try:
        response = requests.get(url, headers=headers, timeout=15)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 这里根据目标网站的HTML结构调整提取规则
        title = soup.find('h1').get_text(strip=True) if soup.find('h1') else "无标题"
        # 提取正文:比如取所有类名为`article-content`下的p标签
        content_blocks = soup.select('.article-content p')
        core_content = '\n\n'.join([block.get_text(strip=True) for block in content_blocks])
        
        domain = url.split('//')[1].split('/')[0]
        
        return {
            'title': title,
            'core_content': core_content,
            'source_url': url,
            'source_domain': domain
        }
    except Exception as e:
        print(f"爬取失败 [{url}]: {str(e)}")
        return None

然后写一个 Django 自定义命令,用来触发爬取并保存数据:

# content_scraper/tasks.py
from django.core.management.base import BaseCommand
from content_scraper.fetchers import fetch_and_extract
from content_store.models import ScrapedContent

class Command(BaseCommand):
    help = "爬取指定URL的核心内容并保存到数据库"

    def add_arguments(self, parser):
        parser.add_argument('url', type=str, help="目标网页URL")

    def handle(self, *args, **options):
        url = options['url']
        content_data = fetch_and_extract(url)
        if content_data:
            # 避免重复存储,用get_or_create
            obj, created = ScrapedContent.objects.get_or_create(
                source_url=url,
                defaults=content_data
            )
            if created:
                self.stdout.write(self.style.SUCCESS(f"✅ 成功保存内容: {obj.title}"))
            else:
                self.stdout.write(self.style.WARNING(f"⚠️ 内容已存在: {obj.title}"))
        else:
            self.stdout.write(self.style.ERROR(f"❌ 爬取URL失败: {url}"))

运行方式:python manage.py scrape https://example.com/article

方案B:复杂爬取(Scrapy + Django ORM)

如果需要批量爬取、处理分页/异步请求,Scrapy 还是很合适的,只要把数据通过 Pipeline 存到 Django ORM:

  1. 在content_scraper目录下创建 Scrapy 项目
  2. 配置 Scrapy 的 Pipeline,连接 Django 环境并保存数据:
# content_scraper/scrapy_project/pipelines.py
import os
import django
os.environ.setdefault("DJANGO_SETTINGS_MODULE", "my_content_api.settings")
django.setup()

from content_store.models import ScrapedContent

class DjangoORMStoragePipeline:
    def process_item(self, item, spider):
        try:
            ScrapedContent.objects.get_or_create(
                source_url=item['source_url'],
                defaults={
                    'title': item['title'],
                    'core_content': item['core_content'],
                    'source_domain': item['source_domain']
                }
            )
            spider.logger.info(f"保存内容成功: {item['title']}")
            return item
        except Exception as e:
            spider.logger.error(f"保存失败: {str(e)}")
            raise e
  1. 在 Scrapy 的settings.py里启用这个 Pipeline 即可。

4. DRF API 配置

快速实现可访问的 API,用ModelViewSet就能一键生成列表、详情接口:

序列化器

# content_store/serializers.py
from rest_framework import serializers
from .models import ScrapedContent

class ScrapedContentSerializer(serializers.ModelSerializer):
    class Meta:
        model = ScrapedContent
        fields = ['id', 'title', 'core_content', 'source_url', 'source_domain', 'scraped_at', 'updated_at']

视图

# content_store/views.py
from rest_framework import viewsets
from .models import ScrapedContent
from .serializers import ScrapedContentSerializer

class ScrapedContentViewSet(viewsets.ReadOnlyModelViewSet):
    # 只开放读权限,因为内容是爬取生成的,不需要用户修改
    queryset = ScrapedContent.objects.all()
    serializer_class = ScrapedContentSerializer
    # 添加过滤、搜索功能
    filterset_fields = ['source_domain']
    search_fields = ['title', 'core_content']

路由配置

  • 在content_store/urls.py里注册路由:
from django.urls import path, include
from rest_framework.routers import DefaultRouter
from .views import ScrapedContentViewSet

router = DefaultRouter()
router.register(r'contents', ScrapedContentViewSet)

urlpatterns = [
    path('', include(router.urls)),
]
  • 在项目根urls.py里挂载 API:
from django.contrib import admin
from django.urls import path, include

urlpatterns = [
    path('admin/', admin.site.urls),
    path('api/', include('content_store.urls')),
]

现在你可以通过以下地址访问 API:

  • 列表接口:/api/contents/
  • 详情接口:/api/contents/<id>/
  • 筛选接口:/api/contents/?source_domain=example.com
  • 搜索接口:/api/contents/?search=关键词

5. 后续维护建议

  • 定时爬取:用 Celery + Redis 实现定时任务(比如每天凌晨爬取指定网站更新),或者用 Linux crontab 定时运行爬取命令
  • 内容更新:定期检查source_url对应的页面是否有更新,更新则刷新core_content和updated_at字段
  • 日志与监控:给爬取逻辑添加日志记录,方便排查失败的爬取任务;可以用 Django 的logging模块配置日志输出
  • 性能优化:给ScrapedContent模型的source_domain、scraped_at字段加索引;DRF 开启分页避免一次性返回大量数据

内容的提问来源于stack exchange,提问作者Quontas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:36:49