基于Django Rest Framework构建HTML内容爬取API的项目结构咨询
嘿,我来帮你梳理下这个 Django+DRF 爬虫 API 项目的规划思路,结合你侧重内容提取而非网页本身的需求,咱们一步步拆解:
项目结构与实现方案
我会推荐一个解耦清晰的结构,把爬取逻辑、数据存储、API 服务分开,方便后续维护和扩展。
1. 基础项目结构(Django + DRF)
先搭一个标准的 Django 项目,再拆分出专门的模块负责爬取和 API:
my_content_api/ ├── my_content_api/ # Django 核心配置目录 │ ├── __init__.py │ ├── settings.py │ ├── urls.py │ └── wsgi.py ├── content_store/ # 负责数据模型与 ORM 操作 │ ├── __init__.py │ ├── admin.py │ ├── models.py │ ├── serializers.py # DRF 序列化器 │ ├── views.py # DRF 视图 │ └── urls.py # API 子路由 ├── content_scraper/ # 独立的爬取逻辑模块 │ ├── __init__.py │ ├── parsers/ # 不同网站的内容解析器 │ │ ├── __init__.py │ │ ├── example_site_parser.py │ │ └── another_site_parser.py │ ├── fetchers.py # 请求工具函数 │ └── tasks.py # 爬取任务脚本(可用于定时执行) ├── manage.py └── requirements.txt
content_store:专注于数据持久化和 API 输出,和爬取逻辑完全解耦content_scraper:只负责从网页提取核心内容,把数据交给content_store存储
2. Django ORM 模型设计
根据你要爬取的内容类型(比如文章、产品信息等)设计模型,核心是存储提取后的内容而非原始 HTML:
# content_store/models.py from django.db import models class ScrapedContent(models.Model): title = models.CharField(max_length=255, verbose_name="内容标题") core_content = models.TextField(verbose_name="核心内容") source_url = models.URLField(unique=True, verbose_name="来源URL") source_domain = models.CharField(max_length=100, verbose_name="来源域名") scraped_at = models.DateTimeField(auto_now_add=True, verbose_name="首次爬取时间") updated_at = models.DateTimeField(auto_now=True, verbose_name="最后更新时间") class Meta: ordering = ['-scraped_at'] verbose_name_plural = "爬取内容库" def __str__(self): return f"{self.title} | {self.source_domain}"
source_url设为unique,避免重复爬取同一页面source_domain方便后续按网站分类筛选内容
3. 爬取逻辑与 Django 整合
你觉得 Scrapy 侧重网页爬取,但其实它也能做内容提取;如果需求简单,用requests+BeautifulSoup更轻量,这里两种方案都给你参考:
方案A:轻量版(Requests + BeautifulSoup)
适合单页/少量页面的内容提取,直接在content_scraper里写工具函数:
# content_scraper/fetchers.py import requests from bs4 import BeautifulSoup def fetch_and_extract(url): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } try: response = requests.get(url, headers=headers, timeout=15) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 这里根据目标网站的HTML结构调整提取规则 title = soup.find('h1').get_text(strip=True) if soup.find('h1') else "无标题" # 提取正文:比如取所有类名为`article-content`下的p标签 content_blocks = soup.select('.article-content p') core_content = '\n\n'.join([block.get_text(strip=True) for block in content_blocks]) domain = url.split('//')[1].split('/')[0] return { 'title': title, 'core_content': core_content, 'source_url': url, 'source_domain': domain } except Exception as e: print(f"爬取失败 [{url}]: {str(e)}") return None
然后写一个 Django 自定义命令,用来触发爬取并保存数据:
# content_scraper/tasks.py from django.core.management.base import BaseCommand from content_scraper.fetchers import fetch_and_extract from content_store.models import ScrapedContent class Command(BaseCommand): help = "爬取指定URL的核心内容并保存到数据库" def add_arguments(self, parser): parser.add_argument('url', type=str, help="目标网页URL") def handle(self, *args, **options): url = options['url'] content_data = fetch_and_extract(url) if content_data: # 避免重复存储,用get_or_create obj, created = ScrapedContent.objects.get_or_create( source_url=url, defaults=content_data ) if created: self.stdout.write(self.style.SUCCESS(f"✅ 成功保存内容: {obj.title}")) else: self.stdout.write(self.style.WARNING(f"⚠️ 内容已存在: {obj.title}")) else: self.stdout.write(self.style.ERROR(f"❌ 爬取URL失败: {url}"))
运行方式:python manage.py scrape https://example.com/article
方案B:复杂爬取(Scrapy + Django ORM)
如果需要批量爬取、处理分页/异步请求,Scrapy 还是很合适的,只要把数据通过 Pipeline 存到 Django ORM:
- 在
content_scraper目录下创建 Scrapy 项目 - 配置 Scrapy 的 Pipeline,连接 Django 环境并保存数据:
# content_scraper/scrapy_project/pipelines.py import os import django os.environ.setdefault("DJANGO_SETTINGS_MODULE", "my_content_api.settings") django.setup() from content_store.models import ScrapedContent class DjangoORMStoragePipeline: def process_item(self, item, spider): try: ScrapedContent.objects.get_or_create( source_url=item['source_url'], defaults={ 'title': item['title'], 'core_content': item['core_content'], 'source_domain': item['source_domain'] } ) spider.logger.info(f"保存内容成功: {item['title']}") return item except Exception as e: spider.logger.error(f"保存失败: {str(e)}") raise e
- 在 Scrapy 的
settings.py里启用这个 Pipeline 即可。
4. DRF API 配置
快速实现可访问的 API,用ModelViewSet就能一键生成列表、详情接口:
序列化器
# content_store/serializers.py from rest_framework import serializers from .models import ScrapedContent class ScrapedContentSerializer(serializers.ModelSerializer): class Meta: model = ScrapedContent fields = ['id', 'title', 'core_content', 'source_url', 'source_domain', 'scraped_at', 'updated_at']
视图
# content_store/views.py from rest_framework import viewsets from .models import ScrapedContent from .serializers import ScrapedContentSerializer class ScrapedContentViewSet(viewsets.ReadOnlyModelViewSet): # 只开放读权限,因为内容是爬取生成的,不需要用户修改 queryset = ScrapedContent.objects.all() serializer_class = ScrapedContentSerializer # 添加过滤、搜索功能 filterset_fields = ['source_domain'] search_fields = ['title', 'core_content']
路由配置
- 在
content_store/urls.py里注册路由:
from django.urls import path, include from rest_framework.routers import DefaultRouter from .views import ScrapedContentViewSet router = DefaultRouter() router.register(r'contents', ScrapedContentViewSet) urlpatterns = [ path('', include(router.urls)), ]
- 在项目根
urls.py里挂载 API:
from django.contrib import admin from django.urls import path, include urlpatterns = [ path('admin/', admin.site.urls), path('api/', include('content_store.urls')), ]
现在你可以通过以下地址访问 API:
- 列表接口:
/api/contents/ - 详情接口:
/api/contents/<id>/ - 筛选接口:
/api/contents/?source_domain=example.com - 搜索接口:
/api/contents/?search=关键词
5. 后续维护建议
- 定时爬取:用 Celery + Redis 实现定时任务(比如每天凌晨爬取指定网站更新),或者用 Linux crontab 定时运行爬取命令
- 内容更新:定期检查
source_url对应的页面是否有更新,更新则刷新core_content和updated_at字段 - 日志与监控:给爬取逻辑添加日志记录,方便排查失败的爬取任务;可以用 Django 的
logging模块配置日志输出 - 性能优化:给
ScrapedContent模型的source_domain、scraped_at字段加索引;DRF 开启分页避免一次性返回大量数据
内容的提问来源于stack exchange,提问作者Quontas
相关产品推荐
相关产品推荐

