You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django集成Scrapy遇问题:自定义爬虫无法运行求助

Hey there! Let's figure out why your Scrapy spider works outside Django but breaks when integrated. I’ve spotted several critical issues in your pipeline and model code that are almost certainly causing the problem. Let’s break them down and fix them step by step:


1. Key Issues in Your Code

Spider.py: Accidental Tuples

Notice the trailing commas at the end of each field assignment in parse_detail? Those turn your field values into tuples instead of strings/lists. For example, x['breadcrumbs'] = ..., creates a tuple like (['crumb1', 'crumb2'],) instead of just the list. This will cause errors when trying to save to Django models.

Pipeline.py: Data Loss & Uninitialized Variables

  • You’re re-instantiating a blank detikNewsItem() instead of using the data from the Scrapy item passed into process_item—this discards all the scraped data.
  • self.items is never initialized, so calling self.items.append() will throw an AttributeError.

Models.py: Mismatched Field Types & Broken Methods

  • Fields like breadcrumbs and tag are lists from your spider, but you’re trying to store them directly in a TextField—Django can’t save raw lists to the database.
  • The to_dict method uses json.loads(self.url) on a plain string URL, which will cause a JSON decoding error.
  • Using self.url in __str__ will lead to messy, long output in the Django admin.

2. Fixed Code Examples

Corrected spider.py

import scrapy
from scrapy_splash import SplashRequest

class NewsSpider(scrapy.Spider):
    name = 'detik'
    allowed_domains = ['news.detik.com']
    start_urls = ['https://news.detik.com/indeks/all/?date=02/28/2018']

    def parse(self, response):
        urls = response.xpath("//div/article/a/@href").extract()
        for url in urls:
            url = response.urljoin(url)
            yield scrapy.Request(url=url, callback=self.parse_detail)
        
        # Follow pagination link
        page_next = response.xpath("//a[@class = 'last']/@href").extract_first()
        if page_next:
            page_next = response.urljoin(page_next)
            yield scrapy.Request(url=page_next, callback=self.parse)

    def parse_detail(self,response):
        item = {}
        item['breadcrumbs'] = response.xpath("//div[@class='breadcrumb']/a/text()").extract()
        item['tanggal'] = response.xpath("//div[@class='date']/text()").extract_first()
        item['penulis'] = response.xpath("//div[@class='author']/text()").extract_first()
        item['judul'] = response.xpath("//h1/text()").extract_first()
        item['berita'] = response.xpath("normalize-space(//div[@class='detail_text'])").extract_first()
        item['tag'] = response.xpath("//div[@class='detail_tag']/a/text()").extract()
        item['url'] = response.request.url
        return item

Corrected Pipeline (pipelines.py)

import json
from django.conf import settings
from your_django_app.models import detikNewsItem  # Replace with your actual app name

class DetikAppPipeline(object):
    def process_item(self, item, spider):
        # Serialize list fields to JSON strings for Django storage
        breadcrumbs_json = json.dumps(item['breadcrumbs'])
        tag_json = json.dumps(item['tag'])

        # Create and save the Django model instance with scraped data
        news_entry = detikNewsItem(
            breadcrumbs=breadcrumbs_json,
            tanggal=item['tanggal'],
            penulis=item['penulis'],
            judul=item['judul'],
            berita=item['berita'],
            tag=tag_json,
            url=item['url']
        )
        news_entry.save()
        return item

Corrected Django Model (models.py)

import json
from django.db import models

class detikNewsItem(models.Model):
    breadcrumbs = models.TextField()  # Stores JSON-serialized list
    tanggal = models.TextField()
    penulis = models.TextField()
    judul = models.TextField()
    berita = models.TextField()
    tag = models.TextField()  # Stores JSON-serialized list
    url = models.URLField(unique=True)  # Better than TextField for URL validation

    @property
    def to_dict(self):
        # Deserialize JSON fields back to lists
        return {
            'url': self.url,
            'tanggal': self.tanggal,
            'breadcrumbs': json.loads(self.breadcrumbs),
            'tag': json.loads(self.tag),
            'judul': self.judul,
            'penulis': self.penulis,
            'berita': self.berita
        }

    def __str__(self):
        # Use title for cleaner admin display (fallback to truncated URL)
        return self.judul or f"{self.url[:50]}..."

3. Critical Setup Checks

Don’t forget these steps to ensure Scrapy and Django play nicely:

  • Configure Scrapy to use Django: Add this to your Scrapy project’s settings.py:
    import os
    import sys
    sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
    os.environ['DJANGO_SETTINGS_MODULE'] = 'your_django_project.settings'  # Replace with your project name
    import django
    django.setup()
    
  • Enable the pipeline: Also in settings.py:
    ITEM_PIPELINES = {
        'your_scrapy_project.pipelines.DetikAppPipeline': 300,
    }
    
  • Run Django migrations: Before running the spider, execute python manage.py migrate to create the database table for detikNewsItem.

内容的提问来源于stack exchange,提问作者Vira Xeva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 09:01:43