Django集成Scrapy遇问题:自定义爬虫无法运行求助
Hey there! Let's figure out why your Scrapy spider works outside Django but breaks when integrated. I’ve spotted several critical issues in your pipeline and model code that are almost certainly causing the problem. Let’s break them down and fix them step by step:
1. Key Issues in Your Code
Spider.py: Accidental Tuples
Notice the trailing commas at the end of each field assignment in parse_detail? Those turn your field values into tuples instead of strings/lists. For example, x['breadcrumbs'] = ..., creates a tuple like (['crumb1', 'crumb2'],) instead of just the list. This will cause errors when trying to save to Django models.
Pipeline.py: Data Loss & Uninitialized Variables
- You’re re-instantiating a blank
detikNewsItem()instead of using the data from the Scrapy item passed intoprocess_item—this discards all the scraped data. self.itemsis never initialized, so callingself.items.append()will throw anAttributeError.
Models.py: Mismatched Field Types & Broken Methods
- Fields like
breadcrumbsandtagare lists from your spider, but you’re trying to store them directly in aTextField—Django can’t save raw lists to the database. - The
to_dictmethod usesjson.loads(self.url)on a plain string URL, which will cause a JSON decoding error. - Using
self.urlin__str__will lead to messy, long output in the Django admin.
2. Fixed Code Examples
Corrected spider.py
import scrapy from scrapy_splash import SplashRequest class NewsSpider(scrapy.Spider): name = 'detik' allowed_domains = ['news.detik.com'] start_urls = ['https://news.detik.com/indeks/all/?date=02/28/2018'] def parse(self, response): urls = response.xpath("//div/article/a/@href").extract() for url in urls: url = response.urljoin(url) yield scrapy.Request(url=url, callback=self.parse_detail) # Follow pagination link page_next = response.xpath("//a[@class = 'last']/@href").extract_first() if page_next: page_next = response.urljoin(page_next) yield scrapy.Request(url=page_next, callback=self.parse) def parse_detail(self,response): item = {} item['breadcrumbs'] = response.xpath("//div[@class='breadcrumb']/a/text()").extract() item['tanggal'] = response.xpath("//div[@class='date']/text()").extract_first() item['penulis'] = response.xpath("//div[@class='author']/text()").extract_first() item['judul'] = response.xpath("//h1/text()").extract_first() item['berita'] = response.xpath("normalize-space(//div[@class='detail_text'])").extract_first() item['tag'] = response.xpath("//div[@class='detail_tag']/a/text()").extract() item['url'] = response.request.url return item
Corrected Pipeline (pipelines.py)
import json from django.conf import settings from your_django_app.models import detikNewsItem # Replace with your actual app name class DetikAppPipeline(object): def process_item(self, item, spider): # Serialize list fields to JSON strings for Django storage breadcrumbs_json = json.dumps(item['breadcrumbs']) tag_json = json.dumps(item['tag']) # Create and save the Django model instance with scraped data news_entry = detikNewsItem( breadcrumbs=breadcrumbs_json, tanggal=item['tanggal'], penulis=item['penulis'], judul=item['judul'], berita=item['berita'], tag=tag_json, url=item['url'] ) news_entry.save() return item
Corrected Django Model (models.py)
import json from django.db import models class detikNewsItem(models.Model): breadcrumbs = models.TextField() # Stores JSON-serialized list tanggal = models.TextField() penulis = models.TextField() judul = models.TextField() berita = models.TextField() tag = models.TextField() # Stores JSON-serialized list url = models.URLField(unique=True) # Better than TextField for URL validation @property def to_dict(self): # Deserialize JSON fields back to lists return { 'url': self.url, 'tanggal': self.tanggal, 'breadcrumbs': json.loads(self.breadcrumbs), 'tag': json.loads(self.tag), 'judul': self.judul, 'penulis': self.penulis, 'berita': self.berita } def __str__(self): # Use title for cleaner admin display (fallback to truncated URL) return self.judul or f"{self.url[:50]}..."
3. Critical Setup Checks
Don’t forget these steps to ensure Scrapy and Django play nicely:
- Configure Scrapy to use Django: Add this to your Scrapy project’s
settings.py:import os import sys sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) os.environ['DJANGO_SETTINGS_MODULE'] = 'your_django_project.settings' # Replace with your project name import django django.setup() - Enable the pipeline: Also in
settings.py:ITEM_PIPELINES = { 'your_scrapy_project.pipelines.DetikAppPipeline': 300, } - Run Django migrations: Before running the spider, execute
python manage.py migrateto create the database table fordetikNewsItem.
内容的提问来源于stack exchange,提问作者Vira Xeva
相关产品推荐
相关产品推荐

