You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫获取相对URL问题求助:如何配置仅抓取绝对URL

Fix: Convert Relative URLs to Absolute in Your Scrapy Spider

Hey there! I totally get the frustration when relative URLs break your scraping workflow—let's fix this quickly. Scrapy has a built-in way to handle this without extra hassle.

The Core Issue

Your current code extracts raw href attributes, which can be relative paths (like /about or ../contact). To turn these into usable absolute URLs, we'll use Scrapy's response.urljoin() method, which resolves relative paths using the current page's URL as the base.

Modified Code

Here's your updated spider with the fix applied:

import scrapy
import os

class MySpider(scrapy.Spider):
    name = 'feed_exporter_test'
    custom_settings = {
        'FEED_FORMAT': 'csv',
        'FEED_URI': 'file1.csv'
    }
    filePath='file1.csv'
    
    if os.path.exists(filePath):
        os.remove(filePath)
    else:
        print("Can not delete the file as it doesn't exists")
    
    start_urls = ['https://www.jamoona.com/']
    
    def parse(self, response):
        # Extract all relative href attributes
        relative_urls = response.xpath("//a/@href").extract()
        
        for rel_url in relative_urls:
            # Skip empty or invalid hrefs to avoid errors
            if not rel_url:
                continue
            # Convert relative URL to absolute
            absolute_url = response.urljoin(rel_url)
            yield {'title': absolute_url}

Key Changes Explained

  • response.urljoin(rel_url): This method automatically handles all types of relative paths:
    • Root-relative paths (e.g., /blog → https://www.jamoona.com/blog)
    • Relative paths (e.g., ../about → https://www.jamoona.com/about)
    • Even full absolute URLs will pass through unchanged, so you don't have to worry about checking if the URL is already absolute.
  • Empty href check: Added a quick skip for empty href values to prevent unnecessary processing and potential errors.

Bonus Tip

If you ever need to follow these URLs to scrape further content, you can use response.follow(rel_url) instead—it automatically handles relative URLs too, and returns a new Response object for parsing the target page.

内容的提问来源于stack exchange,提问作者dipu gala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 21:09:09