Scrapy爬虫获取相对URL问题求助:如何配置仅抓取绝对URL
Hey there! I totally get the frustration when relative URLs break your scraping workflow—let's fix this quickly. Scrapy has a built-in way to handle this without extra hassle.
The Core Issue
Your current code extracts raw href attributes, which can be relative paths (like /about or ../contact). To turn these into usable absolute URLs, we'll use Scrapy's response.urljoin() method, which resolves relative paths using the current page's URL as the base.
Modified Code
Here's your updated spider with the fix applied:
import scrapy import os class MySpider(scrapy.Spider): name = 'feed_exporter_test' custom_settings = { 'FEED_FORMAT': 'csv', 'FEED_URI': 'file1.csv' } filePath='file1.csv' if os.path.exists(filePath): os.remove(filePath) else: print("Can not delete the file as it doesn't exists") start_urls = ['https://www.jamoona.com/'] def parse(self, response): # Extract all relative href attributes relative_urls = response.xpath("//a/@href").extract() for rel_url in relative_urls: # Skip empty or invalid hrefs to avoid errors if not rel_url: continue # Convert relative URL to absolute absolute_url = response.urljoin(rel_url) yield {'title': absolute_url}
Key Changes Explained
response.urljoin(rel_url): This method automatically handles all types of relative paths:- Root-relative paths (e.g.,
/blog→https://www.jamoona.com/blog) - Relative paths (e.g.,
../about→https://www.jamoona.com/about) - Even full absolute URLs will pass through unchanged, so you don't have to worry about checking if the URL is already absolute.
- Root-relative paths (e.g.,
- Empty href check: Added a quick skip for empty
hrefvalues to prevent unnecessary processing and potential errors.
Bonus Tip
If you ever need to follow these URLs to scrape further content, you can use response.follow(rel_url) instead—it automatically handles relative URLs too, and returns a new Response object for parsing the target page.
内容的提问来源于stack exchange,提问作者dipu gala

