Scrapy多字段爬取问题:关联公司名称与地点并优化JSON输出
Hey there! Let's get your Scrapy project sorted out. The core issues right now are:
- You're yielding company names and locations as separate entries, so they don't stay linked to each other
- The location field is pulling in extra date text (like "Open until...")
- Your JSON output is split into separate lists instead of paired company-location objects
Here's how to fix everything:
1. Rewrite Your Spider Code to Pair Entries
The key mistake is looping twice over div.result-description—each loop yields a separate item, breaking the link between a company and its location. Instead, loop once per company container, extract both fields in the same iteration, and yield them together.
We'll also filter out the extra date text from locations:
import scrapy class CompanySpider(scrapy.Spider): name = "Company" def start_requests(self): urls = [ 'https://www.f6s.com/programs?type[]=accelerator&sort=open', ] for url in urls: yield scrapy.Request(url=url, callback=self.parse) def parse(self, response): # Iterate over each company's container (one per company) for company_card in response.css('div.result-description'): # Extract clean company name (use extract_first() to get a string, not a list) program_name = company_card.css('div.title a.action.main.noline::text').extract_first() if program_name: program_name = program_name.strip() # Extract location and filter out date text (like "Open until...") location = "Unknown" all_subtitle_texts = company_card.css('div.subtitle span::text').extract() for text in all_subtitle_texts: cleaned_text = text.strip() # Skip any text that starts with "Open" (the date info) if cleaned_text and not cleaned_text.startswith('Open'): location = cleaned_text break # Yield a single item with both fields paired if program_name: yield { 'program': program_name, 'location': location }
What Changed:
- Single loop per company: Each iteration processes one company card, so name and location are always paired.
extract_first()instead ofextract(): Returns a single string instead of a list, making your JSON cleaner.- Location filtering: We skip the "Open until..." date text to keep only the actual location.
2. Get Clean JSON Output
From your terminal logs, I see you already have these settings in settings.py:
FEED_FORMAT = 'json' FEED_URI = 'companies.json'
With the rewritten spider, your companies.json will now have properly paired entries like this:
[ {"program": "K - LAUNCHPAD 2018", "location": "Los Angeles, CA"}, {"program": "Z Nation Lab R...", "location": "Tokyo, Japan"} ]
No more split lists—each object is a complete company entry with its matching location.
Optional: Use Scrapy Items for Better Structure
For larger projects, it's better to define an Item class in items.py to keep your data structured:
# items.py import scrapy class CompanyItem(scrapy.Item): program = scrapy.Field() location = scrapy.Field()
Then update your spider to use this Item:
from myproject.items import CompanyItem # Inside parse()... yield CompanyItem( program=program_name, location=location )
This makes it easier to add pipelines (like data validation) later on.
Bonus: Add Pagination (If You Need More Pages)
If you want to scrape beyond the first page, add this at the end of your parse method to follow the "Next Page" link:
# Look for the next page link next_page = response.css('a.next-page::attr(href)').extract_first() if next_page: # Follow the link and call parse again for the next page yield response.follow(next_page, self.parse)
That's it! Your spider will now output clean, paired company-location data in JSON format.
内容的提问来源于stack exchange,提问作者paperelephant

