You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多字段爬取问题:关联公司名称与地点并优化JSON输出

Fixing Your Scrapy Company-Location Pairing & JSON Output

Hey there! Let's get your Scrapy project sorted out. The core issues right now are:

  • You're yielding company names and locations as separate entries, so they don't stay linked to each other
  • The location field is pulling in extra date text (like "Open until...")
  • Your JSON output is split into separate lists instead of paired company-location objects

Here's how to fix everything:

1. Rewrite Your Spider Code to Pair Entries

The key mistake is looping twice over div.result-description—each loop yields a separate item, breaking the link between a company and its location. Instead, loop once per company container, extract both fields in the same iteration, and yield them together.

We'll also filter out the extra date text from locations:

import scrapy

class CompanySpider(scrapy.Spider):
    name = "Company"
    def start_requests(self):
        urls = [
            'https://www.f6s.com/programs?type[]=accelerator&sort=open',
        ]
        for url in urls:
            yield scrapy.Request(url=url, callback=self.parse)

    def parse(self, response):
        # Iterate over each company's container (one per company)
        for company_card in response.css('div.result-description'):
            # Extract clean company name (use extract_first() to get a string, not a list)
            program_name = company_card.css('div.title a.action.main.noline::text').extract_first()
            if program_name:
                program_name = program_name.strip()

            # Extract location and filter out date text (like "Open until...")
            location = "Unknown"
            all_subtitle_texts = company_card.css('div.subtitle span::text').extract()
            for text in all_subtitle_texts:
                cleaned_text = text.strip()
                # Skip any text that starts with "Open" (the date info)
                if cleaned_text and not cleaned_text.startswith('Open'):
                    location = cleaned_text
                    break

            # Yield a single item with both fields paired
            if program_name:
                yield {
                    'program': program_name,
                    'location': location
                }

What Changed:

  • Single loop per company: Each iteration processes one company card, so name and location are always paired.
  • extract_first() instead of extract(): Returns a single string instead of a list, making your JSON cleaner.
  • Location filtering: We skip the "Open until..." date text to keep only the actual location.

2. Get Clean JSON Output

From your terminal logs, I see you already have these settings in settings.py:

FEED_FORMAT = 'json'
FEED_URI = 'companies.json'

With the rewritten spider, your companies.json will now have properly paired entries like this:

[
  {"program": "K - LAUNCHPAD 2018", "location": "Los Angeles, CA"},
  {"program": "Z Nation Lab R...", "location": "Tokyo, Japan"}
]

No more split lists—each object is a complete company entry with its matching location.

Optional: Use Scrapy Items for Better Structure

For larger projects, it's better to define an Item class in items.py to keep your data structured:

# items.py
import scrapy

class CompanyItem(scrapy.Item):
    program = scrapy.Field()
    location = scrapy.Field()

Then update your spider to use this Item:

from myproject.items import CompanyItem

# Inside parse()...
yield CompanyItem(
    program=program_name,
    location=location
)

This makes it easier to add pipelines (like data validation) later on.

Bonus: Add Pagination (If You Need More Pages)

If you want to scrape beyond the first page, add this at the end of your parse method to follow the "Next Page" link:

# Look for the next page link
next_page = response.css('a.next-page::attr(href)').extract_first()
if next_page:
    # Follow the link and call parse again for the next page
    yield response.follow(next_page, self.parse)

That's it! Your spider will now output clean, paired company-location data in JSON format.

内容的提问来源于stack exchange,提问作者paperelephant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:04:42