You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy技术问题:如何抓取页面所有标签页中的信息?

Multi-Tab Data Scraping Solution for Wellness.py

Hey Alex, let’s get your crawler to pull data from all three tabs (Profile, Reviews, Phone Numbers & Directions) instead of just Profile. I’ll cover two common scenarios you might run into, since how tabs load content varies across sites.


Scenario 1: All Tab Content is Already in the Page DOM (CSS-Hidden)

Lots of sites load all tab content upfront but hide inactive tabs with CSS. If that’s the case, you can grab everything from the initial response without extra requests. Here’s how to either bundle all data into one item or split it into separate items for clarity:

Option 1: Bundle All Data into a Single Item

def parse(self, response):
    # Grab Profile tab data (reusing your existing logic)
    services = response.xpath('.//span[contains(text(),"Services")]')
    education = response.xpath('.//span[contains(text(),"Education")]')
    training = response.xpath('.//span[contains(text(),"Training")]')
    
    profile_data = {
        'First and Last name': response.css('h1::text').get(),
        'About': response.css('.listing-about::text').get(),
        'Services': services.xpath('following-sibling::span[1]/text()').extract(),
        'Primary Specialty': response.css('.normal::text').get(),
        'Address': ' '.join([i.strip() for i in response.css('.office-address span::text').getall()]),
        'Practice': response.css('.years-in-service::text').get(),
        'Education': education.xpath('following-sibling::span[1]/text()').extract(),
        'Training': training.xpath('following-sibling::span[1]/text()').extract(),
        'Consumer Feedback': response.css('.item-rating-container a::text').get()
    }
    
    # Grab Reviews tab data (adjust selectors to match your site's HTML)
    reviews = []
    for review in response.css('.review-item'):
        reviews.append({
            'Reviewer Name': review.css('.reviewer-name::text').get(),
            'Rating': review.css('.review-rating::attr(data-score)').get(),
            'Review Text': review.css('.review-content::text').get().strip() if review.css('.review-content::text').get() else None,
            'Date': review.css('.review-date::text').get()
        })
    
    # Grab Phone Numbers & Directions tab data
    phone_directions = {
        'Phone Number': response.css('.contact-phone::text').get().strip() if response.css('.contact-phone::text').get() else None,
        'Full Address': ' '.join([i.strip() for i in response.css('.directions-address span::text').getall()]),
        'Directions Link': response.css('.get-directions::attr(href)').get(),
        'Hours': [hour.strip() for hour in response.css('.hours-list li::text').getall() if hour.strip()]
    }
    
    # Yield all data together
    yield {
        **profile_data,
        'Reviews': reviews,
        'Contact & Directions': phone_directions
    }

Option 2: Split into Separate Items (Easier for Post-Processing)

If you prefer to save each tab’s data as a distinct item (e.g., one Profile item, one Review per entry, one Contact item), adjust the parse method like this:

def parse(self, response):
    # Yield Profile data first
    services = response.xpath('.//span[contains(text(),"Services")]')
    education = response.xpath('.//span[contains(text(),"Education")]')
    training = response.xpath('.//span[contains(text(),"Training")]')
    yield {
        'item_type': 'profile',
        'First and Last name': response.css('h1::text').get(),
        'About': response.css('.listing-about::text').get(),
        'Services': services.xpath('following-sibling::span[1]/text()').extract(),
        'Primary Specialty': response.css('.normal::text').get(),
        'Address': ' '.join([i.strip() for i in response.css('.office-address span::text').getall()]),
        'Practice': response.css('.years-in-service::text').get(),
        'Education': education.xpath('following-sibling::span[1]/text()').extract(),
        'Training': training.xpath('following-sibling::span[1]/text()').extract(),
        'Consumer Feedback': response.css('.item-rating-container a::text').get()
    }
    
    # Yield each Review as a separate item
    profile_name = response.css('h1::text').get()
    for idx, review in enumerate(response.css('.review-item')):
        yield {
            'item_type': 'review',
            'profile_name': profile_name,
            'review_id': idx + 1,
            'Reviewer Name': review.css('.reviewer-name::text').get(),
            'Rating': review.css('.review-rating::attr(data-score)').get(),
            'Review Text': review.css('.review-content::text').get().strip() if review.css('.review-content::text').get() else None,
            'Date': review.css('.review-date::text').get()
        }
    
    # Yield Contact & Directions data
    yield {
        'item_type': 'contact_directions',
        'profile_name': profile_name,
        'Phone Number': response.css('.contact-phone::text').get().strip() if response.css('.contact-phone::text').get() else None,
        'Full Address': ' '.join([i.strip() for i in response.css('.directions-address span::text').getall()]),
        'Directions Link': response.css('.get-directions::attr(href)').get(),
        'Hours': [hour.strip() for hour in response.css('.hours-list li::text').getall() if hour.strip()]
    }

Scenario 2: Tabs Load Content via AJAX/New URLs

If clicking a tab triggers a new request (either to a separate URL or an API endpoint), you’ll need to follow those links and use callback methods to process each tab’s data:

def parse(self, response):
    # First, scrape the Profile tab
    yield from self.parse_profile(response)
    
    # Follow the Reviews tab link (adjust selector to match your site's tab link)
    reviews_link = response.css('.reviews-tab a::attr(href)').get()
    if reviews_link:
        yield response.follow(reviews_link, callback=self.parse_reviews)
    
    # Follow the Phone & Directions tab link
    contact_link = response.css('.contact-tab a::attr(href)').get()
    if contact_link:
        yield response.follow(contact_link, callback=self.parse_contact_directions)

def parse_profile(self, response):
    # Your original profile method, slightly updated with an item type
    services = response.xpath('.//span[contains(text(),"Services")]')
    education = response.xpath('.//span[contains(text(),"Education")]')
    training = response.xpath('.//span[contains(text(),"Training")]')
    yield {
        'item_type': 'profile',
        'First and Last name': response.css('h1::text').get(),
        'About': response.css('.listing-about::text').get(),
        'Services': services.xpath('following-sibling::span[1]/text()').extract(),
        'Primary Specialty': response.css('.normal::text').get(),
        'Address': ' '.join([i.strip() for i in response.css('.office-address span::text').getall()]),
        'Practice': response.css('.years-in-service::text').get(),
        'Education': education.xpath('following-sibling::span[1]/text()').extract(),
        'Training': training.xpath('following-sibling::span[1]/text()').extract(),
        'Consumer Feedback': response.css('.item-rating-container a::text').get()
    }

def parse_reviews(self, response):
    # Process Reviews page data
    profile_name = response.css('h1::text').get()
    for idx, review in enumerate(response.css('.review-item')):
        yield {
            'item_type': 'review',
            'profile_name': profile_name,
            'reviewer_name': review.css('.reviewer-name::text').get(),
            'rating': review.css('.star-rating::attr(data-value)').get(),
            'review_text': review.css('.review-body::text').get().strip() if review.css('.review-body::text').get() else None,
            'review_date': review.css('.review-date::text').get()
        }

def parse_contact_directions(self, response):
    # Process Contact & Directions page data
    profile_name = response.css('h1::text').get()
    yield {
        'item_type': 'contact_directions',
        'profile_name': profile_name,
        'phone_number': response.css('.contact-phone::text').get().strip() if response.css('.contact-phone::text').get() else None,
        'address': ' '.join([i.strip() for i in response.css('.directions-address span::text').getall()]),
        'google_maps_link': response.css('.maps-link::attr(href)').get(),
        'hours_of_operation': [hour.strip() for hour in response.css('.hours-list li::text').getall() if hour.strip()]
    }

Quick Tips to Get This Working

  1. Update Selectors: All the CSS/XPath selectors in the code above (like .review-item, .contact-phone) are placeholders. Use scrapy shell "your-target-url" to test and tweak selectors until they correctly grab the data you need.
  2. Check for AJAX: If tabs don’t load content in the initial DOM, open your browser’s DevTools > Network tab, click the tab, and look for the request that loads the data. You might be able to call that API directly instead of following the tab link.
  3. Handle Edge Cases: Add strip() checks and if conditions to avoid None values breaking your data output.

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:17:33