Scrapy技术问题:如何抓取页面所有标签页中的信息?
Hey Alex, let’s get your crawler to pull data from all three tabs (Profile, Reviews, Phone Numbers & Directions) instead of just Profile. I’ll cover two common scenarios you might run into, since how tabs load content varies across sites.
Scenario 1: All Tab Content is Already in the Page DOM (CSS-Hidden)
Lots of sites load all tab content upfront but hide inactive tabs with CSS. If that’s the case, you can grab everything from the initial response without extra requests. Here’s how to either bundle all data into one item or split it into separate items for clarity:
Option 1: Bundle All Data into a Single Item
def parse(self, response): # Grab Profile tab data (reusing your existing logic) services = response.xpath('.//span[contains(text(),"Services")]') education = response.xpath('.//span[contains(text(),"Education")]') training = response.xpath('.//span[contains(text(),"Training")]') profile_data = { 'First and Last name': response.css('h1::text').get(), 'About': response.css('.listing-about::text').get(), 'Services': services.xpath('following-sibling::span[1]/text()').extract(), 'Primary Specialty': response.css('.normal::text').get(), 'Address': ' '.join([i.strip() for i in response.css('.office-address span::text').getall()]), 'Practice': response.css('.years-in-service::text').get(), 'Education': education.xpath('following-sibling::span[1]/text()').extract(), 'Training': training.xpath('following-sibling::span[1]/text()').extract(), 'Consumer Feedback': response.css('.item-rating-container a::text').get() } # Grab Reviews tab data (adjust selectors to match your site's HTML) reviews = [] for review in response.css('.review-item'): reviews.append({ 'Reviewer Name': review.css('.reviewer-name::text').get(), 'Rating': review.css('.review-rating::attr(data-score)').get(), 'Review Text': review.css('.review-content::text').get().strip() if review.css('.review-content::text').get() else None, 'Date': review.css('.review-date::text').get() }) # Grab Phone Numbers & Directions tab data phone_directions = { 'Phone Number': response.css('.contact-phone::text').get().strip() if response.css('.contact-phone::text').get() else None, 'Full Address': ' '.join([i.strip() for i in response.css('.directions-address span::text').getall()]), 'Directions Link': response.css('.get-directions::attr(href)').get(), 'Hours': [hour.strip() for hour in response.css('.hours-list li::text').getall() if hour.strip()] } # Yield all data together yield { **profile_data, 'Reviews': reviews, 'Contact & Directions': phone_directions }
Option 2: Split into Separate Items (Easier for Post-Processing)
If you prefer to save each tab’s data as a distinct item (e.g., one Profile item, one Review per entry, one Contact item), adjust the parse method like this:
def parse(self, response): # Yield Profile data first services = response.xpath('.//span[contains(text(),"Services")]') education = response.xpath('.//span[contains(text(),"Education")]') training = response.xpath('.//span[contains(text(),"Training")]') yield { 'item_type': 'profile', 'First and Last name': response.css('h1::text').get(), 'About': response.css('.listing-about::text').get(), 'Services': services.xpath('following-sibling::span[1]/text()').extract(), 'Primary Specialty': response.css('.normal::text').get(), 'Address': ' '.join([i.strip() for i in response.css('.office-address span::text').getall()]), 'Practice': response.css('.years-in-service::text').get(), 'Education': education.xpath('following-sibling::span[1]/text()').extract(), 'Training': training.xpath('following-sibling::span[1]/text()').extract(), 'Consumer Feedback': response.css('.item-rating-container a::text').get() } # Yield each Review as a separate item profile_name = response.css('h1::text').get() for idx, review in enumerate(response.css('.review-item')): yield { 'item_type': 'review', 'profile_name': profile_name, 'review_id': idx + 1, 'Reviewer Name': review.css('.reviewer-name::text').get(), 'Rating': review.css('.review-rating::attr(data-score)').get(), 'Review Text': review.css('.review-content::text').get().strip() if review.css('.review-content::text').get() else None, 'Date': review.css('.review-date::text').get() } # Yield Contact & Directions data yield { 'item_type': 'contact_directions', 'profile_name': profile_name, 'Phone Number': response.css('.contact-phone::text').get().strip() if response.css('.contact-phone::text').get() else None, 'Full Address': ' '.join([i.strip() for i in response.css('.directions-address span::text').getall()]), 'Directions Link': response.css('.get-directions::attr(href)').get(), 'Hours': [hour.strip() for hour in response.css('.hours-list li::text').getall() if hour.strip()] }
Scenario 2: Tabs Load Content via AJAX/New URLs
If clicking a tab triggers a new request (either to a separate URL or an API endpoint), you’ll need to follow those links and use callback methods to process each tab’s data:
def parse(self, response): # First, scrape the Profile tab yield from self.parse_profile(response) # Follow the Reviews tab link (adjust selector to match your site's tab link) reviews_link = response.css('.reviews-tab a::attr(href)').get() if reviews_link: yield response.follow(reviews_link, callback=self.parse_reviews) # Follow the Phone & Directions tab link contact_link = response.css('.contact-tab a::attr(href)').get() if contact_link: yield response.follow(contact_link, callback=self.parse_contact_directions) def parse_profile(self, response): # Your original profile method, slightly updated with an item type services = response.xpath('.//span[contains(text(),"Services")]') education = response.xpath('.//span[contains(text(),"Education")]') training = response.xpath('.//span[contains(text(),"Training")]') yield { 'item_type': 'profile', 'First and Last name': response.css('h1::text').get(), 'About': response.css('.listing-about::text').get(), 'Services': services.xpath('following-sibling::span[1]/text()').extract(), 'Primary Specialty': response.css('.normal::text').get(), 'Address': ' '.join([i.strip() for i in response.css('.office-address span::text').getall()]), 'Practice': response.css('.years-in-service::text').get(), 'Education': education.xpath('following-sibling::span[1]/text()').extract(), 'Training': training.xpath('following-sibling::span[1]/text()').extract(), 'Consumer Feedback': response.css('.item-rating-container a::text').get() } def parse_reviews(self, response): # Process Reviews page data profile_name = response.css('h1::text').get() for idx, review in enumerate(response.css('.review-item')): yield { 'item_type': 'review', 'profile_name': profile_name, 'reviewer_name': review.css('.reviewer-name::text').get(), 'rating': review.css('.star-rating::attr(data-value)').get(), 'review_text': review.css('.review-body::text').get().strip() if review.css('.review-body::text').get() else None, 'review_date': review.css('.review-date::text').get() } def parse_contact_directions(self, response): # Process Contact & Directions page data profile_name = response.css('h1::text').get() yield { 'item_type': 'contact_directions', 'profile_name': profile_name, 'phone_number': response.css('.contact-phone::text').get().strip() if response.css('.contact-phone::text').get() else None, 'address': ' '.join([i.strip() for i in response.css('.directions-address span::text').getall()]), 'google_maps_link': response.css('.maps-link::attr(href)').get(), 'hours_of_operation': [hour.strip() for hour in response.css('.hours-list li::text').getall() if hour.strip()] }
Quick Tips to Get This Working
- Update Selectors: All the CSS/XPath selectors in the code above (like
.review-item,.contact-phone) are placeholders. Usescrapy shell "your-target-url"to test and tweak selectors until they correctly grab the data you need. - Check for AJAX: If tabs don’t load content in the initial DOM, open your browser’s DevTools > Network tab, click the tab, and look for the request that loads the data. You might be able to call that API directly instead of following the tab link.
- Handle Edge Cases: Add
strip()checks andifconditions to avoidNonevalues breaking your data output.
内容的提问来源于stack exchange,提问作者Alex

