Python Scrapy技术求助:无ID标签的<table>数据提取
Hey there! Since you're new to Python and Scrapy, let's break down how to grab that specific "Wheelchair accessible" data from those 8 tables with class sc-fHxwqH ddWfJE step by step.
Step 1: Target All the Tables First
First, we need to select all 8 tables using Scrapy's selectors. CSS selectors are great for this since we're targeting a specific class combination:
tables = response.css('table.sc-fHxwqH.ddWfJE')
This will give you a list of all the matching tables on the page.
Step 2: Find Rows with "Wheelchair accessible"
Next, loop through each table and look for the row (<tr>) that contains a <td> with the exact text "Wheelchair accessible". XPath works really well here for this conditional match:
for table in tables: # Locate rows where one td has our target text target_rows = table.xpath('.//tr[td[text()="Wheelchair accessible"]]')
If the text might have extra spaces (like Wheelchair accessible ), use contains() instead to make the match more flexible:
target_rows = table.xpath('.//tr[td[contains(text(), "Wheelchair accessible")]]')
Step 3: Extract the Corresponding Data
Once you've got the target row, extract the data from the adjacent <td> (assuming the structure is <tr><td>Wheelchair accessible</td><td>Yes/No</td></tr>). Here's how to grab that value:
for row in target_rows: # Get the text from the second td in the row wheelchair_status = row.css('td:nth-child(2)::text').get() if wheelchair_status: # Clean up whitespace and yield the result yield { 'wheelchair_accessible': wheelchair_status.strip() }
Full Example in Your Spider's Parse Method
Putting it all together, your parse function might look like this:
import scrapy class YourSpider(scrapy.Spider): name = 'wheelchair_spider' start_urls = ['https://your-target-website.com'] def parse(self, response): # Grab all target tables tables = response.css('table.sc-fHxwqH.ddWfJE') for table in tables: # Find rows with our target text target_rows = table.xpath('.//tr[td[text()="Wheelchair accessible"]]') for row in target_rows: # Extract and clean the status status = row.css('td:nth-child(2)::text').get() if status: yield { 'wheelchair_accessible': status.strip() }
Pro Tip for Debugging
Use Scrapy Shell to test your selectors before writing full code! Run this in your terminal:
scrapy shell "https://your-target-website.com"
Then you can test each selector step by step (e.g., len(tables) to confirm you have 8 tables, target_rows.get() to see if the row is found) to fix any issues quickly.
内容的提问来源于stack exchange,提问作者sspaeti

