使用Python+Selenium抓取Seek远程职位位置遇超时,求解决方案
Hey there, let's work through why your Selenium script is timing out when trying to grab location tags from Seek's Work From Home jobs page. Here are the key issues and fixes:
1. Your XPath Has Syntax Errors
Looking at your XPath string, you've got messed-up double quotes (like @id=""app"") which break the locator entirely. In Python, if you wrap your XPath in double quotes, internal double quotes need to be escaped with \", or you can wrap the XPath in single quotes to avoid this confusion entirely.
Broken XPath snippet:
"//*[@id=""app""]/div/div/div[4]/div/div[3]/section/div[2]/div/div[2]/div[""1]/div/div[2]/div/div[1]/div[2]/article/div[1]/span[2]/span/strong/span/span"
2. Absolute XPath Is Extremely Fragile
Your XPath uses a hardcoded absolute path with fixed indexes (like div[""1]), which is super brittle. If Seek updates their page layout even slightly—whether it's a new container div or a shifted element—this locator will fail immediately. Instead, use relative locators that target job cards and their location elements directly.
3. Corrected Code Example
Here's a revised version of your script that fixes these issues and uses more reliable locators:
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service # Initialize Chrome driver with modern Selenium syntax service = Service(r"C:\Program Files (x86)\Google\Chrome\Application\chromedriver.exe") options = webdriver.ChromeOptions() # Add basic anti-detection tweaks to avoid being blocked by Seek options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(service=service, options=options) driver.get("https://www.seek.com.au/jobs?where=Work%20from%20home") assert "SEEK" in driver.title # Wait for the main job list container to load first (more reliable than targeting individual elements) WebDriverWait(driver, 25).until( EC.presence_of_element_located((By.CLASS_NAME, '_1wkzzau0')) ) # Use a relative XPath to target ALL location elements in job cards # This locator targets spans with location-specific classes (verified via Chrome DevTools) locations = WebDriverWait(driver, 25).until( EC.visibility_of_all_elements_located( (By.XPATH, '//div[contains(@class, "job-card")]//span[contains(@class, "yvsb870") and contains(@class, "yvsb874")]') ) ) # Extract and print location text for loc in locations: print(loc.text) driver.quit()
Additional Tips
- Verify Locators with Chrome DevTools: Right-click the location element on Seek's page → "Inspect" → Right-click the element in DevTools → "Copy" → "Copy XPath" (always prefer relative XPath over absolute when possible).
- Handle Infinite Scroll: Seek loads jobs as you scroll. If you need more than the initial batch, add code to scroll to the bottom of the page and wait for new jobs to load before scraping.
- Anti-Scraping Workarounds: For larger-scale scraping, you might need to add random delays between actions or rotate user agents to avoid being flagged by Seek's anti-scraping systems.
内容的提问来源于stack exchange,提问作者matthew1992

