You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium抓取Seek远程职位位置遇超时,求解决方案

Fixing Selenium Timeout When Locating Job Locations on Seek

Hey there, let's work through why your Selenium script is timing out when trying to grab location tags from Seek's Work From Home jobs page. Here are the key issues and fixes:

1. Your XPath Has Syntax Errors

Looking at your XPath string, you've got messed-up double quotes (like @id=""app"") which break the locator entirely. In Python, if you wrap your XPath in double quotes, internal double quotes need to be escaped with \", or you can wrap the XPath in single quotes to avoid this confusion entirely.

Broken XPath snippet:

"//*[@id=""app""]/div/div/div[4]/div/div[3]/section/div[2]/div/div[2]/div[""1]/div/div[2]/div/div[1]/div[2]/article/div[1]/span[2]/span/strong/span/span"

2. Absolute XPath Is Extremely Fragile

Your XPath uses a hardcoded absolute path with fixed indexes (like div[""1]), which is super brittle. If Seek updates their page layout even slightly—whether it's a new container div or a shifted element—this locator will fail immediately. Instead, use relative locators that target job cards and their location elements directly.

3. Corrected Code Example

Here's a revised version of your script that fixes these issues and uses more reliable locators:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service

# Initialize Chrome driver with modern Selenium syntax
service = Service(r"C:\Program Files (x86)\Google\Chrome\Application\chromedriver.exe")
options = webdriver.ChromeOptions()
# Add basic anti-detection tweaks to avoid being blocked by Seek
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(service=service, options=options)
driver.get("https://www.seek.com.au/jobs?where=Work%20from%20home")
assert "SEEK" in driver.title

# Wait for the main job list container to load first (more reliable than targeting individual elements)
WebDriverWait(driver, 25).until(
    EC.presence_of_element_located((By.CLASS_NAME, '_1wkzzau0'))
)

# Use a relative XPath to target ALL location elements in job cards
# This locator targets spans with location-specific classes (verified via Chrome DevTools)
locations = WebDriverWait(driver, 25).until(
    EC.visibility_of_all_elements_located(
        (By.XPATH, '//div[contains(@class, "job-card")]//span[contains(@class, "yvsb870") and contains(@class, "yvsb874")]')
    )
)

# Extract and print location text
for loc in locations:
    print(loc.text)

driver.quit()

Additional Tips

  • Verify Locators with Chrome DevTools: Right-click the location element on Seek's page → "Inspect" → Right-click the element in DevTools → "Copy" → "Copy XPath" (always prefer relative XPath over absolute when possible).
  • Handle Infinite Scroll: Seek loads jobs as you scroll. If you need more than the initial batch, add code to scroll to the bottom of the page and wait for new jobs to load before scraping.
  • Anti-Scraping Workarounds: For larger-scale scraping, you might need to add random delays between actions or rotate user agents to avoid being flagged by Seek's anti-scraping systems.

内容的提问来源于stack exchange,提问作者matthew1992

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:02:36