You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取目标网站的XPath表达式?并爬取简历页面提取技能信息

Hey there! Let's break down how to extract skill information from software engineer resumes on that LiveCareer search page. First, a quick heads-up: this page loads content dynamically (via JavaScript), so basic tools like requests + lxml might not capture all resumes—you’ll likely need a browser automation tool like Selenium or Playwright to handle the dynamic loading.

1. XPath to Target Individual Resume Cards

First, we need to select all the resume containers on the page. From inspecting the page structure, each resume is typically wrapped in a card element. A reliable XPath to grab all these cards would be:

//div[contains(@class, 'resume-item') or contains(@class, 'resume-card')]

This targets any div that has either "resume-item" or "resume-card" in its class attribute—adjust the class names if the page updates its structure (always double-check via browser dev tools!).

2. XPath to Extract Skills from a Single Resume

Once you have a resume card, use a relative XPath to pull out the skills section. Here are two common scenarios:

Scenario 1: Skills are in a list (bulleted items)

If skills are listed as <li> elements inside a <ul> within a skills section:

.//div[contains(@class, 'resume-skills')]/ul/li/text()

The . at the start makes this a relative path, so it only looks inside the current resume card.

Scenario 2: Skills are comma-separated text

If skills are grouped into a single text block separated by commas:

.//div[contains(@class, 'resume-skills')]/text()

You’ll need to split the resulting text by commas and clean up whitespace afterward.

3. Practical Scraping Tips
  • Handle Dynamic Loading: The page loads more resumes as you scroll. To capture all results, simulate scrolling to the bottom of the page repeatedly until no new content loads, or wait for AJAX requests to complete.
  • Respect Anti-Scraping Rules: LiveCareer may block frequent requests. Use a realistic User-Agent header, add delays between requests, and consider rotating proxies if you’re scraping at scale.
  • Validate XPath Regularly: Website structures change over time. Always verify your XPath expressions using browser dev tools (right-click → Inspect → Ctrl+F to test XPath in the Elements tab).
Example Code (Selenium)

Here’s a quick Python snippet to put this into action:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# Initialize browser
driver = webdriver.Chrome()
driver.get("https://www.livecareer.com/resume-search/search?jt=software%20engineer")

# Wait for initial resumes to load
WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'resume-card')]"))
)

# Simulate scrolling to load more resumes (repeat a few times)
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)

# Extract all resume cards
resume_cards = driver.find_elements(By.XPATH, "//div[contains(@class, 'resume-card')]")

# Loop through each card and extract skills
for idx, card in enumerate(resume_cards, 1):
    # Get skills (adjust XPath based on actual page structure)
    skill_elements = card.find_elements(By.XPATH, ".//div[contains(@class, 'resume-skills')]/ul/li")
    skills = [skill.text.strip() for skill in skill_elements if skill.text.strip()]
    
    print(f"Resume {idx} Skills: {', '.join(skills) if skills else 'No skills listed'}")

driver.quit()

内容的提问来源于stack exchange,提问作者HaAbs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 13:42:53