如何获取目标网站的XPath表达式?并爬取简历页面提取技能信息
Hey there! Let's break down how to extract skill information from software engineer resumes on that LiveCareer search page. First, a quick heads-up: this page loads content dynamically (via JavaScript), so basic tools like requests + lxml might not capture all resumes—you’ll likely need a browser automation tool like Selenium or Playwright to handle the dynamic loading.
First, we need to select all the resume containers on the page. From inspecting the page structure, each resume is typically wrapped in a card element. A reliable XPath to grab all these cards would be:
//div[contains(@class, 'resume-item') or contains(@class, 'resume-card')]
This targets any div that has either "resume-item" or "resume-card" in its class attribute—adjust the class names if the page updates its structure (always double-check via browser dev tools!).
Once you have a resume card, use a relative XPath to pull out the skills section. Here are two common scenarios:
Scenario 1: Skills are in a list (bulleted items)
If skills are listed as <li> elements inside a <ul> within a skills section:
.//div[contains(@class, 'resume-skills')]/ul/li/text()
The . at the start makes this a relative path, so it only looks inside the current resume card.
Scenario 2: Skills are comma-separated text
If skills are grouped into a single text block separated by commas:
.//div[contains(@class, 'resume-skills')]/text()
You’ll need to split the resulting text by commas and clean up whitespace afterward.
- Handle Dynamic Loading: The page loads more resumes as you scroll. To capture all results, simulate scrolling to the bottom of the page repeatedly until no new content loads, or wait for AJAX requests to complete.
- Respect Anti-Scraping Rules: LiveCareer may block frequent requests. Use a realistic
User-Agentheader, add delays between requests, and consider rotating proxies if you’re scraping at scale. - Validate XPath Regularly: Website structures change over time. Always verify your XPath expressions using browser dev tools (right-click → Inspect → Ctrl+F to test XPath in the Elements tab).
Here’s a quick Python snippet to put this into action:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # Initialize browser driver = webdriver.Chrome() driver.get("https://www.livecareer.com/resume-search/search?jt=software%20engineer") # Wait for initial resumes to load WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'resume-card')]")) ) # Simulate scrolling to load more resumes (repeat a few times) for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # Extract all resume cards resume_cards = driver.find_elements(By.XPATH, "//div[contains(@class, 'resume-card')]") # Loop through each card and extract skills for idx, card in enumerate(resume_cards, 1): # Get skills (adjust XPath based on actual page structure) skill_elements = card.find_elements(By.XPATH, ".//div[contains(@class, 'resume-skills')]/ul/li") skills = [skill.text.strip() for skill in skill_elements if skill.text.strip()] print(f"Resume {idx} Skills: {', '.join(skills) if skills else 'No skills listed'}") driver.quit()
内容的提问来源于stack exchange,提问作者HaAbs

