使用Python Selenium抓取动态网站职位数据遇问题求助
职位抓取问题求助
我的目标是抓取每个职位卡片的信息构建数据库,计划执行以下步骤:
- 获取网站的最大页数;
- 提取每个职位卡片的ID,通过修改基础URL访问单个职位页面;
- 将数据保存到pandas CSV文件或SQL数据库。
目前我尝试抓取第一页的10个职位卡片的标题和ID,但代码要么返回空列表,要么抛出NoSuchElementException错误。以下是我的代码:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import pandas as pd #Instantiate the webdriver driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) #Define url url = 'https://iefponline.iefp.pt/IEFP/pesquisas/search.do' # load the web page driver.get(url) # set maximun time to load the page in seconds driver.implicitly_wait(15) #collect data that are withing the main ID block contents = driver.find_element(By.ID, 'resultados-pesquisa') # Find all elements with the class name 'offer-card horizontal' emp_offers = contents.find_elements(By.CLASS_NAME, 'offer-card') emp_title_list = [] emp_id_list = [] for emp_offer in emp_offers: offer_title = emp_offer.get_attribute('title') emp_title_list.append(offer_title) offer_id = emp_offer.find_element(By.XPATH, './/div[contains(@class, "offer-code")]/span[2]').text emp_id_list.append(offer_id) print(emp_title_list) print(emp_id_list) # Close the WebDriver driver.quit()
返回结果要么是:
['', '', '', '', '', '', '', '', '', ''] [None, None, None, None, None, None, None, None, None, None]
要么抛出报错信息:
"DevTools listening on ws://127.0.0.1:65003/devtools/browser/327b17e5-a97d-4d84-9ae0-c1c03122286a Traceback (most recent call last): File "c:\Users\dbelt\Documents\scrape\selenium_iefp.py", line 36, in <module> offer_id = emp_offer.find_element(By.XPATH, './/div[contains(@class, "offer-code")]/span[2]').text File "C:\Users\dbelt\anaconda3\lib\site-packages\selenium\webdriver\remote\webelement.py", line 416, in find_element return self._execute(Command.FIND_CHILD_ELEMENT, {"using": by, "value": value})["value"] File "C:\Users\dbelt\anaconda3\lib\site-packages\selenium\webdriver\remote\webelement.py", line 394, in _execute return self._parent.execute(command, params) File "C:\Users\dbelt\anaconda3\lib\site-packages\selenium\webdriver\remote\webdriver.py", line 344, in execute self.error_handler.check_response(response) File "C:\Users\dbelt\anaconda3\lib\site-packages\selenium\webdriver\remote\errorhandler.py", line 229, in check_response raise exception_class(message, screen, stacktrace) selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"xpath","selector":".//div[contains(@class, "offer-code")]/span[2]"} (Session info: chrome=117.0.5938.134); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception Stacktrace: GetHandleVerifier [0x0062CFE3+45267] (No symbol) [0x005B9741] (No symbol) [0x004ABE1D] (No symbol) [0x004DED30] (No symbol) [0x004DF1FB] (No symbol) [0x004D8041] (No symbol) [0x004FB084] (No symbol) [0x004D7F96] (No symbol) [0x004FB2B4] (No symbol) [0x0050DDDA] (No symbol) [0x004FAE36] (No symbol) [0x004D674E] (No symbol) [0x004D78ED] GetHandleVerifier [0x008E5659+2897737] GetHandleVerifier [0x0092E78B+3197051] GetHandleVerifier [0x00928571+3171937] GetHandleVerifier [0x006B5E40+606000] (No symbol) [0x005C338C] (No symbol) [0x005BF508] (No symbol) [0x005BF62F] (No symbol) [0x005B1D27] BaseThreadInitThunk [0x757B7BA9+25] RtlInitializeExceptionChain [0x7711B79B+107] RtlClearBits [0x7711B71F+191]"
此外,我发现该网站的XPATH路径异常冗长,且很多类名包含空格,不清楚如何在find_elements中正确使用这类类名,特此求助。
解决方案
1. 替换等待策略,确保元素完全加载
隐式等待无法应对动态渲染的页面,改用显式等待精准等待目标元素加载完成,避免过早定位导致的空值或报错:
- 等待主结果区块
resultados-pesquisa出现 - 等待所有职位卡片加载完毕
2. 正确处理带空格的类名
By.CLASS_NAME仅支持单个类名匹配,遇到offer-card horizontal这类多类名元素,可选择两种方式:
- 选取其中一个唯一类名(如
offer-card) - 使用
By.CSS_SELECTOR,语法为.offer-card.horizontal(多个类名用.连接)
3. 修正标题与ID的定位逻辑
- 职位标题:不要直接取元素的
title属性,而是定位卡片内的标题文本元素(如.offer-title) - 职位ID:改用更简洁可靠的CSS选择器替代冗长XPATH,避免层级偏移问题
修正后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd # 初始化浏览器驱动 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) url = 'https://iefponline.iefp.pt/IEFP/pesquisas/search.do' driver.get(url) # 初始化显式等待对象,最长等待20秒 wait = WebDriverWait(driver, 20) # 等待主结果区块加载完成 contents = wait.until(EC.presence_of_element_located((By.ID, 'resultados-pesquisa'))) # 等待所有职位卡片加载完成,用CSS选择器定位多类名元素 emp_offers = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, '.offer-card.horizontal'))) emp_title_list = [] emp_id_list = [] for emp_offer in emp_offers: # 定位标题元素获取文本 offer_title = emp_offer.find_element(By.CSS_SELECTOR, '.offer-title').text emp_title_list.append(offer_title) # 定位职位ID,用CSS选择器精准匹配 offer_id = emp_offer.find_element(By.CSS_SELECTOR, '.offer-code span:nth-child(2)').text emp_id_list.append(offer_id) # 打印结果 print(emp_title_list) print(emp_id_list) # 保存数据到CSV文件 df = pd.DataFrame({'职位标题': emp_title_list, '职位ID': emp_id_list}) df.to_csv('职位信息.csv', index=False, encoding='utf-8-sig') # 关闭浏览器 driver.quit()
额外提示
- 若页面存在懒加载,需添加滚动页面的逻辑,确保所有卡片加载完毕
- 抓取多页时,定位分页按钮元素,循环点击并重复抓取逻辑
- 添加随机延迟(如
time.sleep(random.uniform(1,3))),避免请求过于频繁触发反爬机制
内容的提问来源于stack exchange,提问作者dbeltranor
相关产品推荐
相关产品推荐

