使用Selenium无法抓取动态网站人物链接的技术求助
问题
尝试爬取网站https://www.dlapiper.com/en-us/people#t=All&sort=relevancy&numberOfResults=100&f:CountriesID=[United%20Kingdom]上英国地区的个人资料及联系人链接,使用Selenium+ChromeDriver实现。常规脚本在其他动态网站正常运行,但在此网站无法获取人物搜索区块的链接——Chrome开发者工具中能看到这些链接元素,但延长time.sleep时长也无效。原脚本如下:
links = [] driver = webdriver.Chrome() driver.get('https://www.dlapiper.com/en-gb/people#t=All&sort=%40lastname%20ascending&f:CountriesID=[United%20Kingdom]') time.sleep(5) cookies_button = driver.find_element(By.ID, "onetrust-reject-all-handler") cookies_button.click() time.sleep(5) html = driver.page_source time.sleep(5) soup = BeautifulSoup(html, 'html.parser') parse = soup.find_all('a') for item in parse: links.append(item.get('href')) print(links)
解决思路与修正代码
核心问题分析
time.sleep是固定时长等待,无法适配页面动态加载的实际节奏,可能元素还未渲染完成就执行了后续代码。- 直接提取
page_source并用BeautifulSoup解析,可能抓不到JavaScript动态渲染的元素——部分动态内容不会立即写入静态源码。
修正方案
改用显式等待替代固定sleep,直接通过Selenium定位目标元素,同时处理页面滚动加载逻辑(该网站需滚动到底部才能加载所有100条结果):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time links = [] driver = webdriver.Chrome() driver.get('https://www.dlapiper.com/en-gb/people#t=All&sort=%40lastname%20ascending&f:CountriesID=[United%20Kingdom]') # 显式等待并点击拒绝cookie按钮 wait = WebDriverWait(driver, 15) cookie_btn = wait.until(EC.element_to_be_clickable((By.ID, "onetrust-reject-all-handler"))) cookie_btn.click() # 滚动加载所有内容:循环滚动到底部,直到没有新元素加载 last_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待加载新内容 time.sleep(2) # 检查页面高度是否变化,判断是否加载完成 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 显式等待目标人物链接加载完成,直接通过Selenium定位 person_links = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "a[href*='/en-gb/people/']"))) # 提取链接并去重 for link in person_links: href = link.get_attribute('href') if href not in links: links.append(href) print(f"共获取到{len(links)}条人物链接") print(links) driver.quit()
关键说明
- 显式等待:通过
WebDriverWait等待元素可交互或存在,比固定sleep更可靠,避免因加载延迟导致的元素未找到问题。 - 滚动加载处理:该网站采用滚动触发加载机制,必须滚动到底部才能渲染所有100条结果,否则只能获取初始加载的少量内容。
- 直接用Selenium定位:跳过BeautifulSoup解析静态源码的步骤,直接获取动态渲染后的DOM元素,确保能拿到目标链接。
内容的提问来源于stack exchange,提问作者faintglimmer
相关产品推荐
相关产品推荐

