求助:使用Selenium和Beautiful Soup抓取JS渲染医生表格的问题
问题描述
我尝试抓取网站:https://www.globusmedical.com/patient-education-musculoskeletal-system-conditions/resources/find-a-surgeon/
该网站采用JavaScript渲染机制,加载页面后查看源码无法找到医生列表表格,但直接检查对应元素时能看到完整的医生信息。
我希望多次点击“Load More”按钮直到按钮消失,再使用BeautifulSoup解析页面内容。
目前遇到两个问题:
- 打印的
page_source中没有任何医生相关信息,这是什么原因? - 该如何编写while循环实现重复点击“Load More”按钮直到按钮消失?
附上当前代码:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver import ActionChains import time import requests doctor_dict = {} #configure webdriver options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(options = options) driver.get("https://www.globusmedical.com/patient-education-musculoskeletal-system-conditions/resources/find-a-surgeon/") time.sleep(5) clickable = driver.find_element(By.XPATH,'//button[@class="js-eml-load-more-button eml-load-more-button eml-btn btn btn--primary"]') driver.execute_script("arguments[0].click();", clickable) # items = driver.find_element(By.CLASS_NAME,"eml-location grid--item") soup = BeautifulSoup(driver.page_source, 'html.parser') print(soup.prettify()) driver.quit()
问题解答
问题1:page_source无医生信息的原因
- 页面加载不充分:
time.sleep(5)的固定等待时间不一定能覆盖页面动态加载的耗时,尤其是headless模式下,部分资源加载可能更慢,导致初始医生列表还未渲染完成。 - 点击后未等待新内容加载:你点击按钮后立即获取
page_source,此时新的医生数据还没通过AJAX请求加载并渲染到页面上。 - 元素定位可能存在局限性:使用精确的class属性定位按钮,若页面动态渲染时class有细微变化,会导致点击逻辑失效,无法触发医生列表加载。
问题2:循环点击“Load More”直到按钮消失的实现方法
核心逻辑是在循环中尝试定位按钮,找到则点击并等待加载;找不到则说明加载完成,退出循环。用try-except捕获元素不存在的异常,避免程序中断。
修正后的完整代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException import time doctor_dict = {} # 配置浏览器选项 options = webdriver.ChromeOptions() options.add_argument("--headless=new") # 模拟真实浏览器UA,规避反爬检测 options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 禁用自动化特征检测 options.add_argument("--disable-blink-features=AutomationControlled") driver = webdriver.Chrome(options=options) try: driver.get("https://www.globusmedical.com/patient-education-musculoskeletal-system-conditions/resources/find-a-surgeon/") # 等待页面基础框架加载完成 time.sleep(3) # 循环点击Load More按钮 while True: try: # 用模糊class定位,避免精确匹配失效 load_more_btn = driver.find_element(By.XPATH, '//button[contains(@class, "js-eml-load-more-button")]') # 用JS执行点击,避免元素不可点击的问题 driver.execute_script("arguments[0].click();", load_more_btn) # 等待新内容加载完成,时间可根据网络情况调整 time.sleep(2) except NoSuchElementException: # 按钮消失,加载完成,退出循环 break # 解析所有加载完成的页面内容 soup = BeautifulSoup(driver.page_source, 'html.parser') # 定位所有医生卡片元素 doctor_cards = soup.find_all(class_="eml-location") # 提取医生信息(根据页面实际元素结构调整) for card in doctor_cards: name = card.find(class_="eml-location__name").get_text(strip=True) if card.find(class_="eml-location__name") else "未知" hospital = card.find(class_="eml-location__hospital").get_text(strip=True) if card.find(class_="eml-location__hospital") else "未知" doctor_dict[name] = hospital # 打印抓取结果 print("抓取到的医生信息:") for name, hospital in doctor_dict.items(): print(f"{name}: {hospital}") finally: # 确保浏览器进程关闭 driver.quit()
额外优化建议
- 替换
time.sleep为显式等待,更高效稳定:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待按钮可点击,超时时间10秒 load_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(@class, "js-eml-load-more-button")]')) ) - 可以添加页面滚动操作,确保按钮处于可视区域,避免点击失败:
driver.execute_script("arguments[0].scrollIntoView(true);", load_more_btn)
内容的提问来源于stack exchange,提问作者user3628240
相关产品推荐
相关产品推荐

