Selenium网页爬取结果与UI不一致,无法获取网站及参考编号
FCA注册页面爬取问题解决方案
1. 地址爬取冗余内容处理
原CSS选择器未绑定明确的标题上下文,导致匹配到DOM中存在但UI隐藏的冗余元素。改用基于标题关联的XPATH定位,精准获取可见地址:
# 定位"Registered address"标题,再取其紧邻的第一个段落元素 address_title = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//h4[text()='Registered address']")) ) address = address_title.find_element(By.XPATH, "./following-sibling::p[1]").text # 若仍有多余空行/无效内容,可过滤处理 # address = '\n'.join(line.strip() for line in address.split('\n') if line.strip()) print(address)
2. Website信息获取
Website位于"Other information"折叠面板内,需先展开面板再精准定位:
# 展开"Other information"折叠面板(用JS点击避免元素遮挡) other_info_panel = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//h3[text()='Other information']")) ) driver.execute_script("arguments[0].click();", other_info_panel) # 定位Website对应的内容 website = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//h4[text()='Website']/following-sibling::p[1]")) ).text print(website)
3. Reference Number获取
直接通过标签关联的XPATH定位顶部的参考编号:
ref_number = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//span[text()='Reference number']/following-sibling::span[1]")) ).text print(ref_number)
问题根源说明
- 地址冗余:原选择器范围过大,匹配到了页面中隐藏的DOM节点;
- Website失败:未触发折叠面板展开,且选择器未精准指向目标内容;
- Reference Number:原代码未针对该元素的结构做针对性定位。
内容的提问来源于stack exchange,提问作者Masond3
相关产品推荐
相关产品推荐

