You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium网页爬取结果与UI不一致,无法获取网站及参考编号

FCA注册页面爬取问题解决方案

1. 地址爬取冗余内容处理

原CSS选择器未绑定明确的标题上下文,导致匹配到DOM中存在但UI隐藏的冗余元素。改用基于标题关联的XPATH定位,精准获取可见地址:

# 定位"Registered address"标题,再取其紧邻的第一个段落元素
address_title = WebDriverWait(driver, 10).until(
    EC.visibility_of_element_located((By.XPATH, "//h4[text()='Registered address']"))
)
address = address_title.find_element(By.XPATH, "./following-sibling::p[1]").text
# 若仍有多余空行/无效内容,可过滤处理
# address = '\n'.join(line.strip() for line in address.split('\n') if line.strip())
print(address)

2. Website信息获取

Website位于"Other information"折叠面板内,需先展开面板再精准定位:

# 展开"Other information"折叠面板(用JS点击避免元素遮挡)
other_info_panel = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.XPATH, "//h3[text()='Other information']"))
)
driver.execute_script("arguments[0].click();", other_info_panel)

# 定位Website对应的内容
website = WebDriverWait(driver, 10).until(
    EC.visibility_of_element_located((By.XPATH, "//h4[text()='Website']/following-sibling::p[1]"))
).text
print(website)

3. Reference Number获取

直接通过标签关联的XPATH定位顶部的参考编号:

ref_number = WebDriverWait(driver, 10).until(
    EC.visibility_of_element_located((By.XPATH, "//span[text()='Reference number']/following-sibling::span[1]"))
).text
print(ref_number)

问题根源说明

  • 地址冗余:原选择器范围过大,匹配到了页面中隐藏的DOM节点;
  • Website失败:未触发折叠面板展开,且选择器未精准指向目标内容;
  • Reference Number:原代码未针对该元素的结构做针对性定位。

内容的提问来源于stack exchange,提问作者Masond3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 17:50:34