You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium抓取JustDial仅获10条结果,请求代码问题排查

解决JustDial动态加载爬虫仅获取前10条结果的问题

问题根源

你的代码存在以下几个核心问题,导致无法加载全部结果:

  • 固定的time.sleep(5)无法适配不同网络/页面加载速度,经常在新内容或加载按钮未就绪时就执行下一步
  • 未等待"加载更多"按钮处于可交互状态就尝试点击,容易触发异常提前终止循环
  • 缺少对页面内容加载完成的校验逻辑

修复后的代码

import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

start_time = time.time()
url = "https://www.justdial.com/Pune/Stationery-Shops/nct-10453443"
driver = webdriver.Chrome()
driver.get(url)
wait = WebDriverWait(driver, 10)  # 设置10秒超时时间,可按需调整

while True:
    # 滚动至页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    
    try:
        # 等待"加载更多"按钮可点击后再执行点击操作
        more_button = wait.until(EC.element_to_be_clickable((By.ID, "moreH")))
        more_button.click()
        # 点击后短暂等待内容渲染
        time.sleep(3)
    except:
        # 无更多内容可加载时退出循环
        break

# 解析最终页面内容
html = driver.page_source
soup = BeautifulSoup(html, "html.parser")

# 提取并打印结果
for item in soup.find_all("a", {"class": "resultbox_title_anchor"}):
    print(item.text)

print(f"总耗时: {time.time()-start_time:.2f}秒")
driver.quit()  # 记得关闭浏览器释放资源

额外优化建议

  • 如果仍然无法加载全部内容,可尝试分段滚动(每次滚动页面高度的80%,重复几次后再滚到底),模拟人类浏览行为,避免触发反爬机制
  • 检查网站DOM结构是否更新,确认"加载更多"按钮的IDmoreH是否仍然有效
  • 可添加随机休眠时间(如time.sleep(random.uniform(2,5))),进一步降低被识别为爬虫的概率

内容的提问来源于stack exchange,提问作者vinayak sony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 07:30:08