You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup+Selenium爬取TCS职位页时无法获取完整HTML

问题分析与解决方法

核心问题原因

  • 页面未完全加载:调用driver.get(url)后立即获取page_source,但目标页面是AngularJS渲染的动态页面,异步数据还没加载完成,此时DOM中还没有你要找的div.row.custom-row.searched-job.ng-scope元素。
  • 定位选择器不稳定:ng-scope是Angular自动生成的动态类,并非页面固定标识,依赖这个类定位元素容易失效。

修改后的代码

import requests
from bs4 import BeautifulSoup as bs
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome(executable_path=r'D:\Web Scraping\chromedriver.exe')

url = "https://ibegin.tcs.com/iBegin/jobs/search"

driver.get(url)

# 等待目标元素加载完成,最长等待10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'div.row.custom-row.searched-job')))

# 此时页面已加载完成,再获取源码解析
soup = bs(driver.page_source, 'html.parser')

# 使用更稳定的类选择器,去掉动态的ng-scope
target_divs = soup.find_all('div', {'class': 'row custom-row searched-job'})

# 遍历输出结果
for div in target_divs:
    print(div.prettify())

driver.quit()

关键修改说明

  • 引入WebDriverWait显式等待,确保目标元素存在后再进行解析,避免因异步加载导致的元素缺失。
  • 移除动态的ng-scope类作为定位条件,保留页面固定的row custom-row searched-job类,提升定位稳定性。
  • 用find_all替代find,可以获取所有符合条件的元素,避免单个元素查找失败返回空的情况。

内容的提问来源于stack exchange,提问作者Taha Habib

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 02:10:57