You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium无法在目标网页定位指定关键词问题求助

解决Selenium无法定位页面中存在的"Annual Audited Accounts"文本问题

常见问题原因及对应解决方案

1. 文本位于iframe内

部分网站会用iframe加载子内容,需先切换到对应iframe再执行查找操作:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
url = 'https://www.bursamalaysia.com/market_information/announcements/company_announcement/announcement_details?ann_id=3434107'
driver.get(url)

try:
    # 定位并切换到iframe(需根据页面实际iframe属性调整定位规则)
    iframe = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//iframe"))
    )
    driver.switch_to.frame(iframe)
    
    # 查找目标文本
    WebDriverWait(driver, 10).until(
        EC.text_to_be_present_in_element((By.TAG_NAME, "body"), "Annual Audited Accounts")
    )
    print("Found: Annual Audited Accounts")
    
    # 定位并打印PDF链接信息
    pdf_link = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//a[contains(@href, '.pdf')]"))
    )
    print(f"PDF链接已找到:{pdf_link.get_attribute('href')}")
    
except Exception as e:
    print(f"错误信息:{str(e)}")
    print("Text not found")
finally:
    # 切回主文档
    driver.switch_to.default_content()
    driver.quit()

2. 文本需滚动触发加载

部分内容需滚动到可见区域才会渲染,可先滚动页面再查找:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
url = 'https://www.bursamalaysia.com/market_information/announcements/company_announcement/announcement_details?ann_id=3434107'
driver.get(url)

try:
    # 滚动到页面底部触发内容加载
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    driver.implicitly_wait(3)
    
    WebDriverWait(driver, 10).until(
        EC.text_to_be_present_in_element((By.TAG_NAME, "body"), "Annual Audited Accounts")
    )
    print("Found: Annual Audited Accounts")
    
    # 定位PDF链接
    pdf_link = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, "//a[contains(text(), 'PDF') or contains(@href, '.pdf')]"))
    )
    print(f"PDF链接可访问:{pdf_link.get_attribute('href')}")
    
except Exception as e:
    print(f"错误信息:{str(e)}")
    print("Text not found")
finally:
    driver.quit()

3. 用JavaScript直接遍历查找文本

若Selenium内置方法失效,可通过JS遍历页面元素查找目标文本:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait

driver = webdriver.Chrome()
url = 'https://www.bursamalaysia.com/market_information/announcements/company_announcement/announcement_details?ann_id=3434107'
driver.get(url)

try:
    # 用JS查找文本是否存在
    text_found = WebDriverWait(driver, 10).until(lambda d: d.execute_script("""
        function checkText() {
            const allElements = document.getElementsByTagName('*');
            for (let elem of allElements) {
                if (elem.textContent.includes('Annual Audited Accounts')) {
                    return true;
                }
            }
            return false;
        }
        return checkText();
    """))
    
    if text_found:
        print("Found: Annual Audited Accounts")
        # 用JS获取PDF链接
        pdf_href = driver.execute_script("""
            const pdfLinks = document.querySelectorAll('a[href$=".pdf"]');
            return pdfLinks.length > 0 ? pdfLinks[0].href : null;
        """)
        print(f"PDF链接:{pdf_href}" if pdf_href else "未找到PDF链接")
    else:
        print("Text not found")
        
except Exception as e:
    print(f"错误信息:{str(e)}")
    print("Text not found")
finally:
    driver.quit()

4. 直接获取页面文本做匹配

跳过Selenium的等待条件,直接获取页面完整文本进行检查:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait

driver = webdriver.Chrome()
url = 'https://www.bursamalaysia.com/market_information/announcements/company_announcement/announcement_details?ann_id=3434107'
driver.get(url)

try:
    # 等待页面完全加载
    WebDriverWait(driver, 15).until(lambda d: d.execute_script("return document.readyState === 'complete'"))
    
    # 获取页面所有文本并检查
    full_text = driver.find_element(By.TAG_NAME, 'body').text
    if "Annual Audited Accounts" in full_text:
        print("Found: Annual Audited Accounts")
        # 查找PDF链接
        pdf_links = driver.find_elements(By.XPATH, "//a[contains(@href, '.pdf')]")
        print(f"找到PDF链接:{pdf_links[0].get_attribute('href')}" if pdf_links else "无PDF链接")
    else:
        print("Text not found")
        
except Exception as e:
    print(f"错误信息:{str(e)}")
    print("Text not found")
finally:
    driver.quit()

内容的提问来源于stack exchange,提问作者Aman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 00:12:50