You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Webscraping点击按钮功能失效,请求技术协助

解决网页抓取中点击“View”按钮后停止运行的问题

Hey there, let's figure out why your scraping script is freezing up when trying to click the "View" button. I'll break down the likely issues and show you how to fix them:

主要问题分析与修复步骤

  • 重复创建浏览器实例,资源浪费还易触发反爬
    Right now you're initializing a new Chrome driver every time inside the loop. That's not just inefficient—it can leave hanging browser processes, slow down your script, and even make the site flag you as suspicious. Move the driver setup outside the loop, and only quit it once all iterations are done.

  • 模糊的XPATH导致按钮定位失败
    Your current XPATH for the View button (//tr[contains(.,'{}')]/td/a) is too vague. The site might have extra spaces, capitalization differences (like "Bac Giang" instead of lowercase), or other text in the row that throws off the contains check. Instead, target the specific cell with the factory name first, then grab the corresponding View button in that row. For example:

    view_button_xpath = "//td[normalize-space(text())='{}']/following-sibling::td/a[text()='View']"
    

    The normalize-space handles extra spaces, and specifying the button text makes the locator way more reliable.

  • 别再盲目用time.sleep了
    Fixed sleep times are a gamble—if the site loads slower than 3 or 5 seconds, your script will fail; if it loads faster, you're just wasting time. Replace those sleeps with explicit waits that wait for the exact condition you need, like waiting for the button to be clickable (not just visible):

    wait.until(EC.element_to_be_clickable((By.XPATH, view_button_xpath.format(input_text)))).click()
    
  • 空的except块让你完全摸不到问题
    Right now you're catching every error and just ignoring it with pass—that's why you have no clue where it's breaking! Replace that with code that prints the error, so you can see exactly what's going wrong (like a NoSuchElementException or TimeoutException):

    except Exception as e:
        print(f"Failed to process {i}: {str(e)}")
        driver.quit()  # Make sure to clean up even if there's an error
    
  • 弹窗定位可能不准确
    The line wait.until(EC.visibility_of_element_located((By.ID, "window01"))).click() is a bit unclear. If window01 is a popup overlay or a close button, double-check that the ID is correct. You might need to wait for the popup to be interactive instead of just visible.

修改后的代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

matched_suppliers = ['bac giang lng garment corporation']
length = len(matched_suppliers)
aggregator = []

# Initialize driver once outside the loop
driver = webdriver.Chrome(executable_path=r"C:\webdrivers\chromedriver.exe")
wait = WebDriverWait(driver, 20)

try:
    for idx, factory_name in enumerate(matched_suppliers):
        rem = length - idx
        print(f'This is index: {idx}, element: {factory_name}, with remaining : {rem} elements')
        
        driver.get('https://portal.betterwork.org/transparency/compliance')
        
        # Wait for loaders to disappear
        wait.until(EC.invisibility_of_element((By.ID, "loader-wrapper")))
        wait.until(EC.invisibility_of_element((By.CSS_SELECTOR, "div.k-loading-mask")))
        
        # Open advanced search and input factory name
        driver.find_element(By.ID, "pnlHdAdvanceSearch").click()
        search_input = driver.find_element(By.ID, "txtSearchFactory")
        search_input.clear()  # Clear previous input just in case
        search_input.send_keys(factory_name)
        driver.find_element(By.ID, "btnSearchData").click()
        
        # Wait for results and click View button with precise locator
        view_button_xpath = "//td[normalize-space(text())='{}']/following-sibling::td/a[text()='View']"
        wait.until(EC.element_to_be_clickable((By.XPATH, view_button_xpath.format(factory_name.lower())))).click()
        
        # Handle popup (adjust if window01 is not the right element)
        wait.until(EC.element_to_be_clickable((By.ID, "window01"))).click()
        
        # Wait for the issues grid to load and extract text
        wait.until(EC.visibility_of_element_located((By.ID, "gridInfoList")))
        issue_text = driver.find_element(By.ID, "gridInfoList").text
        
        print(f"Issues for {factory_name}:\n{issue_text}")
        aggregator.append([factory_name, issue_text])
        
finally:
    # Make sure driver is closed no matter what
    driver.quit()

# Create DataFrame
bw_issues = pd.DataFrame(aggregator, columns=['Factory', 'Reason'])
print(bw_issues)

内容的提问来源于stack exchange,提问作者Filippo Sebastio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 16:32:41