Webscraping点击按钮功能失效,请求技术协助
Hey there, let's figure out why your scraping script is freezing up when trying to click the "View" button. I'll break down the likely issues and show you how to fix them:
主要问题分析与修复步骤
重复创建浏览器实例,资源浪费还易触发反爬
Right now you're initializing a new Chrome driver every time inside the loop. That's not just inefficient—it can leave hanging browser processes, slow down your script, and even make the site flag you as suspicious. Move the driver setup outside the loop, and only quit it once all iterations are done.模糊的XPATH导致按钮定位失败
Your current XPATH for the View button (//tr[contains(.,'{}')]/td/a) is too vague. The site might have extra spaces, capitalization differences (like "Bac Giang" instead of lowercase), or other text in the row that throws off thecontainscheck. Instead, target the specific cell with the factory name first, then grab the corresponding View button in that row. For example:view_button_xpath = "//td[normalize-space(text())='{}']/following-sibling::td/a[text()='View']"The
normalize-spacehandles extra spaces, and specifying the button text makes the locator way more reliable.别再盲目用
time.sleep了
Fixed sleep times are a gamble—if the site loads slower than 3 or 5 seconds, your script will fail; if it loads faster, you're just wasting time. Replace those sleeps with explicit waits that wait for the exact condition you need, like waiting for the button to be clickable (not just visible):wait.until(EC.element_to_be_clickable((By.XPATH, view_button_xpath.format(input_text)))).click()空的
except块让你完全摸不到问题
Right now you're catching every error and just ignoring it withpass—that's why you have no clue where it's breaking! Replace that with code that prints the error, so you can see exactly what's going wrong (like a NoSuchElementException or TimeoutException):except Exception as e: print(f"Failed to process {i}: {str(e)}") driver.quit() # Make sure to clean up even if there's an error弹窗定位可能不准确
The linewait.until(EC.visibility_of_element_located((By.ID, "window01"))).click()is a bit unclear. Ifwindow01is a popup overlay or a close button, double-check that the ID is correct. You might need to wait for the popup to be interactive instead of just visible.
修改后的代码示例
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd matched_suppliers = ['bac giang lng garment corporation'] length = len(matched_suppliers) aggregator = [] # Initialize driver once outside the loop driver = webdriver.Chrome(executable_path=r"C:\webdrivers\chromedriver.exe") wait = WebDriverWait(driver, 20) try: for idx, factory_name in enumerate(matched_suppliers): rem = length - idx print(f'This is index: {idx}, element: {factory_name}, with remaining : {rem} elements') driver.get('https://portal.betterwork.org/transparency/compliance') # Wait for loaders to disappear wait.until(EC.invisibility_of_element((By.ID, "loader-wrapper"))) wait.until(EC.invisibility_of_element((By.CSS_SELECTOR, "div.k-loading-mask"))) # Open advanced search and input factory name driver.find_element(By.ID, "pnlHdAdvanceSearch").click() search_input = driver.find_element(By.ID, "txtSearchFactory") search_input.clear() # Clear previous input just in case search_input.send_keys(factory_name) driver.find_element(By.ID, "btnSearchData").click() # Wait for results and click View button with precise locator view_button_xpath = "//td[normalize-space(text())='{}']/following-sibling::td/a[text()='View']" wait.until(EC.element_to_be_clickable((By.XPATH, view_button_xpath.format(factory_name.lower())))).click() # Handle popup (adjust if window01 is not the right element) wait.until(EC.element_to_be_clickable((By.ID, "window01"))).click() # Wait for the issues grid to load and extract text wait.until(EC.visibility_of_element_located((By.ID, "gridInfoList"))) issue_text = driver.find_element(By.ID, "gridInfoList").text print(f"Issues for {factory_name}:\n{issue_text}") aggregator.append([factory_name, issue_text]) finally: # Make sure driver is closed no matter what driver.quit() # Create DataFrame bw_issues = pd.DataFrame(aggregator, columns=['Factory', 'Reason']) print(bw_issues)
内容的提问来源于stack exchange,提问作者Filippo Sebastio

