Selenium+Python无头Chrome调用driver.get()无限等待问题求助
driver.get() for tsetmc.com Let’s break down what’s happening here and walk through targeted fixes—since your setup works perfectly for codal.ir, the issue is almost certainly tied to how tsetmc.com interacts with headless Chrome or gaps in your browser configuration.
Common Causes & Practical Fixes
1. Verify Chrome/Chromium and ChromeDriver Version Compatibility
Headless mode is far more sensitive to version mismatches than regular GUI mode. First, confirm your browser and driver versions match exactly:
- Run
chromium-browser --version(orgoogle-chrome --version) on your server to get the browser version. - Ensure your ChromeDriver is built for that exact version—even minor patch differences can cause stalls or crashes.
2. Upgrade to the New Headless Mode + Mimic a Real Browser
Older headless mode has distinct behavioral quirks that many sites detect and block. Replace the old --headless flag with the newer, more stealthy version, plus add flags to make your headless browser act like a regular user:
chrome_options.add_argument("--headless=new") # Chrome 112+ required; behaves like regular Chrome chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--window-size=1920,1080") chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False)
These tweaks bypass most basic anti-headless checks that might be stalling the page load.
3. Replace Static Sleep with Explicit Waits
time.sleep(10) is unreliable and hides what’s actually happening. Use Selenium’s explicit waits to wait for a critical page element, which helps diagnose if the page is stuck loading or just slow:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException # After driver.get() try: # Wait for the page body to load (adjust selector if you need a specific element) WebDriverWait(driver, 30).until( EC.presence_of_element_located((By.TAG_NAME, "body")) ) print("Page loaded successfully") except TimeoutException: print("Timed out waiting for page") # Capture a screenshot to see what the headless browser is stuck on driver.save_screenshot("tsetmc_timeout.png")
The screenshot can reveal captchas, loading spinners, or blocked resources you wouldn’t see otherwise.
4. Rule Out Profile or Dependency Issues
- Temporarily remove the
user-data-dir=seleniumflag—cached cookies or session data in the profile might be causing conflicts with headless mode. - On your Ubuntu Server, install missing libraries that Chrome needs (even for headless):
sudo apt-get install -y libxss1 libappindicator1 libindicator7 libgconf-2-4 - Specify the Chrome binary path explicitly to avoid executable confusion:
chrome_options.binary_location = "/usr/bin/chromium-browser" # Or "/usr/bin/google-chrome"
Modified Working Code Example
Here’s your code updated with all the recommended fixes:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.common.exceptions import TimeoutException chrome_options = Options() # Core headless configuration chrome_options.add_argument("--headless=new") chrome_options.add_argument('--no-sandbox') chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--window-size=1920,1080") # Anti-detection tweaks chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # Optional: Uncomment to test without saved profile # chrome_options.add_argument("user-data-dir=selenium") print("Opening browser") driver = webdriver.Chrome("/usr/lib/chromium-browser/chromedriver", options=chrome_options) print("Sending request to tsetmc.com") driver.get("http://www.tsetmc.com/Loader.aspx?ParTree=15131F") try: WebDriverWait(driver, 30).until( EC.presence_of_element_located((By.TAG_NAME, "body")) ) print("Page loaded successfully") response = driver.page_source print("Got page source, quitting...") except TimeoutException: print("Timed out waiting for page load") driver.save_screenshot("tsetmc_timeout.png") finally: driver.quit()
Final Debugging Step
If the issue persists, enable verbose logging to get granular details about what’s going on during the driver.get() call:
chrome_options.add_argument("--verbose") chrome_options.add_argument("--log-path=chrome_headless.log")
Check the chrome_headless.log file for errors related to network requests, blocked resources, or browser initialization failures.
内容的提问来源于stack exchange,提问作者Arman Babaei

