Python/Selenium无头模式下MacroTrends下载CSV文件失败求助
问题:无头模式下Selenium无法下载MacroTrends的CSV文件
尝试通过Python/Selenium/ChromeDriver在无头模式下,点击MacroTrends网站的「Download Data」按钮下载CSV文件。手动下载时虽有控制台错误Uncaught TypeError: Cannot read properties of undefined (reading 'goal')但能成功获取文件,然而代码执行后断言文件不存在。
原代码
import time from pathlib import Path from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait CWD = Path.cwd() chrome_options = webdriver.ChromeOptions() chrome_prefs = { "download.default_directory": f"{CWD}", "safebrowsing.enabled": False, "profile.default_content_settings.popups": 0, "download.prompt_for_download": False, "download.directory_upgrade": True, } chrome_options.add_experimental_option("prefs", chrome_prefs) chrome_options.add_argument("--headless") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--disable-extensions") ticker = "META" url = f"https://www.macrotrends.net/assets/php/stock_price_history.php?t={ticker}" with webdriver.Chrome(options=chrome_options) as driver: start_time = time.time() print(f"url: {url}") driver.get(url) download_button = None try: wd_wait = WebDriverWait(driver, 10) download_button = wd_wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "chart_buttons"))) time_diff = time.time() - start_time print(f"page load done in {time_diff:.2f} seconds.") except: print("ERROR: page load timeout!") raise SystemExit() start_time = time.time() print(f"download_button type: {type(download_button)}") driver.execute_script("arguments[0].click();", download_button) time_diff = time.time() - start_time print(f"click done in {time_diff:.2f} seconds.") csv_file = CWD.joinpath(f"MacroTrends_Data_Download_{ticker}.csv") assert csv_file.is_file(), f"{csv_file} not found!" print(f"download done in {time_diff:.2f} seconds.")
运行日志
python .\mtrends_download.py DevTools listening on ws://127.0.0.1:59481/devtools/browser/0ef013e3-b442-45d8-b96b-a230e3b27047 url: https://www.macrotrends.net/assets/php/stock_price_history.php?t=META page load done in 1.34 seconds. download_button type: <class 'selenium.webdriver.remote.webelement.WebElement'> [0610/160508.114:INFO:CONSOLE(344)] "Uncaught TypeError: Cannot read properties of undefined (reading 'goal')", source: https://www.macrotrends.net/assets/php/stock_price_history.php?t=META (344) click done in 0.20 seconds. Traceback (most recent call last): File "{SCRIPT_PATH}\mtrends_download.py", line 47, in <module> assert csv_file.is_file(), f"{csv_file} not found!" AssertionError: {SCRIPT_PATH}\MacroTrends_Data_Download_META.csv not found!
解决方案
1. 精确定位下载按钮
原代码定位的是按钮的父容器chart_buttons,而非实际触发下载的<a>标签。改用更精准的定位方式:
- 使用
By.LINK_TEXT直接匹配按钮文本:
download_button = wd_wait.until(EC.element_to_be_clickable((By.LINK_TEXT, "Download Data")))
- 或使用
By.XPATH定位:
download_button = wd_wait.until(EC.element_to_be_clickable((By.XPATH, "//div[@class='chart_buttons']/a[contains(text(), 'Download Data')]")))
2. 增加文件下载等待逻辑
无头模式下文件下载需要一定时间,直接断言会因文件未完成写入而失败。替换原断言为循环等待逻辑:
csv_file = CWD.joinpath(f"MacroTrends_Data_Download_{ticker}.csv") # 最多等待30秒,直到文件出现 max_wait = 30 wait_elapsed = 0 while not csv_file.is_file() and wait_elapsed < max_wait: time.sleep(1) wait_elapsed += 1 assert csv_file.is_file(), f"{csv_file} 下载超时或未找到!" print(f"download done in {time.time() - start_time:.2f} seconds.")
3. 优化无头模式配置
添加窗口尺寸参数,避免因无头模式默认视口过小导致元素交互异常;同时使用新版无头模式提升兼容性:
chrome_options.add_argument("--headless=new") # 新版无头模式兼容性更好 chrome_options.add_argument("--window-size=1920,1080") # 设置窗口尺寸
修改后的完整代码
import time from pathlib import Path from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait CWD = Path.cwd() chrome_options = webdriver.ChromeOptions() chrome_prefs = { "download.default_directory": f"{CWD}", "safebrowsing.enabled": False, "profile.default_content_settings.popups": 0, "download.prompt_for_download": False, "download.directory_upgrade": True, } chrome_options.add_experimental_option("prefs", chrome_prefs) chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--disable-extensions") chrome_options.add_argument("--window-size=1920,1080") ticker = "META" url = f"https://www.macrotrends.net/assets/php/stock_price_history.php?t={ticker}" with webdriver.Chrome(options=chrome_options) as driver: start_time = time.time() print(f"url: {url}") driver.get(url) download_button = None try: wd_wait = WebDriverWait(driver, 10) download_button = wd_wait.until(EC.element_to_be_clickable((By.LINK_TEXT, "Download Data"))) time_diff = time.time() - start_time print(f"page load done in {time_diff:.2f} seconds.") except: print("ERROR: page load timeout!") raise SystemExit() start_time = time.time() print(f"download_button type: {type(download_button)}") download_button.click() time_diff = time.time() - start_time print(f"click done in {time_diff:.2f} seconds.") # 等待文件下载完成 csv_file = CWD.joinpath(f"MacroTrends_Data_Download_{ticker}.csv") max_wait = 30 wait_elapsed = 0 while not csv_file.is_file() and wait_elapsed < max_wait: time.sleep(1) wait_elapsed += 1 assert csv_file.is_file(), f"{csv_file} not found!" print(f"download done in {time.time() - start_time:.2f} seconds.")
内容的提问来源于stack exchange,提问作者Ganesh C.
相关产品推荐
相关产品推荐

