无法抓取USDA网站动态加载的TXT文件,Selenium代码报错求助
解决USDA网站TXT数据抓取问题
先处理你的Selenium报错
你遇到的ValueError: Timeout value connect was <object object at ...>错误,核心原因是Chrome浏览器、ChromeDriver、Selenium版本不兼容,再加上旧版headless模式配置和executable_path设置不当导致的。按以下步骤修正:
1. 确保版本匹配
- 查看Chrome浏览器版本(设置→关于Chrome),下载对应版本的ChromeDriver
- 升级Selenium到最新稳定版:
pip install --upgrade selenium
2. 修正代码配置
新版Selenium(4.6+)支持自动管理ChromeDriver,无需手动指定executable_path;同时新版Chrome的headless模式需要用--headless=new参数。修正后的代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time URL = 'https://mymarketnews.ams.usda.gov/viewReport/2960' def extract_text_file_links(driver): elements = driver.find_elements(By.CSS_SELECTOR, 'a[href$=".txt"]') return [element.get_attribute('href') for element in elements] def main(): options = webdriver.ChromeOptions() # 新版Chrome headless模式配置 options.add_argument('--headless=new') # 禁用不必要的资源加载和弹窗 options.add_argument('--disable-gpu') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') # Selenium 4.6+ 自动管理ChromeDriver,无需指定executable_path driver = webdriver.Chrome(options=options) # 设置页面加载超时 driver.set_page_load_timeout(30) try: driver.get(URL) # 兜底等待页面加载,避免动态渲染未完成 time.sleep(5) WebDriverWait(driver, 30).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'a[href$=".txt"]')) ) text_file_links = extract_text_file_links(driver) # 筛选2017年10月至2020年7月的目标链接 target_links = [link for link in text_file_links if any( f"{year}{month:02d}" in link for year in range(2017, 2021) for month in range(10 if year==2017 else 1, 8 if year==2020 else 13) )] for link in target_links: print(link) # 后续下载示例:用requests库请求链接并保存文件 # import requests # for link in target_links: # filename = link.split('/')[-1] # with open(filename, 'wb') as f: # f.write(requests.get(link).content) finally: driver.quit() if __name__ == "__main__": main()
替代方案:不用Selenium,直接爬静态页面
如果页面的TXT链接是静态渲染的(查看页面源码能找到),用requests+BeautifulSoup更高效,完全避免浏览器驱动的问题:
import requests from bs4 import BeautifulSoup URL = 'https://mymarketnews.ams.usda.gov/viewReport/2960' HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } def get_target_txt_links(): response = requests.get(URL, headers=HEADERS, timeout=30) soup = BeautifulSoup(response.text, 'html.parser') # 提取所有TXT格式链接 all_links = [a['href'] for a in soup.find_all('a', href=True) if a['href'].endswith('.txt')] # 筛选2017年10月至2020年7月的链接 target_links = [link for link in all_links if any( f"{year}{month:02d}" in link for year in range(2017, 2021) for month in range(10 if year==2017 else 1, 8 if year==2020 else 13) )] return target_links if __name__ == "__main__": links = get_target_txt_links() for link in links: print(link) # 下载示例 # for link in links: # filename = link.split('/')[-1] # with open(filename, 'wb') as f: # f.write(requests.get(link, headers=HEADERS).content)
注意事项
- 若网站有反爬机制,需在请求头中添加合理的
User-Agent,必要时加入请求延迟(time.sleep(1)) - 批量下载文件时避免请求过于频繁,防止IP被封禁
内容的提问来源于stack exchange,提问作者Conor Larkin
相关产品推荐
相关产品推荐

