You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法抓取USDA网站动态加载的TXT文件,Selenium代码报错求助

解决USDA网站TXT数据抓取问题

先处理你的Selenium报错

你遇到的ValueError: Timeout value connect was <object object at ...>错误,核心原因是Chrome浏览器、ChromeDriver、Selenium版本不兼容,再加上旧版headless模式配置和executable_path设置不当导致的。按以下步骤修正:

1. 确保版本匹配

  • 查看Chrome浏览器版本(设置→关于Chrome),下载对应版本的ChromeDriver
  • 升级Selenium到最新稳定版:pip install --upgrade selenium

2. 修正代码配置

新版Selenium(4.6+)支持自动管理ChromeDriver,无需手动指定executable_path;同时新版Chrome的headless模式需要用--headless=new参数。修正后的代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

URL = 'https://mymarketnews.ams.usda.gov/viewReport/2960'

def extract_text_file_links(driver):
    elements = driver.find_elements(By.CSS_SELECTOR, 'a[href$=".txt"]')
    return [element.get_attribute('href') for element in elements]

def main():
    options = webdriver.ChromeOptions()
    # 新版Chrome headless模式配置
    options.add_argument('--headless=new')
    # 禁用不必要的资源加载和弹窗
    options.add_argument('--disable-gpu')
    options.add_argument('--no-sandbox')
    options.add_argument('--disable-dev-shm-usage')
    
    # Selenium 4.6+ 自动管理ChromeDriver,无需指定executable_path
    driver = webdriver.Chrome(options=options)
    # 设置页面加载超时
    driver.set_page_load_timeout(30)

    try:
        driver.get(URL)
        # 兜底等待页面加载,避免动态渲染未完成
        time.sleep(5)
        WebDriverWait(driver, 30).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'a[href$=".txt"]'))
        )
        
        text_file_links = extract_text_file_links(driver)
        # 筛选2017年10月至2020年7月的目标链接
        target_links = [link for link in text_file_links if any(
            f"{year}{month:02d}" in link 
            for year in range(2017, 2021) 
            for month in range(10 if year==2017 else 1, 8 if year==2020 else 13)
        )]
        
        for link in target_links:
            print(link)
        # 后续下载示例:用requests库请求链接并保存文件
        # import requests
        # for link in target_links:
        #     filename = link.split('/')[-1]
        #     with open(filename, 'wb') as f:
        #         f.write(requests.get(link).content)

    finally:
        driver.quit()

if __name__ == "__main__":
    main()

替代方案:不用Selenium,直接爬静态页面

如果页面的TXT链接是静态渲染的(查看页面源码能找到),用requests+BeautifulSoup更高效,完全避免浏览器驱动的问题:

import requests
from bs4 import BeautifulSoup

URL = 'https://mymarketnews.ams.usda.gov/viewReport/2960'
HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

def get_target_txt_links():
    response = requests.get(URL, headers=HEADERS, timeout=30)
    soup = BeautifulSoup(response.text, 'html.parser')
    # 提取所有TXT格式链接
    all_links = [a['href'] for a in soup.find_all('a', href=True) if a['href'].endswith('.txt')]
    # 筛选2017年10月至2020年7月的链接
    target_links = [link for link in all_links if any(
        f"{year}{month:02d}" in link 
        for year in range(2017, 2021) 
        for month in range(10 if year==2017 else 1, 8 if year==2020 else 13)
    )]
    return target_links

if __name__ == "__main__":
    links = get_target_txt_links()
    for link in links:
        print(link)
    # 下载示例
    # for link in links:
    #     filename = link.split('/')[-1]
    #     with open(filename, 'wb') as f:
    #         f.write(requests.get(link, headers=HEADERS).content)

注意事项

  • 若网站有反爬机制,需在请求头中添加合理的User-Agent,必要时加入请求延迟(time.sleep(1))
  • 批量下载文件时避免请求过于频繁,防止IP被封禁

内容的提问来源于stack exchange,提问作者Conor Larkin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 15:20:03