You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python调用curl获取Investing.com指定区间股票历史数据

解决Investing.com历史数据查询参数无效的问题

你的问题核心在于请求参数不匹配网站接口要求,以及请求格式和头信息错误,导致服务器忽略了你的查询参数,返回默认数据。以下是修正方案:

关键问题分析

  1. 参数名称错误:Investing.com的历史数据表单使用的参数名不是from_date/to_date/interval,而是st_date/end_date/interval_sec,还必须包含action: historical_data字段。
  2. 请求格式错误:网站接受的是application/x-www-form-urlencoded格式的表单数据,而非JSON,你设置的Content-Type: application/json会让服务器无法解析参数。
  3. Curl参数构造错误:-A后的User-Agent应该拆分为单独的列表元素(避免空格导致参数识别错误),同时需要添加Referer头确保请求合法。
  4. 反爬机制:部分情况下需要先获取页面Cookie,再提交查询请求,否则服务器会拒绝自定义查询。

修正后的完整代码

from subprocess import run, PIPE
from htmlement import HTMLement
from urllib.parse import urlencode

def get_data(tabel):
    try:
        thead = tabel.find("thead")
        trh = thead.find("tr")
        header = []
        for th in trh:
            tmp = th.find('div').find('button').find('span')
            header.append(tmp.text)
        print(header)

        body = tabel.find("tbody")
        for tr in body.findall("tr"):
            row = []
            cnt = 0
            for td in tr:
                if cnt == 0:
                    tijd = td.find('time')
                    row = [tijd.text]
                else:
                    row.append(td.text)
                cnt += 1
            print(row)

    except:
        print('table ; data ; invalid')
    return

def get_document(val):
    url = f'https://www.investing.com{val}-historical-data'
    print('url:', url)
    
    # 修正参数名,匹配网站表单要求
    params = {
        "st_date": "07/01/2024",
        "end_date": "08/31/2024",
        "interval_sec": "Daily",
        "action": "historical_data"
    }
    
    # 构造正确的Curl参数:先获取页面Cookie,再提交表单
    # 第一步:获取页面Cookie,保存到临时文件
    cookie_jar = "investing_cookies.txt"
    run(['curl.exe', '-s', '-A', 'Chrome/91.0.4472.114', '-c', cookie_jar, url], stdout=PIPE)
    
    # 第二步:用Cookie提交表单查询,使用urlencode编码表单数据
    encoded_params = urlencode(params)
    cnfg = [
        'curl.exe', '-s',
        '-A', 'Chrome/91.0.4472.114',
        '-b', cookie_jar,
        '-H', f'Referer: {url}',
        '-d', encoded_params,
        url
    ]
    print('config:', cnfg, '\n')

    htm_doc = run(cnfg, stdout=PIPE).stdout.decode('utf-8')
    try:
        parser = HTMLement("table", attrs={"class": "freeze-column-w-1 w-full overflow-x-auto text-xs leading-4"})
        parser.feed(htm_doc)
        table = parser.close()
        get_data(table)
    except:
        print('html ; table ; invalid')
    return

if __name__ == '__main__':
    get_document('/equities/aarons')
    exit()

主要修改点说明

  1. 参数名修正:将from_date改为st_date,to_date改为end_date,interval改为interval_sec,新增action: historical_data字段,完全匹配网站表单提交的参数。
  2. 请求格式调整:使用urllib.parse.urlencode编码表单数据,而非JSON序列化,去掉错误的Content-Type: application/json头,让Curl自动设置正确的表单头。
  3. Cookie处理:先请求一次目标页面获取Cookie并保存,提交查询时携带Cookie,绕过网站的基础反爬校验。
  4. Curl参数拆分:将-A和User-Agent拆分为两个独立列表元素,避免空格导致的参数识别错误;添加Referer头,模拟浏览器的正常请求流程。

内容的提问来源于stack exchange,提问作者MarcusAurelius

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 17:35:54