Colab中13 F文件爬取无结果问题求助
解决13F文件爬取返回0条结果的问题
核心问题分析
你的代码存在三个关键问题:无效API令牌、查询参数缺失关键过滤条件、结果解析逻辑错误,导致请求成功但无法获取到13F filings数据。
分步解决方案
1. 替换为有效的SEC-API令牌
你当前使用的token=010101是占位符,必须替换为从SEC-API平台注册获取的有效令牌。无有效令牌时,API仅返回空结构或测试数据,无法获取真实13F文件。
2. 修正查询参数,添加13F文件类型过滤
13F文件的官方表单类型为13F-HR(常规报告)或13F-HR/A(修正报告),你的查询未指定该条件,导致搜索范围覆盖所有SEC文件,无法精准定位目标数据。同时调整时间范围的写法以符合API要求,修正后的参数结构:
params = { "query": { "query_string": { "query": "formType:(\"13F-HR\" OR \"13F-HR/A\") AND name:\"Bridgewater Associates, LP\"", "filter": "filedAt:[2021-01-01 TO 2021-03-15]" } }, "from": 0, "size": 10, "sort": [ { "filedAt": { "order": "desc" } } ] }
3. 修正结果解析逻辑
SEC-API的Search API返回的results数组中,每个元素直接对应一条13F filing记录,无需嵌套解析filings['docs']。原代码的解析逻辑错误,导致无法提取到数据,修正后的解析部分:
# Parse the response JSON data data = json.loads(response.content) # Get the filing documents from the response filings = [] if 'results' in data: filings = data['results'] # 直接取results数组作为filings列表 print(f"Number of filings found: {len(filings)}")
4. 可选:添加调试输出定位问题
在解析前打印完整响应数据,能快速确认API返回的结构是否符合预期:
print("Response data:", json.dumps(data, indent=2))
修正后的完整代码
!pip install requests # 原代码未用到pexpect,可移除该安装命令 import json import requests # 替换为你的有效SEC-API令牌 endpoint = "https://api.sec-api.io?token=YOUR_VALID_TOKEN" # Set the search parameters params = { "query": { "query_string": { "query": "formType:(\"13F-HR\" OR \"13F-HR/A\") AND name:\"Bridgewater Associates, LP\"", "filter": "filedAt:[2021-01-01 TO 2021-03-15]" } }, "from": 0, "size": 10, "sort": [ { "filedAt": { "order": "desc" } } ] } # Send the search request to the API endpoint response = requests.post(endpoint, json=params) if response.status_code == 200: print("Request successful!") else: print("Request failed with status code:", response.status_code) exit() # Parse the response JSON data data = json.loads(response.content) # 调试输出,查看响应结构(可选) # print("Response data:", json.dumps(data, indent=2)) # Get the filing documents from the response filings = [] if 'results' in data: filings = data['results'] print(f"Number of filings found: {len(filings)}") # Loop through each filing document and get the holdings data for filing in filings: filing_url = filing['linkToHtml'] filing_date = filing['filedAt'] holdings = filing.get('holdings', []) # 用get避免缺失字段抛出KeyError # Loop through each holding and extract the required information for holding in holdings: name_of_issuer = holding['nameOfIssuer'] title_of_class = holding['titleOfClass'] cusip = holding['cusip'] ticker = holding.get('ticker', 'N/A') # 部分记录可能无ticker字段 cik = holding['cik'] value = holding['value'] shares = holding['shrsOrPrnAmt']['sshPrnamt'] share_type = holding['shrsOrPrnAmt']['sshPrnamtType'] investment_discretion = holding['investmentDiscretion'] # Print the extracted information print(f"Name of Issuer: {name_of_issuer}") print(f"Title of Class: {title_of_class}") print(f"CUSIP: {cusip}") print(f"Ticker: {ticker}") print(f"CIK: {cik}") print(f"Value: {value}") print(f"Shares: {shares}") print(f"Share Type: {share_type}") print(f"Investment Discretion: {investment_discretion}") print(f"Filing Date: {filing_date}") print(f"Filing URL: {filing_url}") print("-------------")
内容的提问来源于stack exchange,提问作者P A
相关产品推荐
相关产品推荐

