SEC对冲基金13F报告爬虫出现IndexError: list index out of range错误求助
IndexError: list index out of range in SEC 13F Web Scraper Problem Description
I'm new to programming, and with a friend's help, I built a web scraper to pull hedge fund 13F reports from the SEC website. It worked fine before, but recently I hit an error at this line:
response_two = get_request(sec_url + tags[0]['href'])
The error is IndexError: list index out of range, and I can't figure out why the index suddenly stopped working. I tried troubleshooting with the SEC site's browser console but didn't get anywhere. Here's my full code:
import requests import re import csv import lxml import numpy as np import pandas as pd from bs4 import BeautifulSoup sec_url = 'https://www.sec.gov' def get_request(url): return requests.get(url) def create_url(cik): return 'https://www.sec.gov/cgi-bin/browse-edgar?CIK={}&owner=exclude&action=getcompany&type=13F-HR'.format(cik) def get_user_input(): cik = input("Enter CIK number:") return cik requested_cik = get_user_input() # Find mutual fund by CIK number on EDGAR response = get_request(create_url(requested_cik)) soup = BeautifulSoup(response.text, "html.parser") tags = soup.findAll('a', id="documentsbutton") # Find latest 13F report for mutual fund response_two = get_request(sec_url + tags[0]['href']) soup_two = BeautifulSoup(response_two.text, "html.parser") tags_two = soup_two.findAll('a', attrs={'href': re.compile('xml')}) xml_url = tags_two[3].get('href') response_xml = get_request(sec_url + xml_url) soup_xml = BeautifulSoup(response_xml.content, "lxml") # DataFrame df = pd.DataFrame() df['companies'] = soup_xml.body.findAll(re.compile('nameofissuer')) df['value'] = soup_xml.body.findAll(re.compile('value')) for row in df.index: df.loc[row, 'value'] = df.loc[row, 'value'].text df.loc[row, 'companies'] = df.loc[row, 'companies'].text df['value'] = df['value'].astype(float) df = df.groupby('companies').sum() df = df.sort_values('value',ascending=False) for row in df.index: df.loc[row, 'allocation'] = df.loc[row, 'value']/df['value'].sum()*100 df['allocation'] = df['allocation'].astype(int) df = df.drop('value', axis=1) df
Solution
Let's break down why this is happening and fix it step by step:
1. Why the Error Happens
The IndexError means tags = soup.findAll('a', id="documentsbutton") is returning an empty list. That can happen for a few key reasons:
- The SEC updated their page structure, so the
documentsbuttonID no longer exists - Your request is being blocked by SEC's anti-scraping measures (no valid request headers)
- The CIK number you entered is invalid, or there are no 13F reports for that entity
2. Fix #1: Add Error Checking for Empty Lists
First, we need to make sure we don't try to access an index that doesn't exist. Add this check right after fetching tags:
tags = soup.findAll('a', id="documentsbutton") if not tags: print(f"Oops! No 13F reports found for CIK {requested_cik}, or the SEC page structure changed.") exit() # Now proceed safely response_two = get_request(sec_url + tags[0]['href'])
Do the same for the XML link later in your code—tags_two[3] has the same risk of index errors:
tags_two = soup_two.findAll('a', attrs={'href': re.compile('xml')}) if len(tags_two) < 4: print("Error: Couldn't locate the 13F XML file link.") exit() xml_url = tags_two[3].get('href')
3. Fix #2: Add Valid Request Headers (Critical for SEC)
The SEC requires scrapers to send a valid User-Agent header with your contact info, otherwise they'll block your requests or return a non-functional page. Update your get_request function:
def get_request(url): headers = { 'User-Agent': 'Your Full Name/1.0 (your.email@example.com)' # Replace with your actual info! } return requests.get(url, headers=headers)
This is probably the most common reason your scraper stopped working—SEC likely tightened their anti-scraping rules since you last used it.
4. Fix #3: Verify Your Request is Working
Add checks to confirm you're getting a valid page response:
response = get_request(create_url(requested_cik)) print(f"Response Status Code: {response.status_code}") # Print the first 500 characters of the page to check if it's what you expect print(response.text[:500]) soup = BeautifulSoup(response.text, "html.parser")
If the status code isn't 200, or the printed text looks like a captcha/error page, your request is being blocked. You may need to add more headers (like Accept-Language) or delay requests to avoid rate limits.
5. Fix #4: Check for Page Structure Changes
If your request returns a valid 200 status but tags is still empty, the SEC changed their page structure. Go to the EDGAR page for your CIK in a browser, right-click the "Documents" button, and inspect the element to find the new selector. For example, if the button now uses a class instead of an ID, update your code:
# If the button's class is "documents-button" instead of ID "documentsbutton" tags = soup.find_all('a', class_='documents-button')
Final Notes
Start with adding the request headers and error checking—those two fixes will resolve most cases of this error. If the page structure changed, adjust your selectors based on what you find in the browser's inspector.
内容的提问来源于stack exchange,提问作者Ryan Reeves

