You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SEC对冲基金13F报告爬虫出现IndexError: list index out of range错误求助

IndexError: list index out of range in SEC 13F Web Scraper

Problem Description

I'm new to programming, and with a friend's help, I built a web scraper to pull hedge fund 13F reports from the SEC website. It worked fine before, but recently I hit an error at this line:

response_two = get_request(sec_url + tags[0]['href'])

The error is IndexError: list index out of range, and I can't figure out why the index suddenly stopped working. I tried troubleshooting with the SEC site's browser console but didn't get anywhere. Here's my full code:

import requests
import re
import csv
import lxml
import numpy as np
import pandas as pd
from bs4 import BeautifulSoup

sec_url = 'https://www.sec.gov'
def get_request(url):
    return requests.get(url)
def create_url(cik):
    return 'https://www.sec.gov/cgi-bin/browse-edgar?CIK={}&owner=exclude&action=getcompany&type=13F-HR'.format(cik)
def get_user_input():
    cik = input("Enter CIK number:")
    return cik

requested_cik = get_user_input()
# Find mutual fund by CIK number on EDGAR
response = get_request(create_url(requested_cik))
soup = BeautifulSoup(response.text, "html.parser")
tags = soup.findAll('a', id="documentsbutton")
# Find latest 13F report for mutual fund
response_two = get_request(sec_url + tags[0]['href'])
soup_two = BeautifulSoup(response_two.text, "html.parser")
tags_two = soup_two.findAll('a', attrs={'href': re.compile('xml')})
xml_url = tags_two[3].get('href')
response_xml = get_request(sec_url + xml_url)
soup_xml = BeautifulSoup(response_xml.content, "lxml")

# DataFrame
df = pd.DataFrame()
df['companies'] = soup_xml.body.findAll(re.compile('nameofissuer'))
df['value'] = soup_xml.body.findAll(re.compile('value'))

for row in df.index:
    df.loc[row, 'value'] = df.loc[row, 'value'].text
    df.loc[row, 'companies'] = df.loc[row, 'companies'].text

df['value'] = df['value'].astype(float)
df = df.groupby('companies').sum()
df = df.sort_values('value',ascending=False)

for row in df.index:
    df.loc[row, 'allocation'] = df.loc[row, 'value']/df['value'].sum()*100

df['allocation'] = df['allocation'].astype(int)
df = df.drop('value', axis=1)
df

Solution

Let's break down why this is happening and fix it step by step:

1. Why the Error Happens

The IndexError means tags = soup.findAll('a', id="documentsbutton") is returning an empty list. That can happen for a few key reasons:

  • The SEC updated their page structure, so the documentsbutton ID no longer exists
  • Your request is being blocked by SEC's anti-scraping measures (no valid request headers)
  • The CIK number you entered is invalid, or there are no 13F reports for that entity

2. Fix #1: Add Error Checking for Empty Lists

First, we need to make sure we don't try to access an index that doesn't exist. Add this check right after fetching tags:

tags = soup.findAll('a', id="documentsbutton")
if not tags:
    print(f"Oops! No 13F reports found for CIK {requested_cik}, or the SEC page structure changed.")
    exit()
# Now proceed safely
response_two = get_request(sec_url + tags[0]['href'])

Do the same for the XML link later in your code—tags_two[3] has the same risk of index errors:

tags_two = soup_two.findAll('a', attrs={'href': re.compile('xml')})
if len(tags_two) < 4:
    print("Error: Couldn't locate the 13F XML file link.")
    exit()
xml_url = tags_two[3].get('href')

3. Fix #2: Add Valid Request Headers (Critical for SEC)

The SEC requires scrapers to send a valid User-Agent header with your contact info, otherwise they'll block your requests or return a non-functional page. Update your get_request function:

def get_request(url):
    headers = {
        'User-Agent': 'Your Full Name/1.0 (your.email@example.com)'  # Replace with your actual info!
    }
    return requests.get(url, headers=headers)

This is probably the most common reason your scraper stopped working—SEC likely tightened their anti-scraping rules since you last used it.

4. Fix #3: Verify Your Request is Working

Add checks to confirm you're getting a valid page response:

response = get_request(create_url(requested_cik))
print(f"Response Status Code: {response.status_code}")
# Print the first 500 characters of the page to check if it's what you expect
print(response.text[:500])
soup = BeautifulSoup(response.text, "html.parser")

If the status code isn't 200, or the printed text looks like a captcha/error page, your request is being blocked. You may need to add more headers (like Accept-Language) or delay requests to avoid rate limits.

5. Fix #4: Check for Page Structure Changes

If your request returns a valid 200 status but tags is still empty, the SEC changed their page structure. Go to the EDGAR page for your CIK in a browser, right-click the "Documents" button, and inspect the element to find the new selector. For example, if the button now uses a class instead of an ID, update your code:

# If the button's class is "documents-button" instead of ID "documentsbutton"
tags = soup.find_all('a', class_='documents-button')

Final Notes

Start with adding the request headers and error checking—those two fixes will resolve most cases of this error. If the page structure changed, adjust your selectors based on what you find in the browser's inspector.


内容的提问来源于stack exchange,提问作者Ryan Reeves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:48:49