You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中可见列无法访问?KeyError问题求助

问题与解决建议

问题描述

代码功能为从多个爬取的数据集提取并拼接数据,拼接后的数据集可见"Insider Name"列,但访问该列时触发KeyError报错。

运行代码

import pandas as pd
import numpy as np # for numeric python functions
from pylab import * # for easy matplotlib plotting
from bs4 import BeautifulSoup
import requests
url1='http://openinsider.com/screener?s=&o=&pl=&ph=&ll=&lh=&fd=30&fdr=&td=0&tdr=&fdlyl=&fdlyh=&daysago=&xp=1&vl=&vh=&ocl=&och=&sic1=-1&sicl=100&sich=9999&grp=0&nfl=&nfh=&nil=&nih=&nol=&noh=&v2l=&v2h=&oc2l=&oc2h=&sortcol=0&cnt=100&page=1'
df1 = pd.read_html(url1)
table=df1[11]
#the table works - now lets make it look at change owned to find the largest value
#sorting
n = np.quantile(table['Qty'], [0.50])
print("99th percentile: ",n)
q=table.sort_values('Qty', ascending = False)
page = requests.get(url1)
name=q['Ticker'].str.replace('\d+', '')
name1 = (table['Ticker'])
n = name1.count()
#Buyers for the company
All = []
url = 'http://openinsider.com/'
for entry in name1:
  table2 = pd.read_html(url+entry)
  dfn=table2[11]
  All.append(dfn)
All = pd.concat(All)
print(All.columns)#<- my sanity check
print(All['Insider Name'])#<- where the problem lies

报错信息

KeyError: 'Insider Name'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
/usr/local/lib/python3.7/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3361                 return self._engine.get_loc(casted_key)
   3362             except KeyError as err:
-> 3363                 raise KeyError(key) from err
   3364 
   3365         if is_scalar(key) and isna(key) and not self.hasnans:

KeyError: 'Insider Name'

解决建议

  • 排查列名的隐藏字符:print(All.columns)显示的列名可能存在前后空格、换行符或不可见字符,导致字面匹配失败。可以用以下代码查看列名的原始形态:

    for col in All.columns:
        print(repr(col))  # 输出列名的精确字符串表示,包含隐藏字符
    

    若发现多余字符,执行All.columns = All.columns.str.strip()去除前后空格,或用正则替换特殊字符。

  • 统一拼接数据集的列名:循环中每个dfn=table2[11]对应的表格可能存在列名不一致(如大小写差异、拼写变体),导致拼接后列名混乱。可以在拼接前强制统一列名:

    for entry in name1:
        table2 = pd.read_html(url+entry)
        dfn=table2[11]
        # 根据实际列顺序,定义统一的列名列表
        dfn.columns = ['Insider Name', 'Ticker', 'Transaction Type', ...]
        All.append(dfn)
    
  • 强制转换列名为字符串类型:拼接后列名可能不是字符串类型,执行All.columns = All.columns.astype(str)转换后再尝试访问。

  • 临时用位置索引访问:若暂时无法定位列名问题,可通过列的位置索引访问,例如All.iloc[:, 0](假设"Insider Name"是第1列),同时继续排查列名异常。

内容的提问来源于stack exchange,提问作者Ok-Investments

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 03:21:30