Python爬取Kworb榜单触发IndexError: list index out of range问题求助
解决Kworb榜单爬取的IndexError问题
你遇到的IndexError: list index out of range是因为用stripped_strings提取文本时,部分行的文本列表长度不足5,导致取c[4]时超出范围。下面给你两种解决方案:
方案一:直接用pandas的read_html(推荐)
你写的第一段代码其实已经能正确爬取数据,而且代码更简洁高效,完全可以直接用这个方法:
import pandas as pd df = pd.read_html('https://kworb.net/spotify/country/hk_weekly.html', attrs={'id':'spotifyweekly'})[0] df[['Artist','Song']] = df['Artist and Title'].str.split(' - ', n=1, expand=True) df[['Pos','Artist','Song']].to_excel('yourfile.xlsx', index=False)
方案二:修复BeautifulSoup的写法
如果一定要用BeautifulSoup,不要用stripped_strings,直接定位每个td元素提取文本,同时判断元素数量避免索引越界:
import pandas as pd import requests from bs4 import BeautifulSoup soup = BeautifulSoup(requests.get("https://kworb.net/spotify/country/hk_weekly.html").content, 'html.parser') data = [] for e in soup.select('#spotifyweekly tr:has(td)'): tds = e.find_all('td') # 确保有足够的列再提取数据 if len(tds) >= 5: data.append({ 'Pos': tds[0].get_text(strip=True), 'Artist': tds[2].get_text(strip=True), 'Song': tds[4].get_text(strip=True) }) pd.DataFrame(data).to_excel('yourfile_bs.xlsx', index=False)
错误原因说明
stripped_strings会把行内所有文本节点按顺序提取,包括td里的子元素文本,导致生成的列表长度和你预期的列数不匹配;部分行可能因为结构特殊(比如合并单元格、空列),td数量不足,直接按索引取值就会触发越界错误。直接定位td元素提取文本,能更准确对应表格列,加上长度判断就能避免报错。
内容的提问来源于stack exchange,提问作者Hi Hi try
相关产品推荐
相关产品推荐

