BeautifulSoup批量抓取多链接时出现IndexError列表索引越界如何解决
你的报错是因为直接通过固定下标读取links列表元素,当页面的社交链接数量少于3个时,访问不存在的下标就会触发索引越界错误。
修正思路
- 先提取所有符合条件的社交链接的href属性,生成独立的链接列表
- 按照你预设的3个链接列的要求,给长度不足的列表自动填充
None,如果链接数量超过3个可根据需要调整长度上限 - 后续如果需要适配不确定的最大链接数量,可以先遍历所有URL统计最大链接数,再动态生成CSV表头、自动补全到对应长度即可
修正后完整代码
from bs4 import BeautifulSoup import requests from csv import writer urls = ['https://url.com/1','https://url.com/2', 'https://url.com/3'] # 预设的社交链接列数量,需要调整列数直接改这里即可自动适配表头 MAX_SOCIAL_LINKS = 3 with open('multi.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) # 动态生成表头,调整链接数量时无需手动修改 header = ['Name', 'Location'] + [f'Link{i+1}' if i>0 else 'Link' for i in range(MAX_SOCIAL_LINKS)] thewriter.writerow(header) for url in urls: my_url = requests.get(url) html = my_url.content soup = BeautifulSoup(html,'html.parser') lists = soup.find_all('div', class_="profile-info-holder") for l in lists: name = l.find('div', class_="profile-name").text.strip() location = l.find('div', class_="profile-location").text.strip() # 提取所有社交链接的href属性 links = l.find_all('a', class_="intercept", href=True) social_links = [a.get('href') for a in links] # 不足指定数量的部分填充None,超过的部分可根据需求选择是否截断 social_links = social_links[:MAX_SOCIAL_LINKS] + [None] * (MAX_SOCIAL_LINKS - len(social_links)) info = [name, location] + social_links thewriter.writerow(info)
内容的提问来源于stack exchange,提问作者alexeidos
相关产品推荐
相关产品推荐

