使用Beautiful Soup网页爬取时遭遇KeyError:0问题求助
KeyError:0 错误分析与解决(BeautifulSoup爬取黄页网站)
错误原因
- 你使用
soup.find('div', class_='result')仅获取到单个结果标签,而非所有搜索结果的列表。此时listings是BeautifulSoup的Tag对象,不支持索引访问。 - 用
enumerate(listings)遍历Tag时,实际是遍历该标签的子节点,而非你需要的商家结果。此时listings[index]会被解析为尝试获取Tag的属性(比如tag['0']),直接触发KeyError:0。 - 循环中重复用
listings[index]访问元素完全冗余,enumerate已经返回了每个结果的value,直接使用即可。
解决步骤
- 获取所有结果集合:将
find('div', class_='result')改为find_all('div', class_='result'),让listings存储所有商家结果的Tag列表。 - 简化循环访问:循环中直接使用
value代替listings[index],避免错误的索引操作。 - 添加异常处理:针对部分商家可能缺失电话、地址等信息的情况,增加
try-except块防止程序崩溃。
修正后的完整代码
import pandas as pd import requests from bs4 import BeautifulSoup info = pd.DataFrame() # 可选:添加请求头规避反爬 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } html_text = requests.get( 'https://www.yellowpages.com/search?search_terms=Plumbers&geo_location_terms=St+Tammany%2C+LA', headers=headers ).text soup = BeautifulSoup(html_text, 'lxml') # 用find_all获取所有商家结果 listings = soup.find('div', class_='search-results center-ads').find_all('div', class_='result') # 遍历结果,直接使用循环变量value for index, value in enumerate(listings): try: info_block = value.div.div.find('div', class_='info') primary_info = info_block.find('div', class_='info-section info-primary') secondary_info = info_block.find('div', class_='info-section info-secondary') NAME = primary_info.h2.a.text CATEGORY = primary_info.find('div', class_='categories').text PHONE = secondary_info.find('div', class_='phone').text ADDRESS = secondary_info.p.text LINK = primary_info.h2.a.get('href') print(f'Name: {NAME}\nCategory: {CATEGORY}\nPhone {PHONE}\nAddress: {ADDRESS},\nLink: www.yellowpages.com{LINK}\n---') # 将数据存入DataFrame info = pd.concat([info, pd.DataFrame({ 'Name': [NAME], 'Category': [CATEGORY], 'Phone': [PHONE], 'Address': [ADDRESS], 'Link': [f'www.yellowpages.com{LINK}'] })], ignore_index=True) except AttributeError: # 跳过缺失关键信息的条目 print(f'条目{index+1}缺失信息,已跳过\n---') continue # 可选:保存结果到CSV # info.to_csv('plumbers_st_tammany.csv', index=False)
额外提示
- 黄页网站有反爬机制,建议添加
time.sleep(1)控制请求间隔,避免被封禁。 - 若需要爬取多页数据,可分析分页URL规律,循环构造请求链接。
内容的提问来源于stack exchange,提问作者MOSIMOS
相关产品推荐
相关产品推荐

