Python BeautifulSoup爬虫重复输出同列表问题求助
网页爬虫重复输出结果列表的问题排查与解决
问题原因分析
- 打印位置错误:代码中
print(modeling_company)写在for循环内部,每次抓取到一条数据并追加到列表后,就会打印整个列表。如果soup.find_all('div', class_='gz-list-card-wrapper')返回了50+个元素,就会连续打印50+次列表(每次列表长度递增),看起来像是重复输出同一列表。 - 网站内容变化:目标网站可能更新了页面结构或返回了重复的公司卡片,导致
cls的长度远大于预期,循环次数随之增加。
解决方案
1. 调整打印位置
将打印完整列表的代码移到循环外部,这样只会在所有数据抓取完成后打印一次结果:
# 把循环内的print(modeling_company)删除,移到循环结束后 print(modeling_company)
2. 检查返回的卡片元素
在循环前打印部分卡片内容,确认是否存在重复:
# 打印前5个卡片的文本内容,排查是否重复 for i, cl in enumerate(cls[:5]): print(f"第{i+1}个卡片:{cl.get_text(strip=True)}")
如果发现网站返回了重复内容,需要调整元素定位方式,比如使用更精确的CSS选择器,或检查是否有分页逻辑导致重复抓取。
3. 避免复用核心变量
代码中复用了response和soup变量,容易造成混淆,建议改用不同变量名:
detail_response = requests.get(link, timeout=50) detail_soup = BeautifulSoup(detail_response.content, 'html.parser')
修改后的完整代码
import requests from bs4 import BeautifulSoup url = 'https://business.narimn.org/list/searchalpha/a' response = requests.get(url, timeout=50) soup = BeautifulSoup(response.content, 'html.parser') cls = soup.find_all('div', class_='gz-list-card-wrapper') print(f"找到{len(cls)}个公司卡片") modeling_company = [] for cl in cls: try: a_tag = cl.find('a') if a_tag: link = a_tag['href'] detail_response = requests.get(link, timeout=50) detail_soup = BeautifulSoup(detail_response.content, 'html.parser') company_name = detail_soup.find('h1', class_='gz-pagetitle').get_text() address = detail_soup.find('li', class_='list-group-item gz-card-address') street_ad = address.find('span', class_='gz-street-address').get_text() if address else '' city_ad = address.find('span', class_='gz-address-city').get_text() if address else '' state_ad = address.find('span', itemprop='addressRegion').get_text() if address else '' zip_code = address.find('span', itemprop='postalCode').get_text() if address else '' p_n = detail_soup.find('li', class_='list-group-item gz-card-phone') phone = p_n.find('span', itemprop='telephone').get_text() if p_n else '' modeling_company.append([company_name, street_ad, city_ad, state_ad, zip_code, phone]) # 可选:打印单个抓取的公司名,跟踪进度 print(f"已抓取:{company_name}") except Exception as e: print(f'抓取出错:{e}') # 循环结束后打印完整结果列表 print("\n所有抓取结果:") print(modeling_company)
内容的提问来源于stack exchange,提问作者hadi hassan
相关产品推荐
相关产品推荐

