You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup爬虫重复输出同列表问题求助

网页爬虫重复输出结果列表的问题排查与解决

问题原因分析

  1. 打印位置错误:代码中print(modeling_company)写在for循环内部,每次抓取到一条数据并追加到列表后,就会打印整个列表。如果soup.find_all('div', class_='gz-list-card-wrapper')返回了50+个元素,就会连续打印50+次列表(每次列表长度递增),看起来像是重复输出同一列表。
  2. 网站内容变化:目标网站可能更新了页面结构或返回了重复的公司卡片,导致cls的长度远大于预期,循环次数随之增加。

解决方案

1. 调整打印位置

将打印完整列表的代码移到循环外部,这样只会在所有数据抓取完成后打印一次结果:

# 把循环内的print(modeling_company)删除,移到循环结束后
print(modeling_company)

2. 检查返回的卡片元素

在循环前打印部分卡片内容,确认是否存在重复:

# 打印前5个卡片的文本内容,排查是否重复
for i, cl in enumerate(cls[:5]):
    print(f"第{i+1}个卡片:{cl.get_text(strip=True)}")

如果发现网站返回了重复内容,需要调整元素定位方式,比如使用更精确的CSS选择器,或检查是否有分页逻辑导致重复抓取。

3. 避免复用核心变量

代码中复用了response和soup变量,容易造成混淆,建议改用不同变量名:

detail_response = requests.get(link, timeout=50)
detail_soup = BeautifulSoup(detail_response.content, 'html.parser')

修改后的完整代码

import requests
from bs4 import BeautifulSoup

url = 'https://business.narimn.org/list/searchalpha/a'
response = requests.get(url, timeout=50)
soup = BeautifulSoup(response.content, 'html.parser')

cls = soup.find_all('div', class_='gz-list-card-wrapper')
print(f"找到{len(cls)}个公司卡片")
modeling_company = []

for cl in cls:
    try:
        a_tag = cl.find('a')
        if a_tag:
            link = a_tag['href']
            detail_response = requests.get(link, timeout=50)
            detail_soup = BeautifulSoup(detail_response.content, 'html.parser')
            
            company_name = detail_soup.find('h1', class_='gz-pagetitle').get_text()
            address = detail_soup.find('li', class_='list-group-item gz-card-address')
            
            street_ad = address.find('span', class_='gz-street-address').get_text() if address else ''
            city_ad = address.find('span', class_='gz-address-city').get_text() if address else ''
            state_ad = address.find('span', itemprop='addressRegion').get_text() if address else ''
            zip_code = address.find('span', itemprop='postalCode').get_text() if address else ''
            
            p_n = detail_soup.find('li', class_='list-group-item gz-card-phone')
            phone = p_n.find('span', itemprop='telephone').get_text() if p_n else ''
            
            modeling_company.append([company_name, street_ad, city_ad, state_ad, zip_code, phone])
            # 可选:打印单个抓取的公司名,跟踪进度
            print(f"已抓取:{company_name}")
    except Exception as e:
        print(f'抓取出错:{e}')

# 循环结束后打印完整结果列表
print("\n所有抓取结果:")
print(modeling_company)

内容的提问来源于stack exchange,提问作者hadi hassan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:40:24