You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Beautiful Soup高效实现多页面网页抓取并解决运行卡顿问题?

问题诊断与优化方案

1. 脚本无输出/死循环问题

你的脚本确实陷入了死循环,核心原因是分页逻辑的缩进错误:

  • 你写的count=count+1等翻页代码在while(count<2)循环的外部,循环内部count永远等于1,永远满足count<2的判断条件,会一直重复请求同一个页面,自然不会有输出。
  • 另外你的urls数组目前仅配置了1个目标地址,要爬5个URL需要先把剩下4个地址补充到这个数组里。
    正常情况下抓取5个列表页,耗时通常在几秒到十几秒区间,具体取决于你的网络速度和目标网站的响应速度。

2. j循环优化方案

你当前的j循环是完全冗余的无效逻辑,且固定索引取属性的写法容错性极低,优化点如下:

  • 首先删除for j in info['link']这层循环:你现在的逻辑是每新增一个房源链接,就遍历所有已经存入的历史链接重复执行解析,不仅完全没用,还会指数级增加运行耗时。
  • 你当前并没有请求房源详情页,用列表页的soup根本拿不到单个房源的详细字段,要获取chars-column下的属性,需要请求每个房源的link地址,再解析详情页的soup。
  • 不要用固定下标取value-chars的内容,一旦某个房源缺少某个属性,索引就会错位,正确写法是遍历属性行,把属性名和值对应存储:
# 优化后的详情页字段提取逻辑示例
for tag in soup.findAll('div', attrs={'class':'list-announcement-block'}):
    # 先提取列表页已有字段,原有逻辑不变
    name = tag.find('a', attrs={'itemprop':'name'})
    description = tag.find('div', attrs={'class':'announcement-block__description'})
    link = name['href'] if name else None
    date = tag.find('div', attrs={'class':'announcement-block__date'})
    price = tag.find('meta', attrs={'itemprop':'price'})
    price2 = tag.find('div', attrs={'class':'announcement-block__price _premium'})
    
    info['name'].append(name['content'] if name else 'N/A')
    info['description'].append(description.get_text().strip() if description else 'N/A')
    info['date'].append(date.get_text().strip() if date else 'N/A')
    info['price'].append(price['content'] if price else price2.get_text().strip() if price2 else 'N/A')
    
    detail_link = 'http://www.unegui.mn'+link if link else None
    if not detail_link:
        # 缺省值填充
        for key in ['floor','balcony','garage']: # 所有详情字段补N/A
            info[key].append('N/A')
        info['link'].append('N/A')
        continue
    info['link'].append(detail_link)
    # 请求单个房源详情页
    detail_req = Request(detail_link, headers=get_headers())
    detail_html = urlopen(detail_req, context=ctx).read()
    detail_soup = bsoup(detail_html, 'html.parser')
    # 初始化字段映射字典
    char_map = {}
    for char_col in detail_soup.findAll('ul', attrs={'class':'chars-column'}):
        for item in char_col.find_all('li'):
            key = item.find('span', class_='key-chars').get_text(strip=True) if item.find('span', class_='key-chars') else None
            value = item.find('span', class_='value-chars').get_text(strip=True) if item.find('span', class_='value-chars') else 'N/A'
            if key:
                char_map[key] = value
    # 按需从char_map里取对应字段,不用固定索引,不会错位
    info['floor'].append(char_map.get('楼层', 'N/A'))
    info['balcony'].append(char_map.get('阳台', 'N/A'))
    info['garage'].append(char_map.get('车库', 'N/A'))
    # 其余字段按照网站实际属性名同理取值即可
  • 修正分页逻辑的缩进,把翻页代码放到while循环内部:
for i in urls:
    count=1
    y=i
    while(count<2): # 要爬多页就修改这个判断阈值
        http_request = Request(i, headers=get_headers())
        html_file = urlopen(http_request, context=ctx)
        html_text = html_file.read()
        soup = bsoup(html_text, 'html.parser')
        # 原有列表页解析逻辑不变
        # ...
        # 翻页逻辑放在while内部
        count=count+1
        page = '?page='+str(count)
        i=y+page

内容的提问来源于stack exchange,提问作者WX1505

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 01:06:00