如何使用Beautiful Soup高效实现多页面网页抓取并解决运行卡顿问题?
问题诊断与优化方案
1. 脚本无输出/死循环问题
你的脚本确实陷入了死循环,核心原因是分页逻辑的缩进错误:
- 你写的
count=count+1等翻页代码在while(count<2)循环的外部,循环内部count永远等于1,永远满足count<2的判断条件,会一直重复请求同一个页面,自然不会有输出。 - 另外你的
urls数组目前仅配置了1个目标地址,要爬5个URL需要先把剩下4个地址补充到这个数组里。
正常情况下抓取5个列表页,耗时通常在几秒到十几秒区间,具体取决于你的网络速度和目标网站的响应速度。
2. j循环优化方案
你当前的j循环是完全冗余的无效逻辑,且固定索引取属性的写法容错性极低,优化点如下:
- 首先删除
for j in info['link']这层循环:你现在的逻辑是每新增一个房源链接,就遍历所有已经存入的历史链接重复执行解析,不仅完全没用,还会指数级增加运行耗时。 - 你当前并没有请求房源详情页,用列表页的soup根本拿不到单个房源的详细字段,要获取
chars-column下的属性,需要请求每个房源的link地址,再解析详情页的soup。 - 不要用固定下标取
value-chars的内容,一旦某个房源缺少某个属性,索引就会错位,正确写法是遍历属性行,把属性名和值对应存储:
# 优化后的详情页字段提取逻辑示例 for tag in soup.findAll('div', attrs={'class':'list-announcement-block'}): # 先提取列表页已有字段,原有逻辑不变 name = tag.find('a', attrs={'itemprop':'name'}) description = tag.find('div', attrs={'class':'announcement-block__description'}) link = name['href'] if name else None date = tag.find('div', attrs={'class':'announcement-block__date'}) price = tag.find('meta', attrs={'itemprop':'price'}) price2 = tag.find('div', attrs={'class':'announcement-block__price _premium'}) info['name'].append(name['content'] if name else 'N/A') info['description'].append(description.get_text().strip() if description else 'N/A') info['date'].append(date.get_text().strip() if date else 'N/A') info['price'].append(price['content'] if price else price2.get_text().strip() if price2 else 'N/A') detail_link = 'http://www.unegui.mn'+link if link else None if not detail_link: # 缺省值填充 for key in ['floor','balcony','garage']: # 所有详情字段补N/A info[key].append('N/A') info['link'].append('N/A') continue info['link'].append(detail_link) # 请求单个房源详情页 detail_req = Request(detail_link, headers=get_headers()) detail_html = urlopen(detail_req, context=ctx).read() detail_soup = bsoup(detail_html, 'html.parser') # 初始化字段映射字典 char_map = {} for char_col in detail_soup.findAll('ul', attrs={'class':'chars-column'}): for item in char_col.find_all('li'): key = item.find('span', class_='key-chars').get_text(strip=True) if item.find('span', class_='key-chars') else None value = item.find('span', class_='value-chars').get_text(strip=True) if item.find('span', class_='value-chars') else 'N/A' if key: char_map[key] = value # 按需从char_map里取对应字段,不用固定索引,不会错位 info['floor'].append(char_map.get('楼层', 'N/A')) info['balcony'].append(char_map.get('阳台', 'N/A')) info['garage'].append(char_map.get('车库', 'N/A')) # 其余字段按照网站实际属性名同理取值即可
- 修正分页逻辑的缩进,把翻页代码放到while循环内部:
for i in urls: count=1 y=i while(count<2): # 要爬多页就修改这个判断阈值 http_request = Request(i, headers=get_headers()) html_file = urlopen(http_request, context=ctx) html_text = html_file.read() soup = bsoup(html_text, 'html.parser') # 原有列表页解析逻辑不变 # ... # 翻页逻辑放在while内部 count=count+1 page = '?page='+str(count) i=y+page
内容的提问来源于stack exchange,提问作者WX1505
相关产品推荐
相关产品推荐

