Python爬虫运行后打印列表仅返回第一个值的问题求助
问题根因
代码仅返回单条数据是缩进错误、变量误用等低级问题叠加导致,具体问题如下:
- 核心逻辑错误:
transform函数中构造商家字典、追加列表的代码没有缩进放在for item in data循环体内。循环遍历所有商家节点时只会反复覆盖title/address/phones三个变量,等整个循环结束才会执行一次列表追加,最终列表只会保留最后遍历到的1条数据。 - 变量命名不统一:
extract函数入参定义为大写URL,内部请求时却使用了未定义的小写url;且同个页面重复发起两次GET请求,第二次请求未携带请求头,极易被站点反爬拦截。 - 语法错误:请求头内的
User-Agent字符串、地址清洗的replace参数被错误换行截断,会直接触发运行报错或请求格式异常。 - 逻辑缺失:未做分页遍历,当前代码仅请求第1页内容,即使修复逻辑错误也无法获取全部分类下的所有商家数据。
- 反爬兼容问题:请求手机号接口时未携带请求头,大概率被拦截返回异常数据。
修正后可运行代码
import requests from bs4 import BeautifulSoup import time my_list = [] # 统一全局请求头,避免重复定义 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36' } def extract(page_url): resp = requests.get(page_url, headers=headers) # 复用第一次请求的响应内容解析,删除冗余重复请求 soup = BeautifulSoup(resp.content, "html.parser") return soup.select("[data-tooltip-phones]") def transform(item_list): for item in item_list: phone_url = "https://yellowpages.com.eg" + item["data-tooltip-phones"] title = item.find_previous(class_="item-title").text.strip() address = item.find_previous(class_="address-text").text.strip().replace('\n', '') # 手机号接口携带请求头,加短延时避免触发频率拦截 phones = requests.get(phone_url, headers=headers).json() time.sleep(1) # 单条数据处理完成后立即追加到列表,逻辑放在循环体内 business = { 'name': title, 'address': address, 'telephone': phones } my_list.append(business) # 分页爬取,当前页无商家条目时自动终止 page_num = 1 while True: current_url = f'https://yellowpages.com.eg/en/category/charcoal/p{page_num}' print(f'正在爬取第{page_num}页') current_page_items = extract(current_url) if not current_page_items: break transform(current_page_items) page_num += 1 # 页面请求间隔2秒,降低被反爬封禁的概率 time.sleep(2) print(f'爬取完成,共获取{len(my_list)}条商家信息') print(my_list)
内容的提问来源于stack exchange,提问作者Fouad Elmahdy
相关产品推荐
相关产品推荐

