Python网络爬虫For循环迭代异常,药房查询字典输出信息重复

问题原因
- 全局定义的列表
B会累加所有药房的抓取结果,导致不同药房的内容混在一起 - 处理单个搜索结果页时,遍历了全部药房名称更新字典,所有名称对应的value都会被覆盖为当前的
B,最终所有药房的返回值完全相同 - 药房名称和对应的搜索结果没有做绑定处理,逻辑错位
修复方案
写法1:边请求边处理(更省内存)
import requests import bs4 g = [] for y in range(2): y = input('enter the name of pharmacies') g.append(y) C = {} for name in g: url = 'https://google.com/search?q=' + name request_result = requests.get(url) soup = bs4.BeautifulSoup(request_result.text, "html.parser") # 每个药房单独初始化结果列表,避免内容混杂 res_list = [] heading_object = soup.find_all('div', {'class': 'AVsepf'}) for info in heading_object: res_list.append(info.getText()) C[name] = res_list
写法2:保留先存所有soup的逻辑
import requests import bs4 g = [] for y in range(2): y = input('enter the name of pharmacies') g.append(y) soupe = [] for text in g: url = 'https://google.com/search?q=' + text request_result = requests.get(url) soup = bs4.BeautifulSoup(request_result.text,"html.parser") soupe.append(soup) C = {} # 用zip绑定药房名称和对应的搜索soup for name, sup in zip(g, soupe): res_list = [] heading_object = sup.find_all('div', {'class': 'AVsepf'}) for info in heading_object: res_list.append(info.getText()) C[name] = res_list
内容的提问来源于stack exchange,提问作者user11644013
相关产品推荐
相关产品推荐

