Python爬虫:循环多次输出列表,仅需最终完整列表的问题
问题:爬虫循环中多次打印列表,仅需最终完整列表
原代码
from urllib.request import urlopen from bs4 import BeautifulSoup url = 'http://wholepeople.com/best-thrift-stores-in-nyc/#:~:text=1%20Beacon%27s%20Closet.%20We%20think%20any%20thrifter%20in,...%2010%20Emma%20Rogue.%20...%20More%20items...%20' page = urlopen(url) soup = BeautifulSoup(page, 'html.parser') h2Tags = soup.find_all('h2') # 原代码缺少列表初始化,会触发NameError storeList = [] for store in h2Tags[:12]: allStores = (store.text) storeList.append(allStores) print(storeList)
当前输出
['Beacon’s Closet'] ['Beacon’s Closet', 'Laced Up'] ['Beacon’s Closet', 'Laced Up', '2ND Street'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop', 'Crossroads Trading'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop', 'Crossroads Trading', 'The Attic'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop', 'Crossroads Trading', 'The Attic', 'Emma Rogue'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop', 'Crossroads Trading', 'The Attic', 'Emma Rogue', 'Flamingo’s Vintage Pound'] ['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop', 'Crossroads Trading', 'The Attic', 'Emma Rogue', 'Flamingo’s Vintage Pound', 'Monk Vintage Clothing']
问题原因
你把print(storeList)写在了for循环内部,每完成一次循环(往列表里添加一个元素)就会执行一次打印,所以会输出12次逐步增长的列表。
解决方法
把print(storeList)移到for循环外面,等所有元素都添加完成后,再打印最终的完整列表。同时补上原代码缺失的storeList = []初始化语句,避免报错。
修改后的代码:
from urllib.request import urlopen from bs4 import BeautifulSoup url = 'http://wholepeople.com/best-thrift-stores-in-nyc/#:~:text=1%20Beacon%27s%20Closet.%20We%20think%20any%20thrifter%20in,...%2010%20Emma%20Rogue.%20...%20More%20items...%20' page = urlopen(url) soup = BeautifulSoup(page, 'html.parser') h2Tags = soup.find_all('h2') storeList = [] for store in h2Tags[:12]: allStores = store.text storeList.append(allStores) # 循环结束后打印完整列表 print(storeList)
运行后只会输出最终的完整列表:
['Beacon’s Closet', 'Laced Up', '2ND Street', 'Buffalo Exchange', 'Tired Thrift', 'L Train Vintage', 'Cure Thrift Shop', 'Crossroads Trading', 'The Attic', 'Emma Rogue', 'Flamingo’s Vintage Pound', 'Monk Vintage Clothing']
内容的提问来源于stack exchange,提问作者arosesthorn
相关产品推荐
相关产品推荐

