如何遍历赛事页面列表批量爬取所有赛事票务信息?
利物浦球票销售信息爬取问题
我正尝试将票务销售信息整理为易读列表或可筛选表格,目前已成功编写脚本获取各赛事的页面链接,但后续针对单页面的爬取仅能拿到第一条结果,期望获取约4组数据,使用find_all也无效。
已实现的赛事链接爬取代码及结果
代码
import requests from bs4 import BeautifulSoup url = "https://www.liverpoolfc.com/tickets/tickets-availability" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") pages = [] for link in soup.find_all("a", class_="ticket-card fixture"): href = link.get("href") if href: pages.append(href) print("Pages:") for page in set(pages): print("- " + page)
运行结果
Pages: - /tickets/tickets-availability/wolverhampton-wanderers-v-liverpool-fc-4-feb-2023-0300pm-245 - /tickets/tickets-availability/liverpool-fc-v-arsenal-8-apr-2023-0300pm-236 - /tickets/tickets-availability/liverpool-fc-v-manchester-united-4-mar-2023-0300pm-235 - /tickets/tickets-availability/liverpool-fc-v-real-madrid-21-feb-2023-0800pm-238 - /tickets/tickets-availability/liverpool-fc-v-tottenham-hotspur-29-apr-2023-0300pm-232 - /tickets/tickets-availability/liverpool-fc-v-nottingham-forest-22-apr-2023-0300pm-234 - /tickets/tickets-availability/liverpool-fc-v-fulham-18-mar-2023-0300pm-237 - /tickets/tickets-availability/newcastle-united-v-liverpool-fc-18-feb-2023-0530pm-246 - /tickets/tickets-availability/liverpool-fc-v-brentford-6-may-2023-0300pm-231 - /tickets/tickets-availability/liverpool-fc-v-aston-villa-20-may-2023-0300pm-230 - /tickets/tickets-availability/liverpool-fc-v-everton-13-feb-2023-0800pm-233 - /tickets/tickets-availability/crystal-palace-v-liverpool-fc-25-feb-2023-0745pm-247
单页面爬取的问题代码及结果
问题代码
import requests from bs4 import BeautifulSoup url = "https://www.liverpoolfc.com/tickets/tickets-availability/liverpool-fc-v-everton-13-feb-2023-0800pm-233" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Find all the elements with the desired class ticket_sales = soup.find_all(class_="accorMenu") # Create a list to store the extracted information sales_list = [] # Check if any ticket sales were found if ticket_sales: # Iterate over each ticket sale for accorMenuList in ticket_sales: # Extract the desired information from the ticket sale saletype = soup.find("span", class_="saletype").text.strip() salename = soup.find("span", class_="salename").text.strip() prereqs = soup.find("span", class_="prereqs").text.strip() status = soup.find("span", class_="status").text.strip() whenavailable = soup.find("span", class_="whenavailable").text.strip() # Store the extracted information in a dictionary sale_info = { "saletype": saletype, "salename": salename, "prereqs": prereqs, "status": status, "whenavailable": whenavailable } # Add the dictionary to the list of sales sales_list.append(sale_info) # Print the list of sales for sale in sales_list: print("Saletype:", sale["saletype"]) print("Salename:", sale["salename"]) print("Prereqs:", sale["prereqs"]) print("Status:", sale["status"]) print("Whenavailable:", sale["whenavailable"]) print("---") else: # If no ticket sales were found, print a message print("No ticket sales found.")
运行结果
Saletype: match ticket - Salename: Hospitality Prereqs: Status: available Whenavailable: Mon 6 Feb 2023, 11:00am ---
问题原因及修复方案
问题原因
循环中始终使用soup.find()方法,该方法会从整个HTML文档中查找第一个匹配的元素,导致每次循环都获取到相同的第一条数据,无法遍历每个accorMenu容器下的对应票务信息。
修复后的代码
import requests from bs4 import BeautifulSoup url = "https://www.liverpoolfc.com/tickets/tickets-availability/liverpool-fc-v-everton-13-feb-2023-0800pm-233" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # 定位所有票务条目容器 ticket_sales = soup.find_all(class_="accorMenu") sales_list = [] if ticket_sales: for menu_item in ticket_sales: # 从当前容器内查找子元素,而非整个文档 saletype = menu_item.find("span", class_="saletype").text.strip() if menu_item.find("span", class_="saletype") else "无" salename = menu_item.find("span", class_="salename").text.strip() if menu_item.find("span", class_="salename") else "无" prereqs = menu_item.find("span", class_="prereqs").text.strip() if menu_item.find("span", class_="prereqs") else "无" status = menu_item.find("span", class_="status").text.strip() if menu_item.find("span", class_="status") else "无" whenavailable = menu_item.find("span", class_="whenavailable").text.strip() if menu_item.find("span", class_="whenavailable") else "无" sale_info = { "票种类型": saletype, "票种名称": salename, "购买要求": prereqs, "状态": status, "开售时间": whenavailable } sales_list.append(sale_info) # 输出整理后的票务信息 for idx, sale in enumerate(sales_list, 1): print(f"第{idx}组票务信息:") print(f"票种类型: {sale['票种类型']}") print(f"票种名称: {sale['票种名称']}") print(f"购买要求: {sale['购买要求']}") print(f"状态: {sale['状态']}") print(f"开售时间: {sale['开售时间']}") print("---") else: print("未找到任何票务销售信息。")
修复说明
- 将
soup.find()替换为menu_item.find(),确保从当前遍历的accorMenu容器内查找子元素,从而获取每组对应的票务数据。 - 增加空值判断,避免因页面元素缺失导致代码报错。
- 将输出字段改为中文,更符合阅读习惯。
内容的提问来源于stack exchange,提问作者Wings
相关产品推荐
相关产品推荐

