Python爬取速卖通无结果问题排查与解决求助
速卖通爬虫无法获取商品解决方案
问题概述
我是Web scraping新手,尝试编写Python代码爬取速卖通(AliExpress)网站,虽网传其对爬虫友好,但运行代码后始终显示“未找到商品”,无法获取商品URL。已添加超时、异常处理相关代码仍无效果,请求解决方案。
原代码如下:
import requests from bs4 import BeautifulSoup import csv import time base_url = "https://ar.aliexpress.com/?gatewayAdapt=glo2ara" urls = [base_url] products = [] while len(urls) != 0: current_url = urls.pop() try: response = requests.get(current_url, timeout=20) response.raise_for_status() soup = BeautifulSoup(response.content, "html.parser") link_elements = soup.select("a[href]") for link_element in link_elements: url = link_element['href'] if base_url in url: urls.append(url) product = {} product["url"] = current_url # Corrected selector to target the image element image_element = soup.select_one(".product img") if image_element: product["image"] = image_element["src"] else: product["image"] = "Image not found" title_element = soup.select_one(".product_title") if title_element: product["title"] = title_element.text else: product["title"] = "Title not found" # Add other fields as needed (e.g., price) products.append(product) except requests.exceptions.RequestException as e: print(f"An error occurred while processing {current_url}: {e}") time.sleep(0.1) if products: # Print the list of products print(products) # Writing to CSV file csv_file_path = 'products.csv' with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=products[0].keys()) # Writing header writer.writeheader() # Writing rows writer.writerows(products) print(f"Data written to {csv_file_path}") else: print("No products found.")
核心问题与解决步骤
- 缺失请求头被反爬拦截:速卖通会校验请求中的
User-Agent等标识,默认requests库的请求头会被识别为爬虫。需添加模拟浏览器的请求头,比如Chrome的UA。 - 选择器与页面结构不匹配:原代码使用的
.product、.product_title并非速卖通当前页面的商品元素类名,需要根据实际页面结构调整选择器。 - URL筛选逻辑错误:原代码将所有包含base_url的链接都加入爬取队列,但大部分是导航、分类页面,需筛选包含
item.htm的商品详情页URL。 - 动态内容加载限制:速卖通部分商品数据由JavaScript动态渲染,静态
requests请求可能无法获取完整内容,若基础调整后仍无效,可考虑使用Selenium/Playwright模拟浏览器渲染。
修正后的代码
import requests from bs4 import BeautifulSoup import csv import time # 模拟浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Language": "ar-EG,ar;q=0.9,en-US;q=0.8,en;q=0.7" } base_url = "https://ar.aliexpress.com" urls = [base_url] products = [] visited_urls = set() # 避免重复爬取同一页面 while urls: current_url = urls.pop() if current_url in visited_urls: continue visited_urls.add(current_url) try: response = requests.get(current_url, headers=headers, timeout=20) response.raise_for_status() soup = BeautifulSoup(response.content, "html.parser") # 提取所有链接,筛选商品详情页(包含item.htm)和站内其他页面 for link in soup.select("a[href]"): href = link.get("href") if not href: continue # 处理相对链接 full_url = href if href.startswith("http") else f"{base_url}{href}" # 筛选商品页或站内分类/导航页 if base_url in full_url: if "item.htm" in full_url: # 处理商品详情页 product = {"url": full_url} # 访问商品详情页获取信息 try: prod_response = requests.get(full_url, headers=headers, timeout=20) prod_soup = BeautifulSoup(prod_response.content, "html.parser") # 提取商品标题(根据当前页面结构调整) title_elem = prod_soup.select_one("h1.product-title-text") product["title"] = title_elem.get_text(strip=True) if title_elem else "Title not found" # 提取商品图片(根据当前页面结构调整) img_elem = prod_soup.select_one("img.magnifier-image") product["image"] = img_elem.get("src") if img_elem else "Image not found" products.append(product) except Exception as e: print(f"Failed to fetch product {full_url}: {e}") else: # 添加非商品站内页面到队列(避免重复) if full_url not in visited_urls: urls.append(full_url) time.sleep(1) # 延长休眠时间,避免频繁请求被封 except requests.exceptions.RequestException as e: print(f"Error processing {current_url}: {e}") if products: print(products) # 写入CSV csv_file_path = 'products.csv' with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=products[0].keys()) writer.writeheader() writer.writerows(products) print(f"Data saved to {csv_file_path}") else: print("No products found.")
额外说明
- 速卖通的页面结构可能随时变化,若选择器失效,需重新通过浏览器开发者工具查看元素类名或标签结构。
- 频繁请求可能触发反爬机制,建议进一步增加休眠时间,或使用代理IP分散请求来源。
- 若动态渲染内容无法通过静态请求获取,可替换为Selenium模拟浏览器操作,确保页面完全加载后再提取数据。
内容的提问来源于stack exchange,提问作者nadamalki
相关产品推荐
相关产品推荐

