爬取Musinsa商品链接的Python爬虫无输出,请求排查错误
爬虫代码错误排查与修正
问题代码
import urllib.request from bs4 import BeautifulSoup import time link_list = list() header = {'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36'} url = "https://www.musinsa.com/brands/poloralphlauren" req = urllib.request.Request(url=url, headers=header) sourcecode = urllib.request.urlopen(req) soup = BeautifulSoup(sourcecode, "html.parser") for href in soup.find("div", class_="article_info").find_all("list_info"): link_list = link_list.append(href.find("a")["href"]) time.sleep(0.1) print(link_list)
错误点分析
元素定位错误
find_all("list_info")用法错误:list_info是元素的class属性值,不是HTML标签名,需用find_all(class_="list_info")匹配。- 原页面中
div.article_info并非商品列表的父容器,直接调用soup.find("div", class_="article_info")会返回None,后续调用find_all会触发AttributeError,导致代码中断。
列表操作错误
link_list = link_list.append(...)写法错误:append()方法是原地修改列表,返回值为None,赋值后link_list会变成None,后续循环无法执行。正确写法是直接调用link_list.append(...)。
修正后的代码
import urllib.request from bs4 import BeautifulSoup import urllib.parse import time link_list = list() header = {'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36'} url = "https://www.musinsa.com/brands/poloralphlauren" req = urllib.request.Request(url=url, headers=header) sourcecode = urllib.request.urlopen(req) soup = BeautifulSoup(sourcecode, "html.parser") # 定位实际商品列表项:页面中商品项为li标签,class为li_box for item in soup.find_all("li", class_="li_box"): # 提取商品链接标签 a_tag = item.find("a", class_="img-block") # 避免元素不存在或无href属性报错 if a_tag and "href" in a_tag.attrs: # 拼接完整URL(处理相对路径) full_link = urllib.parse.urljoin(url, a_tag["href"]) link_list.append(full_link) time.sleep(0.1) print(link_list)
补充说明
- 用
urllib.parse.urljoin处理相对路径,确保获取完整可访问的商品URL。 - 增加判断条件避免因页面元素异常导致代码崩溃。
- 网站页面结构可能随更新变化,若后续失效需重新检查HTML结构调整选择器。
内容的提问来源于stack exchange,提问作者Null물을 흘린다
相关产品推荐
相关产品推荐

