如何使用Python-BeautifulSoup抓取页面中的多个href链接?
使用Python BeautifulSoup抓取指定链接的正确方法
你的代码存在两处关键错误:
- 错误使用
findAll('div', {'a class': 'xxx'}),这个写法是查找带有a class属性的div元素,而非div下的目标a标签; make_list是div元素的列表,无法直接调用findAll('href')——href是a标签的属性,不是独立标签。
正确实现方式
方法一:直接定位目标a标签
直接筛选所有符合class的a标签,再提取href属性:
# 查找所有目标a标签 link_items = base_soup.find_all('a', class_='link--muted no--text--decoration result-item') # 提取有效href属性(过滤空值) links = [item.get('href') for item in link_items if item.get('href')] print(links)
方法二:先定位父div再找a标签(更精准)
如果页面中存在其他相同class的a标签,可先定位到包含目标链接的父div,再查找内部的a标签:
# 定位父div元素 target_div = base_soup.find('div', class_='cBox-body cBox-body--eyeCatcher') if target_div: # 在父div内查找目标a标签 link_items = target_div.find_all('a', class_='link--muted no--text--decoration result-item') # 提取href属性 links = [item.get('href') for item in link_items if item.get('href')] print(links)
关键说明
- 使用
class_作为指定class的参数(避免与Python内置class关键字冲突); - 用
tag.get('href')获取属性值,比tag['href']更安全——当标签无href属性时,get()会返回None而不是抛出异常; - 列表推导式可快速生成有效链接列表。
内容的提问来源于stack exchange,提问作者test kullanıcısı
相关产品推荐
相关产品推荐

