使用BeautifulSoup4与Python爬取href返回None?求代码错误排查
问题排查与修复方案
问题原因
你当前代码里,soup.select('.titleline')选中的是带有.titleline类的<span>元素,而不是包含链接的<a>标签。<span>本身没有href属性,所以调用get('href')会返回None。
修复方法
有两种简单的修复方式:
方式一:修改选择器直接选中<a>标签
把选择器改成.titleline a,直接定位到标题行里的链接元素:
# 原代码行 # links = soup.select('.titleline') # 修改为 links = soup.select('.titleline a')
修改后原函数无需改动,就能正常获取到href值。
方式二:在循环中提取子元素的href
如果需要保留原选择器,在循环里先找到<span>下的<a>标签再取href:
def fit_hn(links): hn = [] for item in links: # 找到当前span下的a标签,再获取href a_tag = item.find('a') href = a_tag.get('href') if a_tag else None hn.append(href) return hn
优化后的完整代码(方式一示例)
import requests from bs4 import BeautifulSoup import pprint res = requests.get('https://news.ycombinator.com/newest') soup = BeautifulSoup(res.text, 'html.parser') links = soup.select('.titleline a') def fit_hn(links): hn = [] for item in links: href = item.get('href') hn.append(href) return hn pprint.pprint(fit_hn(links))
内容的提问来源于stack exchange,提问作者Deshan Sing
相关产品推荐
相关产品推荐

