使用网页抓取工具获取href链接失败问题求助
问题解决:抓取Hacker News链接返回None的原因及修复方法
问题原因
你用soup.select('.titleline')获取的是包含<a>标签的父容器元素(Hacker News页面里是<span class="titleline">),而href属性属于容器内部的<a>标签。直接在父容器上调用.get('href'),因父元素无该属性会返回None;而.getText()能正常获取标题,是因为父容器包含了子元素<a>的文本内容。
修复方案
有两种简单的修复方式:
方案1:直接选中<a>标签(推荐)
修改选择器,直接定位到.titleline下的<a>元素,这样每个links元素本身就是带href的链接标签:
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/news') soup = BeautifulSoup(res.text, 'html.parser') # 直接选中.titleline下的<a>标签 links = soup.select('.titleline a') votes = soup.select('.score') def create_custom_hn(links, votes): hn = [] for item in links: title = item.getText() href = item.get('href') print(href) hn.append({'title': title, 'link': href}) return hn print(create_custom_hn(links, votes))
方案2:在父容器中查找<a>子元素
如果要保留原选择器,需先从.titleline元素里找到子元素<a>,再获取它的href属性:
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/news') soup = BeautifulSoup(res.text, 'html.parser') links = soup.select('.titleline') votes = soup.select('.score') def create_custom_hn(links, votes): hn = [] for idx, item in enumerate(links): title = links[idx].getText() # 先找到子元素<a>再提取href href = links[idx].find('a').get('href') print(href) hn.append({'title': title, 'link': href}) return hn print(create_custom_hn(links, votes))
内容的提问来源于stack exchange,提问作者Mikee1
相关产品推荐
相关产品推荐

