使用Beautiful Soup爬取HTML无法获取网站URL,求问题排查
问题分析与解决:无法提取Hacker News文章URL
你提取文章URL的代码逻辑有误。Hacker News页面中,titleline类对应的是一个<span>标签,而文章的URL是嵌套在这个span内部的<a>标签里的,你现在直接去获取这个span的href属性,自然拿不到结果——因为span标签本身并不具备href属性。
修正后的URL提取代码
把你原来的URL提取部分替换成以下代码:
# 先找到titleline标签(也就是包含标题的span) article_tag = soup.find(class_='titleline') # 在这个span内部找到嵌套的a标签 article_url_tag = article_tag.find('a') # 获取a标签的href属性 print(article_url_tag.get('href'))
完整修正代码
from bs4 import BeautifulSoup import requests response = requests.get("https://news.ycombinator.com/") yc_webpage = response.text soup = BeautifulSoup(yc_webpage, 'html.parser') article_tag = soup.find(class_='titleline') article_text = article_tag.get_text() print(article_text) article_score_tag = soup.find(class_='score') article_score_text = article_score_tag.get_text() print(article_score_text) # 修正后的URL提取逻辑 article_url_tag = article_tag.find('a') print(article_url_tag.get('href'))
这样就能正确提取到文章的URL了。
内容的提问来源于stack exchange,提问作者PolymathCarlos
相关产品推荐
相关产品推荐

