使用BeautifulSoup爬取Hacker News时无法获取文章href的问题求助
问题分析与解决方法
你的核心问题是**span.titleline标签本身不含href属性**,这个属性是嵌套在该span内部的<a>标签上的。你之前尝试直接获取a标签但未成功,是查找逻辑有误。
修正后的代码
from bs4 import BeautifulSoup import requests response = requests.get("https://news.ycombinator.com/") yc_web_page = response.text soup = BeautifulSoup(yc_web_page, "html.parser") articles = soup.find_all(name="span", class_="titleline") article_texts = [] article_links = [] for article_tag in articles: # 定位span内部的a标签 a_tag = article_tag.find(name="a") if a_tag: article_text = a_tag.get_text() article_texts.append(article_text) article_link = a_tag.get("href") article_links.append(article_link) # 处理点赞数:部分置顶文章可能没有score标签,需确保列表长度匹配 article_upvotes = [] for score in soup.find_all(name="span", class_="score"): upvote_num = int(score.getText().split()[0]) article_upvotes.append(upvote_num) # 补全无点赞数的文章(赋值为0) while len(article_upvotes) < len(article_texts): article_upvotes.append(0) largest_number = max(article_upvotes) largest_index = article_upvotes.index(largest_number) print(article_texts[largest_index]) print(article_links[largest_index]) print(article_upvotes[largest_index])
关键修正点
- 遍历
span.titleline时,通过article_tag.find(name="a")获取内部的a标签,再从a标签提取href和文章标题 - 增加a标签存在性判断,避免网页结构变动导致报错
- 补充点赞数列表的长度匹配处理,防止因置顶文章无score标签引发索引越界
内容的提问来源于stack exchange,提问作者Thomas Hudson
相关产品推荐
相关产品推荐

