使用BeautifulSoup爬取HN时列表异常及类型错误求助
Hacker News 爬虫问题修复指南
问题1:输出列表长度不稳定
原因
- 页面底部的「More」条目被误收入文章列表,导致
article_texts长度多出1条; - 单独遍历
subtext获取点赞数,未与文章条目一一绑定,数据匹配逻辑混乱; - 手动添加
-22到点赞数列表的操作完全破坏了数据一致性; - 请求未带浏览器头,可能被网站限制,返回不完整的页面内容。
解决办法
- 过滤「More」条目:收集文章时判断锚文本是否为「More」,是则跳过;
- 绑定文章与点赞数:通过文章条目找到对应父节点,再定位
subtext,确保每篇文章的点赞数精准匹配; - 添加请求头模拟浏览器访问,避免被网站限制。
问题2:遍历时报错 TypeError: String indices must be integers
原因
原代码用ordered_scores.index()获取排序索引,但遇到相同点赞数的文章时,index()只会返回第一个匹配项的位置,导致多篇文章抢占同一索引,最终ORDERED_ARTICLES中残留初始的空字符串,遍历时空字符串无法用字典索引访问。
解决办法
直接用Python内置的sorted()函数对文章列表排序,指定排序依据为点赞数,反向排序即可,无需手动维护索引和空列表,逻辑更简洁可靠。
修正后的完整代码
from bs4 import BeautifulSoup import requests import pprint URL = "https://news.ycombinator.com/news" # 添加请求头模拟浏览器访问 HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(URL, headers=HEADERS) yc_webpage = response.text yc_soup = BeautifulSoup(yc_webpage, "html.parser") # 直接定位每一篇文章的容器(tr),确保文章和点赞数一一对应 article_rows = yc_soup.find_all("tr", class_="athing") all_articles = [] for row in article_rows: # 获取文章标题和链接 title_tag = row.find("span", class_="titleline").find("a") article_title = title_tag.get_text(strip=True) article_link = title_tag.get("href") # 获取对应行的点赞数:找到下一个兄弟tr(subtext所在行) subtext_row = row.find_next_sibling("tr") score_tag = subtext_row.find("span", class_="score") # 处理无点赞数的情况(新发布的文章) article_score = int(score_tag.get_text(strip=True).replace(" points", "")) if score_tag else 0 all_articles.append({ "title": article_title, "link": article_link, "score": article_score }) # 按点赞数从高到低排序 ordered_articles = sorted(all_articles, key=lambda x: x["score"], reverse=True) # 输出结果 for idx, article in enumerate(ordered_articles, start=1): print(f"{idx}. {article['title']} ({article['link']})") print(f"{article['score']} Upvotes 🔼\n")
内容的提问来源于stack exchange,提问作者Python
相关产品推荐
相关产品推荐

