Python网络爬虫如何提取同类别下多个<a>标签中指定的帖子链接
解决方法
你要提取的帖子链接对应的<a>标签是包裹标题span.title的父级元素,直接通过标题节点反向查找父级a标签即可,不会匹配到另一个无关的a标签。
基础写法(直接补全find后的参数)
link = post.find("span", {"class": "title"}).find_parent("a")["href"]
稳妥写法(增加空值判断,避免节点不存在时报错)
补全后的完整函数如下:
def extract_job(post): title_span = post.find("span", {"class": "title"}) title = title_span.text.strip() if title_span else None company = post.find("span", {"class": "company"}).text.strip() if post.find("span", {"class": "company"}) else None location = post.find("span", {"class": "region"}).text.strip() if post.find("span", {"class": "region"}) else None # 核心逻辑:通过标题span找父级a标签 link_tag = title_span.find_parent("a") if title_span else None link = link_tag.get("href") if link_tag else None return { "title": title, "company": company, "location": location, "link": link }
其他可选方案
如果目标<a>标签有独有的属性(比如自定义class、rel属性等),也可以直接筛选a标签:
- 若目标a带有
class="post-link"的属性,写法如下:
link = post.find("a", {"class": "post-link"}).get("href")
- 也可以用CSS选择器直接匹配包含标题span的a标签(要求BeautifulSoup版本≥4.7.0):
link = post.select_one("a:has(span.title)").get("href")
内容的提问来源于stack exchange,提问作者Seowoo Jang
相关产品推荐
相关产品推荐

