Python网页爬取问题:无法从bs4元素提取姓名与学号
解决方法:提取姓名和学号
嘿,我来帮你搞定这个提取问题!你现在拿到的是BeautifulSoup的Tag对象列表,得先把里面的文本内容抽出来,再拆分姓名和学号,咱们一步步来:
步骤1:修正目标标签的获取方式
你代码里的all[0]应该是个小失误,而且find_all返回的是Tag对象列表,其实咱们直接用find就能拿到单个目标p标签(页面里这个样式的p标签应该只有一个):
import requests from bs4 import BeautifulSoup r = requests.get("http://cbcs.fastvturesults.com/student/1sp15me001") c = r.content soup = BeautifulSoup(c, "html.parser") # 获取目标p标签 target_p = soup.find("p", {"class": "card-title page-title mb-0 mt-0"})
步骤2:提取姓名
p标签里的姓名是标签的第一个文本节点,咱们可以用.contents获取子节点,再去掉多余空格:
name = target_p.contents[0].strip() # 输出结果:'Agnello Fernandes A'
步骤3:提取学号
学号在span标签里,先找到这个span,提取文本后再去掉前后的括号:
student_id = target_p.find("span", class_="text-muted").text.strip().strip("()") # 输出结果:'1sp15me001'
完整可运行代码
把这些步骤整合起来的完整代码:
import requests from bs4 import BeautifulSoup r = requests.get("http://cbcs.fastvturesults.com/student/1sp15me001") c = r.content soup = BeautifulSoup(c, "html.parser") target_p = soup.find("p", {"class": "card-title page-title mb-0 mt-0"}) name = target_p.contents[0].strip() student_id = target_p.find("span", class_="text-muted").text.strip().strip("()") print("姓名:", name) print("学号:", student_id)
这样就能精准提取到你想要的内容啦,核心就是把BeautifulSoup的Tag对象里的文本内容提取出来,再做简单的格式化处理~
内容的提问来源于stack exchange,提问作者Sayeed
相关产品推荐
相关产品推荐

