如何用BeautifulSoup爬取同class名的首个div文本获取漫画章节数
解决BeautifulSoup获取漫画章节编号的问题
原代码的问题
你的CSS选择器语法有误:
- 类名组合时需要用
.连接,而非空格:section.pure-g entryBar应改为section.pure-g.entryBar(表示同时拥有pure-g和entryBar类的section) - 同理,
div.pure-1 md-1-5要改成div.pure-1.md-1-5
修正后的代码实现
要直接获取首个目标div的文本并提取章节编号,可使用select_one(返回第一个匹配元素):
chp_number = None # 精准定位目标div target_div = soup.select_one('section.pure-g.entryBar > div.pure-1.md-1-5') if target_div: div_text = target_div.text.strip() if "Ch" in div_text: # 提取纯章节编号,比如从"Ch 178"中取出数字部分 chp_number = div_text.replace("Ch ", "") print(chp_number)
说明
select_one比select后取索引[0]更直接,避免空列表索引报错- 修正后的选择器准确匹配到你需要的首个div,无需遍历所有元素
- 最后通过字符串处理提取纯章节编号,过滤掉"Ch"前缀
内容的提问来源于stack exchange,提问作者moosepowa
相关产品推荐
相关产品推荐

