You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping:如何获取BeautifulSoup中兄弟<li>标签的文本?

解决思路与代码调整

在BeautifulSoup里,next_sibling会捕获节点间的空白文本(换行、空格这类),所以你直接拿到的不是目标<li>标签,以下几种方法可以解决:

方法1:用find_next_sibling()精准定位

直接指定查找下一个<li>兄弟节点,代码最简洁:

from bs4 import BeautifulSoup
import requests

results = requests.get('https://www.journalismjobs.com/job-listings')
soup = BeautifulSoup(results.text, 'html.parser')
jobs = soup.find_all('a', class_='job-item')

for job in jobs:
    job_title = job.find('h3', class_='job-item-title').text
    job_company = job.find('div', class_='job-item-company').text
    job_details = job.find('ul', class_='job-item-details')
    job_location = job_details.li.text.strip()
    # 调整这一行
    job_type = job_details.li.find_next_sibling('li').text.strip()
    job_desc = job.find('div', class_='job-item-description').text.strip()

方法2:批量提取所有<li>再取值

既然两个目标元素都在同一个<ul>下,直接获取所有<li>标签,按索引取对应内容:

from bs4 import BeautifulSoup
import requests

results = requests.get('https://www.journalismjobs.com/job-listings')
soup = BeautifulSoup(results.text, 'html.parser')
jobs = soup.find_all('a', class_='job-item')

for job in jobs:
    job_title = job.find('h3', class_='job-item-title').text
    job_company = job.find('div', class_='job-item-company').text
    job_details = job.find('ul', class_='job-item-details')
    # 获取所有li标签
    detail_items = job_details.find_all('li')
    job_location = detail_items[0].text.strip()
    job_type = detail_items[1].text.strip()
    job_desc = job.find('div', class_='job-item-description').text.strip()

方法3:跳过空白节点(不推荐,代码冗余)

如果一定要用next_sibling,需要循环跳过空白文本节点,直到找到<li>:

from bs4 import BeautifulSoup
import requests

results = requests.get('https://www.journalismjobs.com/job-listings')
soup = BeautifulSoup(results.text, 'html.parser')
jobs = soup.find_all('a', class_='job-item')

for job in jobs:
    job_title = job.find('h3', class_='job-item-title').text
    job_company = job.find('div', class_='job-item-company').text
    job_details = job.find('ul', class_='job-item-details')
    job_location = job_details.li.text.strip()
    # 跳过空白节点
    next_node = job_details.li.next_sibling
    while next_node and next_node.name != 'li':
        next_node = next_node.next_sibling
    job_type = next_node.text.strip() if next_node else '无数据'
    job_desc = job.find('div', class_='job-item-description').text.strip()

推荐前两种方法,逻辑清晰且代码简洁,能稳定获取目标内容。

内容的提问来源于stack exchange,提问作者Jean-Paul Azzopardi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 15:39:20