You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过URL而非页面ID获取Confluence页面或子版块内容?

解决方法:从Confluence URL直接定位并获取指定内容

1. 先解析URL提取关键参数

Confluence的URL自带定位内容所需的全部信息,写个简单的解析逻辑就能自动提取空间键、页面ID(或标题)、锚点/评论ID:

from urllib.parse import urlparse, parse_qs

def parse_confluence_url(url):
    parsed = urlparse(url)
    path_parts = parsed.path.strip('/').split('/')
    query_params = parse_qs(parsed.query)
    
    result = {}
    # 处理整页/子版块URL
    if 'pages' in path_parts:
        pages_idx = path_parts.index('pages')
        if pages_idx + 1 < len(path_parts):
            result['page_id'] = path_parts[pages_idx + 1]
        if 'spaces' in path_parts:
            spaces_idx = path_parts.index('spaces')
            if spaces_idx + 1 < len(path_parts):
                result['space_key'] = path_parts[spaces_idx + 1]
        # 提取锚点(子版块标识)
        if parsed.fragment:
            result['anchor'] = parsed.fragment
    # 处理评论URL
    if 'focusedCommentId' in query_params:
        result['comment_id'] = query_params['focusedCommentId'][0]
    
    return result

2. 无需手动找页面ID:通过空间键+标题获取页面信息

如果URL没直接带页面ID,或者你不想用ID,可以调用Confluence的页面列表接口,通过空间键和页面标题查询页面详情:

import requests

def get_page_by_space_and_title(space_key, page_title, auth_token, domain):
    url = f"https://{domain}/api/v2/pages"
    params = {
        'spaceKey': space_key,
        'title': page_title,
        'limit': 1
    }
    headers = {
        'Authorization': f'Bearer {auth_token}',
        'Accept': 'application/json'
    }
    response = requests.get(url, headers=headers, params=params)
    response.raise_for_status()
    pages = response.json().get('results', [])
    return pages[0] if pages else None

调用时,页面标题可以从URL最后一段(把连字符替换成空格)提取,或者直接用解析后的路径部分拼接。

3. 定位页面内的子版块(锚点内容)

拿到页面的HTML内容后,用BeautifulSoup解析并定位锚点对应的元素:

from bs4 import BeautifulSoup

def get_anchor_content(page_html, anchor_id):
    soup = BeautifulSoup(page_html, 'html.parser')
    anchor_element = soup.find(id=anchor_id)
    if not anchor_element:
        return None
    # 提取锚点所在版块的完整内容(到下一个同级标题为止)
    content = []
    current_element = anchor_element
    while current_element and not current_element.name.startswith('h'):
        content.append(str(current_element))
        current_element = current_element.next_sibling
    return '\n'.join(content)

页面的HTML内容可从页面详情接口的body.storage.value字段获取。

4. 直接获取评论内容

如果URL指向评论,提取comment_id后直接调用评论详情接口:

def get_comment(comment_id, auth_token, domain):
    url = f"https://{domain}/api/v2/comments/{comment_id}"
    headers = {
        'Authorization': f'Bearer {auth_token}',
        'Accept': 'application/json'
    }
    response = requests.get(url, headers=headers)
    response.raise_for_status()
    return response.json()

整合逻辑

把上面的步骤串起来,输入URL就能自动获取对应内容:

def get_confluence_content_from_url(url, auth_token, domain):
    parsed_data = parse_confluence_url(url)
    
    # 优先处理评论请求
    if 'comment_id' in parsed_data:
        return get_comment(parsed_data['comment_id'], auth_token, domain)
    
    # 处理页面/子版块请求
    if 'page_id' in parsed_data:
        # 通过页面ID获取页面详情
        page_url = f"https://{domain}/api/v2/pages/{parsed_data['page_id']}?body-format=storage"
        headers = {'Authorization': f'Bearer {auth_token}', 'Accept': 'application/json'}
        page = requests.get(page_url, headers=headers).json()
    else:
        # 通过空间键+标题获取页面详情
        page_title = parsed_data['path_parts'][-1].replace('-', ' ') if len(parsed_data.get('path_parts', [])) > 0 else ''
        page = get_page_by_space_and_title(parsed_data['space_key'], page_title, auth_token, domain)
    
    if not page:
        return "页面未找到"
    
    # 有锚点则返回子版块内容,否则返回整页内容
    if 'anchor' in parsed_data:
        return get_anchor_content(page['body']['storage']['value'], parsed_data['anchor'])
    return page['body']['storage']['value']

内容的提问来源于stack exchange,提问作者user23569219

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 22:33:23