如何通过URL而非页面ID获取Confluence页面或子版块内容?
解决方法:从Confluence URL直接定位并获取指定内容
1. 先解析URL提取关键参数
Confluence的URL自带定位内容所需的全部信息,写个简单的解析逻辑就能自动提取空间键、页面ID(或标题)、锚点/评论ID:
from urllib.parse import urlparse, parse_qs def parse_confluence_url(url): parsed = urlparse(url) path_parts = parsed.path.strip('/').split('/') query_params = parse_qs(parsed.query) result = {} # 处理整页/子版块URL if 'pages' in path_parts: pages_idx = path_parts.index('pages') if pages_idx + 1 < len(path_parts): result['page_id'] = path_parts[pages_idx + 1] if 'spaces' in path_parts: spaces_idx = path_parts.index('spaces') if spaces_idx + 1 < len(path_parts): result['space_key'] = path_parts[spaces_idx + 1] # 提取锚点(子版块标识) if parsed.fragment: result['anchor'] = parsed.fragment # 处理评论URL if 'focusedCommentId' in query_params: result['comment_id'] = query_params['focusedCommentId'][0] return result
2. 无需手动找页面ID:通过空间键+标题获取页面信息
如果URL没直接带页面ID,或者你不想用ID,可以调用Confluence的页面列表接口,通过空间键和页面标题查询页面详情:
import requests def get_page_by_space_and_title(space_key, page_title, auth_token, domain): url = f"https://{domain}/api/v2/pages" params = { 'spaceKey': space_key, 'title': page_title, 'limit': 1 } headers = { 'Authorization': f'Bearer {auth_token}', 'Accept': 'application/json' } response = requests.get(url, headers=headers, params=params) response.raise_for_status() pages = response.json().get('results', []) return pages[0] if pages else None
调用时,页面标题可以从URL最后一段(把连字符替换成空格)提取,或者直接用解析后的路径部分拼接。
3. 定位页面内的子版块(锚点内容)
拿到页面的HTML内容后,用BeautifulSoup解析并定位锚点对应的元素:
from bs4 import BeautifulSoup def get_anchor_content(page_html, anchor_id): soup = BeautifulSoup(page_html, 'html.parser') anchor_element = soup.find(id=anchor_id) if not anchor_element: return None # 提取锚点所在版块的完整内容(到下一个同级标题为止) content = [] current_element = anchor_element while current_element and not current_element.name.startswith('h'): content.append(str(current_element)) current_element = current_element.next_sibling return '\n'.join(content)
页面的HTML内容可从页面详情接口的body.storage.value字段获取。
4. 直接获取评论内容
如果URL指向评论,提取comment_id后直接调用评论详情接口:
def get_comment(comment_id, auth_token, domain): url = f"https://{domain}/api/v2/comments/{comment_id}" headers = { 'Authorization': f'Bearer {auth_token}', 'Accept': 'application/json' } response = requests.get(url, headers=headers) response.raise_for_status() return response.json()
整合逻辑
把上面的步骤串起来,输入URL就能自动获取对应内容:
def get_confluence_content_from_url(url, auth_token, domain): parsed_data = parse_confluence_url(url) # 优先处理评论请求 if 'comment_id' in parsed_data: return get_comment(parsed_data['comment_id'], auth_token, domain) # 处理页面/子版块请求 if 'page_id' in parsed_data: # 通过页面ID获取页面详情 page_url = f"https://{domain}/api/v2/pages/{parsed_data['page_id']}?body-format=storage" headers = {'Authorization': f'Bearer {auth_token}', 'Accept': 'application/json'} page = requests.get(page_url, headers=headers).json() else: # 通过空间键+标题获取页面详情 page_title = parsed_data['path_parts'][-1].replace('-', ' ') if len(parsed_data.get('path_parts', [])) > 0 else '' page = get_page_by_space_and_title(parsed_data['space_key'], page_title, auth_token, domain) if not page: return "页面未找到" # 有锚点则返回子版块内容,否则返回整页内容 if 'anchor' in parsed_data: return get_anchor_content(page['body']['storage']['value'], parsed_data['anchor']) return page['body']['storage']['value']
内容的提问来源于stack exchange,提问作者user23569219
相关产品推荐
相关产品推荐

