如何使用REST API和Python导出Confluence知识库文章?
使用Confluence REST API导出知识库文章的实现方法
前提条件
- 拥有Confluence实例的API访问权限(推荐生成个人访问令牌(PAT),比用户名密码更安全)
- 目标知识库文章/所属空间的查看及导出权限
1. 导出单篇知识库文章
Confluence REST API提供直接导出单篇内容的端点,支持HTML、PDF、XML等格式:
请求端点
GET /rest/api/content/{id}/export/{format}
参数说明:
{id}:文章的内容ID(可从文章URL末尾的数字串获取){format}:导出格式,可选值为html/pdf/xml
示例代码(Python)
import requests # 替换为你的Confluence实例信息 confluence_base_url = "https://your-confluence-domain.com" content_id = "123456" # 目标文章ID export_format = "pdf" personal_access_token = "your-generated-pat" # 构建请求 url = f"{confluence_base_url}/rest/api/content/{content_id}/export/{format}" headers = {"Authorization": f"Bearer {personal_access_token}"} response = requests.get(url, headers=headers, stream=True) # 保存导出文件 if response.status_code == 200: filename = f"confluence_article_{content_id}.{export_format}" with open(filename, "wb") as f: for chunk in response.iter_content(chunk_size=1024): f.write(chunk) print(f"导出成功:{filename}") else: print(f"导出失败,状态码{response.status_code}:{response.text}")
2. 批量导出整个知识库空间
如果你的知识库对应Confluence的一个空间,可使用异步导出端点批量导出所有内容:
步骤说明
- 发起异步导出请求,获取任务ID
- 轮询任务状态,直到导出完成
- 下载生成的导出文件
请求端点
- 发起导出:
POST /rest/api/space/{spaceKey}/export/async - 查询任务状态:
GET /rest/api/space/{spaceKey}/export/async/{taskId}
示例代码(Python)
import requests import time # 替换为你的配置信息 confluence_base_url = "https://your-confluence-domain.com" space_key = "KB" # 知识库空间的Key(可从空间URL获取) export_format = "pdf" personal_access_token = "your-generated-pat" headers = { "Authorization": f"Bearer {personal_access_token}", "Content-Type": "application/json" } # 1. 发起导出任务 init_export_url = f"{confluence_base_url}/rest/api/space/{space_key}/export/async" payload = { "exportType": export_format, "scope": "all", # 可选:all/pages/blogs "includeAttachments": True, "pageOrientation": "portrait" } init_response = requests.post(init_export_url, headers=headers, json=payload) if init_response.status_code != 202: print(f"发起任务失败:{init_response.text}") exit() task_id = init_response.json()["id"] print(f"导出任务已启动,ID:{task_id}") # 2. 轮询任务状态 status_url = f"{confluence_base_url}/rest/api/space/{space_key}/export/async/{task_id}" while True: status_response = requests.get(status_url, headers=headers) status_data = status_response.json() if status_data["status"] == "COMPLETE": download_url = status_data["result"]["url"] break elif status_data["status"] == "FAILED": print(f"导出失败:{status_data['message']}") exit() print(f"导出进行中,当前状态:{status_data['status']},5秒后重试...") time.sleep(5) # 3. 下载导出文件 download_response = requests.get(download_url, headers=headers, stream=True) if download_response.status_code == 200: filename = f"confluence_space_{space_key}_export.{export_format}" with open(filename, "wb") as f: for chunk in download_response.iter_content(chunk_size=1024): f.write(chunk) print(f"批量导出成功:{filename}") else: print(f"下载失败:{download_response.text}")
关键注意事项
- 导出PDF时,会使用Confluence空间配置的PDF模板,如需自定义样式需提前在后台设置
- 异步导出有超时限制,超大空间建议拆分导出或联系管理员调整系统配置
- 个人访问令牌需在Confluence个人设置中生成,权限需包含
read和export相关权限
内容的提问来源于stack exchange,提问作者Kompal
相关产品推荐
相关产品推荐

