You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取Confluence页面中的表格数据?

从Confluence提取表格数据的Python实现方法

方法一:使用Confluence官方REST API

Confluence提供REST API可直接获取页面内容,你可以通过API拿到页面的HTML格式内容,再解析其中的表格。

步骤:

  1. 准备认证信息:
    • Confluence Cloud:需在Atlassian账号设置中生成API令牌,认证方式为用户名:API令牌做Base64编码。
    • Confluence Server/Data Center:直接使用用户名和密码即可。
  2. 发送GET请求获取页面内容,接口路径为/rest/api/content/{pageId}?expand=body.storage
  3. 用BeautifulSoup解析返回的HTML内容,提取表格数据

示例代码:

import requests
from bs4 import BeautifulSoup
import base64

# 配置信息
base_url = "https://your-confluence-domain.atlassian.net"
page_id = "123456"  # 替换为目标页面ID
username = "your-email@example.com"
api_token = "your-api-token"

# 生成认证头
auth_str = f"{username}:{api_token}"
auth_bytes = auth_str.encode('ascii')
auth_base64 = base64.b64encode(auth_bytes).decode('ascii')
headers = {
    "Authorization": f"Basic {auth_base64}",
    "Accept": "application/json"
}

# 获取页面内容
response = requests.get(f"{base_url}/rest/api/content/{page_id}?expand=body.storage", headers=headers)
response.raise_for_status()
page_data = response.json()
html_content = page_data['body']['storage']['value']

# 解析表格
soup = BeautifulSoup(html_content, 'html.parser')
tables = soup.find_all('table')

# 处理第一个表格为例
for table in tables:
    rows = table.find_all('tr')
    table_data = []
    for row in rows:
        cols = row.find_all(['th', 'td'])
        row_data = [col.get_text(strip=True) for col in cols]
        table_data.append(row_data)
    print(table_data)

方法二:使用第三方库atlassian-python-api

这个库封装了Confluence API,操作更简洁,无需手动处理HTTP请求。

步骤:

  1. 安装库:pip install atlassian-python-api
  2. 初始化Confluence客户端并完成认证
  3. 获取页面内容,提取表格数据

示例代码:

from atlassian import Confluence
from bs4 import BeautifulSoup

# 初始化客户端
confluence = Confluence(
    url="https://your-confluence-domain.atlassian.net",
    username="your-email@example.com",
    password="your-api-token"  # Confluence Server用密码,Cloud用API令牌
)

# 获取页面内容
page_id = "123456"
page = confluence.get_page_by_id(page_id, expand="body.storage")
html_content = page['body']['storage']['value']

# 解析表格
soup = BeautifulSoup(html_content, 'html.parser')
tables = soup.find_all('table')

for table in tables:
    rows = table.find_all('tr')
    table_data = []
    for row in rows:
        cols = row.find_all(['th', 'td'])
        row_data = [col.get_text(strip=True) for col in cols]
        table_data.append(row_data)
    print(table_data)

注意事项

  • 确保账号拥有目标页面的查看权限,否则会返回403错误。
  • 页面ID可从页面URL中获取:比如URL为https://xxx.atlassian.net/wiki/spaces/SPACE/pages/123456/Page+Title,其中123456就是页面ID。
  • 若表格包含合并单元格,需额外处理colspan和rowspan属性,避免数据缺失。

内容的提问来源于stack exchange,提问作者Saravana Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:22:18