从Confluence提取表格至DataFrame遇401,请求核查代码规范性
代码合规性分析与修正建议
你的代码核心逻辑没问题,但存在几个直接导致401未授权错误的关键问题:
1. 环境变量读取逻辑错误
os.environ['OEwfen9FFrerGreer5GRrrdfd']这里的键应该是你在系统中设置的环境变量名称,而不是直接把API Key字符串填进去。比如你把密钥存在名为CONFLUENCE_API_KEY的环境变量里,代码要写成:
api_key = os.environ['CONFLUENCE_API_KEY']
当前写法会读取不到正确的密钥,直接触发权限验证失败。
2. Confluence服务器URL格式错误
server参数只需要填Confluence的根域名(API入口地址),不需要带具体页面路径。正确写法是:
server = "https://confluence.abc.com"
get_page_by_title方法会通过空间名和页面标题定位目标页面,带页面路径的URL会导致API请求地址错误,无法通过权限校验。
3. 非权限相关的代码优化(可选)
- BeautifulSoup标签查找可以更明确,把
row("th")改为row.find_all("th"),row("td")改为row.find_all("td"),避免语法歧义:tables_raw = [[[cell.text for cell in row.find_all("th") + row.find_all("td")] for row in table.find_all("tr")] for table in BeautifulSoup(body, features="lxml").find_all("table")] - 添加异常捕获,能更清晰地定位错误类型:
try: page = confluence.get_page_by_title("TEST", "page 1", expand="body.storage") except Exception as e: print(f"请求错误: {str(e)}")
修正后的完整代码示例
from atlassian import Confluence import os from bs4 import BeautifulSoup import pandas as pd user = "user_name" # 替换为你的环境变量名称 api_key = os.environ['CONFLUENCE_API_KEY'] # 使用Confluence根域名 server = "https://confluence.abc.com" confluence = Confluence(url=server, username=user, password=api_key) try: page = confluence.get_page_by_title("TEST", "page 1", expand="body.storage") body = page["body"]["storage"]["value"] tables_raw = [[[cell.text for cell in row.find_all("th") + row.find_all("td")] for row in table.find_all("tr")] for table in BeautifulSoup(body, features="lxml").find_all("table")] tables_df = [pd.DataFrame(table) for table in tables_raw] for table_df in tables_df: print(table_df) except Exception as e: print(f"错误信息: {str(e)}")
另外请确认两个权限事项:
- 你的账号拥有目标页面(空间TEST下的page 1)的查看权限
- 使用的API Key是Confluence个人设置中生成的官方令牌,而非账号密码
内容的提问来源于stack exchange,提问作者Saravana Kumar
相关产品推荐
相关产品推荐

