使用BeautifulSoup解析Confluence表格时遇AttributeError问题排查
问题
尝试从Confluence页面读取表格并保存为JSON/CSV文件时,调用Confluence API能正常获取响应,但用BeautifulSoup解析表格时出现错误,错误信息如下:
C:\User>testScript_confulenec None Traceback (most recent call last): File "C:\User\testScript_confulenec.py", line 34, in <module> for response_data in table.find_all('tbody'): AttributeError: 'NoneType' object has no attribute 'find_all'
API返回的页面内容片段(关键部分):
{'results': [{'id': '2533457945', 'type': 'page', ..., 'body': {'view': {'value': '<div class="table-wrap"> <table data-layout="wide" data-local-id="5f6d5d7f-00ee-4788-8999-e16daab2ba6c" class="confluenceTable"><tbody> <tr><td class="confluenceTd"><p>id</p></td>...</tr> ... </tbody></table></div>'}}]}
使用的Python代码:
# This code sample uses the 'requests' library: # http://docs.python-requests.org import requests from requests.auth import HTTPBasicAuth import json from bs4 import BeautifulSoup url = "https://test.net/wiki/rest/api/content?spaceKey=demo&title=checkTable&expand=space,body.view" auth = HTTPBasicAuth("test@gmail.com", "********") headers = { "Accept": "application/json" } response = requests.request( "GET", url, headers=headers, auth=auth ) #Parsing the HTML file soup = BeautifulSoup(response.text, 'html.parser') #selecting the table table = soup.find('table', class_ = 'confluenceTable') print(table) #storing all rows into one variable for response_data in table.find_all('tbody'): rows = response_data.find_all('tr') print(rows)
原因及解决方法
问题核心是你直接把API返回的JSON文本传给了BeautifulSoup,而非提取出嵌套在JSON里的表格HTML内容。Confluence API返回的是JSON格式数据,表格的HTML实际在response.json()['results'][0]['body']['view']['value']路径下,直接解析整个JSON文本自然找不到<table>标签。
修正步骤:
- 先将API响应解析为JSON对象,提取出包含表格的HTML片段。
- 再用BeautifulSoup解析这段HTML,而非原始响应文本。
修正后的代码:
# This code sample uses the 'requests' library: # http://docs.python-requests.org import requests from requests.auth import HTTPBasicAuth import json from bs4 import BeautifulSoup url = "https://test.net/wiki/rest/api/content?spaceKey=demo&title=checkTable&expand=space,body.view" auth = HTTPBasicAuth("test@gmail.com", "********") headers = { "Accept": "application/json" } response = requests.request( "GET", url, headers=headers, auth=auth ) # 解析API响应为JSON对象 response_json = response.json() # 提取表格所在的HTML内容 html_content = response_json['results'][0]['body']['view']['value'] # 解析提取出的HTML soup = BeautifulSoup(html_content, 'html.parser') # 选择目标表格 table = soup.find('table', class_='confluenceTable') print(table) # 遍历表格行(先判断表格是否存在,避免NoneType错误) if table: for tbody in table.find_all('tbody'): rows = tbody.find_all('tr') print(rows) else: print("未找到目标表格")
额外提示:
- 加入
if table:的判断,可避免因表格不存在导致的后续报错。 - 若要转为CSV/JSON,可遍历每行的
<td>标签,提取<p>内的文本,整理成列表或字典后保存。
内容的提问来源于stack exchange,提问作者NiveditaK
相关产品推荐
相关产品推荐

