迁移OneNote至Azure Blob Storage时遇InvalidUri错误求助
OneNote页面迁移至Azure Blob Storage时遇InvalidUri错误
问题背景
编写Python脚本提取OneNote笔记本所有分区的页面内容并迁移至Azure Blob Storage,前约100页运行正常,但处理第21个分区时突然终止,抛出InvalidUri错误。
错误信息
HttpResponseError Traceback (most recent call last) Cell In[39], line 72 70 print(p_title) 71 try: ---> 72 container_client.upload_blob(name=p_title,data='p_content', overwrite=True) 73 except KeyError as ere: 74 print(ere) File .\tester\Lib\site-packages\azure\core\tracing\decorator.py:78, in distributed_trace.<locals>.decorator.<locals>.wrapper_use_tracer(*args, **kwargs) //tester is my kernel 76 span_impl_type = settings.tracing_implementation() 77 if span_impl_type is None: ---> 78 return func(*args, **kwargs) 80 # Merge span is parameter is set, but only if no explicit parent are passed 81 if merge_span and not passed_in_parent: File: .\tester\Lib\site-packages\azure\storage\blob\_container_client.py:1125, in ContainerClient.upload_blob(self, name, data, blob_type, length, metadata, **kwargs) 1123 timeout = kwargs.pop('timeout', None) 1124 encoding = kwargs.pop('encoding', 'UTF-8') -> 1125 blob.upload_blob( 1126 data, 1127 blob_type=blob_type, 1128 length=length, 1129 metadata=metadata, 1130 timeout=timeout, ... ErrorCode:InvalidUri Content: <?xml version="1.0" encoding="utf-8"?> <Error><Code>InvalidUri</Code><Message>The requested URI does not represent any resource on the server. RequestId:4fcbf4ad-501e-0059-1d3e-c2ae87000000 Time:2024-06-19T11:46:07.8701879Z</Message></Error>
相关代码
request_rate = 250 #specify headers that will enter each graphAPI request _headers = { 'Authorization': 'Bearer ' + curr_access_token, 'Content-Type': 'application/json' } def get_pages_within_section(site_id,section_id): #specify url of pages within a specific section url_pages_within_section = f"https://graph.microsoft.com/v1.0/sites/{site_id}/onenote/sections/{section_id}/pages" response = requests.get(url_pages_within_section, headers=_headers).json() return response['value'] def get_page_content(site_id,page_id): # specify endpoint (url of a specific page from a specific section) url_page = f"https://graph.microsoft.com/v1.0/sites/{site_id}/onenote/pages/{page_id}/content" #call to that page to get its contents for i in range(request_rate): pass page_response = requests.get(url_page,headers=_headers) print(page_response) #if everything went well if page_response.status_code == 200: soup = BeautifulSoup(page_response.text,'html.parser') else: # otherwise retry in the dumbest, easiest way possible time.sleep(300) page_response = requests.get(url_page,headers=_headers) page_response.raise_for_status() if page_response.status_code == 200: soup = BeautifulSoup(page_response.text,'html.parser') else: # if you failed again, just give up ;/ page_title = None page_content = None return page_title,page_content #if this is truthy then go on and give us the good stuff if soup.title.string: page_title = soup.title.string if page_title.endswith('.'): page_title = page_title.split('.')[0] else: #just in case if something STILL goes wrong return the following page_title = None page_content = None page_content = page_response.text return page_title,page_content i = 1 #for each section within the notebook for sec_num, section in enumerate(list_sections['value']): #get pages for that section print(sec_num) section_info = get_pages_within_section(site_id=__site_id,section_id=section['id']) if i == 1: i=0 print(section_info) #for each page within that section for page in section_info: #get page title and its .html content p_title, p_content = get_page_content(site_id=__site_id,page_id=page['id']) #then upload a new blob into our previously established container if p_title: # container_client.upload_blob(name=p_title,data=p_content, overwrite=True, content_settings = cnt_settings) print(p_title) try: container_client.upload_blob(name=p_title,data='p_content', overwrite=True) except KeyError as ere: print(ere) else: print(f'came across an empty one!{p_title}')
问题排查与修复
1. 核心错误1:Blob名称包含非法字符
Azure Blob名称不允许包含\ / : * ? " < > |等字符,也不能以.或/结尾。第21个分区的页面标题大概率包含这些非法字符,导致URI无效。
修复方法:添加标题清洗函数,移除或替换非法字符:
def sanitize_blob_name(name): illegal_chars = '<>:"/\\|?*' for char in illegal_chars: name = name.replace(char, '_') # 移除首尾的.和/,避免触发命名限制 name = name.strip('./') # 处理空标题情况 return name or 'untitled_page'
上传前调用该函数:
sanitized_title = sanitize_blob_name(p_title)
2. 核心错误2:上传数据传参错误
代码中data='p_content'会把字符串'p_content'作为内容上传,而非实际页面内容变量p_content,需去掉引号:
container_client.upload_blob(name=sanitized_title, data=p_content, overwrite=True)
3. 额外优化点
- 完善异常捕获:当前仅捕获
KeyError,需添加Azure SDK特定异常捕获,便于排查问题:from azure.core.exceptions import HttpResponseError try: container_client.upload_blob(name=sanitized_title, data=p_content, overwrite=True) except HttpResponseError as e: print(f"上传失败:页面标题[{p_title}],错误详情:{e}") except Exception as e: print(f"未知错误:页面标题[{p_title}],错误:{e}") - 合理控制请求速率:
for i in range(request_rate): pass无法有效控制请求频率,建议替换为time.sleep()避免触发Graph API速率限制:time.sleep(0.5) # 每次请求后等待0.5秒
内容的提问来源于stack exchange,提问作者Jarosław Jaworski
相关产品推荐
相关产品推荐

