You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

迁移OneNote至Azure Blob Storage时遇InvalidUri错误求助

OneNote页面迁移至Azure Blob Storage时遇InvalidUri错误

问题背景

编写Python脚本提取OneNote笔记本所有分区的页面内容并迁移至Azure Blob Storage,前约100页运行正常,但处理第21个分区时突然终止,抛出InvalidUri错误。

错误信息

HttpResponseError                         Traceback (most recent call last)
Cell In[39], line 72
     70 print(p_title)
     71 try:
---> 72     container_client.upload_blob(name=p_title,data='p_content', overwrite=True)
     73 except KeyError as ere:
     74     print(ere)
File .\tester\Lib\site-packages\azure\core\tracing\decorator.py:78, in distributed_trace.<locals>.decorator.<locals>.wrapper_use_tracer(*args, **kwargs) //tester is my kernel
 76 span_impl_type = settings.tracing_implementation()
     77 if span_impl_type is None:
---> 78     return func(*args, **kwargs)
     80 # Merge span is parameter is set, but only if no explicit parent are passed
     81 if merge_span and not passed_in_parent:

File: .\tester\Lib\site-packages\azure\storage\blob\_container_client.py:1125, in ContainerClient.upload_blob(self, name, data, blob_type, length, metadata, **kwargs)
1123 timeout = kwargs.pop('timeout', None)
   1124 encoding = kwargs.pop('encoding', 'UTF-8')
-> 1125 blob.upload_blob(
   1126     data,
   1127     blob_type=blob_type,
   1128     length=length,
   1129     metadata=metadata,
   1130     timeout=timeout,

...
ErrorCode:InvalidUri
Content: <?xml version="1.0" encoding="utf-8"?>
<Error><Code>InvalidUri</Code><Message>The requested URI does not represent any resource on the server.
RequestId:4fcbf4ad-501e-0059-1d3e-c2ae87000000
Time:2024-06-19T11:46:07.8701879Z</Message></Error>

相关代码

request_rate = 250

#specify headers that will enter each graphAPI request
_headers = {
    'Authorization': 'Bearer ' + curr_access_token,
    'Content-Type': 'application/json'
}

def get_pages_within_section(site_id,section_id):
    #specify url of pages within a specific section
    url_pages_within_section = f"https://graph.microsoft.com/v1.0/sites/{site_id}/onenote/sections/{section_id}/pages"
    response = requests.get(url_pages_within_section, headers=_headers).json()
    return response['value']

def get_page_content(site_id,page_id):
    # specify endpoint (url of a specific page from a specific section)
    url_page = f"https://graph.microsoft.com/v1.0/sites/{site_id}/onenote/pages/{page_id}/content"
    #call to that page to get its contents
    for i in range(request_rate):
        pass
    page_response = requests.get(url_page,headers=_headers)
    print(page_response)
    #if everything went well
    if page_response.status_code == 200:
        soup = BeautifulSoup(page_response.text,'html.parser')
    else: 
        # otherwise retry in the dumbest, easiest way possible
        time.sleep(300)
        page_response = requests.get(url_page,headers=_headers)
        page_response.raise_for_status()
        if page_response.status_code == 200:
            soup = BeautifulSoup(page_response.text,'html.parser')
        else:
            # if you failed again, just give up ;/
            page_title = None
            page_content = None
            return page_title,page_content

    #if this is truthy then go on and give us the good stuff
    if soup.title.string:
        page_title = soup.title.string
        if page_title.endswith('.'):
            page_title = page_title.split('.')[0]
    else:
        #just in case if something STILL goes wrong return the following
        page_title = None
        page_content = None
    
    page_content = page_response.text
    return page_title,page_content


i = 1

#for each section within the notebook
for sec_num, section in enumerate(list_sections['value']):
    #get pages for that section
    print(sec_num)
    section_info = get_pages_within_section(site_id=__site_id,section_id=section['id'])
    if i == 1:
        i=0
        print(section_info)
    #for each page within that section
    for page in section_info:
        #get page title and its .html content
        p_title, p_content = get_page_content(site_id=__site_id,page_id=page['id'])
        #then upload a new blob into our previously established container
        if p_title:
            # container_client.upload_blob(name=p_title,data=p_content, overwrite=True, content_settings = cnt_settings)
            print(p_title)
            try:
                container_client.upload_blob(name=p_title,data='p_content', overwrite=True)
            except KeyError as ere:
                print(ere)

        else:
            print(f'came across an empty one!{p_title}')

问题排查与修复

1. 核心错误1:Blob名称包含非法字符

Azure Blob名称不允许包含\ / : * ? " < > |等字符,也不能以.或/结尾。第21个分区的页面标题大概率包含这些非法字符,导致URI无效。

修复方法:添加标题清洗函数,移除或替换非法字符:

def sanitize_blob_name(name):
    illegal_chars = '<>:"/\\|?*'
    for char in illegal_chars:
        name = name.replace(char, '_')
    # 移除首尾的.和/,避免触发命名限制
    name = name.strip('./')
    # 处理空标题情况
    return name or 'untitled_page'

上传前调用该函数:

sanitized_title = sanitize_blob_name(p_title)

2. 核心错误2:上传数据传参错误

代码中data='p_content'会把字符串'p_content'作为内容上传,而非实际页面内容变量p_content,需去掉引号:

container_client.upload_blob(name=sanitized_title, data=p_content, overwrite=True)

3. 额外优化点

  • 完善异常捕获:当前仅捕获KeyError,需添加Azure SDK特定异常捕获,便于排查问题:
    from azure.core.exceptions import HttpResponseError
    
    try:
        container_client.upload_blob(name=sanitized_title, data=p_content, overwrite=True)
    except HttpResponseError as e:
        print(f"上传失败:页面标题[{p_title}],错误详情:{e}")
    except Exception as e:
        print(f"未知错误:页面标题[{p_title}],错误:{e}")
    
  • 合理控制请求速率:for i in range(request_rate): pass无法有效控制请求频率,建议替换为time.sleep()避免触发Graph API速率限制:
    time.sleep(0.5)  # 每次请求后等待0.5秒
    

内容的提问来源于stack exchange,提问作者Jarosław Jaworski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 02:17:07