PycURL上传含非ASCII字符路径文件报error 26问题求助
- Windows 10 操作系统
- Python 3.10.4
- PycURL 7.45.1
需要通过HTTP POST请求向指定API接口上传本地文件,但本地文件完整路径包含德语变音符号(Umlauts,例如字符ö,位于系统用户名路径段),示例文件路径为:"C:\\Users\\NameWith_ö\\folder\\another_folder\\more_folders\\file.pdf"。
已知PycURL仅支持纯ASCII字符串或字节串类型传入参数,因此预先实现编码逻辑,将所有传入PycURL的字符串参数编码为UTF-8格式字节串。该操作解决了此前调用curl.setopt(curl.HTTPPOST, files)时抛出的UnicodeEncodeError异常,但路径编码为UTF-8字节串后,请求执行时抛出(error 26, '')错误,无法完成文件上传。
POST请求核心方法
def post(self, request_url, custom_headers : list = None, upload_fields : dict = None, upload_files : dict = None, further_directives : dict = None, proxy : dict = None): """ Posts a HTTP POST request and returns a response object. :param request_url: The URL that request should point to. :type request_url: string :param custom_headers: An optional list of costum headers that should be sent with the request. :type custom_headers: list :param upload_fields: An optional dict of data that should be inputted into targeted upload fields {key = field name: value = input data}. :type upload_fields: dict :param upload_files: An optional dict of files that should be uploaded to the target site {key = field name: value = dict {key = 'filepath': value = path to file (required), key = 'new_filename': value = new name for uploaded file (optional), key = 'content_type': value = content type (optional, default is 'multipart/form-data')}}. :type upload_files: dict :param further_directives: An optional dict of further directives that should be included in the request (for list of all possible directives see PLACEHOLDER). :type further_directives: dict :param proxy: An optional dict of information used to route the request through a proxy {key = proxy port: value = int, key = proxy domain: value = string} :type proxy: dict :returns: Response object. :rtype: Curlified_Response """ self.prepare_new_request() body_buffer = BytesIO() curl = pycurl.Curl() request_url, custom_headers, further_directives, proxy, upload_fields, upload_files = self.encode_strings(request_url, custom_headers, further_directives, proxy, upload_fields, upload_files) curl.setopt(curl.URL, request_url) if proxy: self.inject_kerberos_proxy(proxy['proxy port'], proxy['proxy domain'], curl)# if custom_headers: curl.setopt(curl.HTTPHEADER, custom_headers)# if further_directives: for item in further_directives.items(): self.include_further_directive(item, curl)# if upload_fields: postfields = urlencode(upload_fields)# curl.setopt(curl.POSTFIELDS, postfields) if upload_files: files = [] for item in upload_files.items(): files.append(self.build_file_upload(item, curl))# curl.setopt(curl.HTTPPOST, files) curl.setopt(curl.HEADERFUNCTION, self.read_header_line) curl.setopt(curl.WRITEFUNCTION, body_buffer.write) curl.perform() self.read_body(body_buffer) curl.close()
文件上传表单项构建方法build_file_upload
def build_file_upload(self, item, curl): """ Gets a tuple of information for a file upload, checks the correctnes of the content and returns the properly built input. """ info_list = [] for i in item[1].items(): if i[0] == 'filepath': info_list.append(curl.FORM_FILE) info_list.append(i[1]) elif i[0] == 'new filename': info_list.append(curl.FORM_FILENAME) info_list.append(i[1]) elif i[0] == 'contenttype': info_list.append(curl.FORM_CONTENTTYPE) info_list.append(i[1]) else: raise ArgumentsError("No known file upload key was provided! Please revise.") if not info_list: raise RequiredArgumentsMissing("One or more required parameters for the file to be uploaded are missing! Please revise.") if 'content_type' not in info_list: info_list.append(curl.FORM_CONTENTTYPE) info_list.append(b'multipart/form-data') return (item[0], tuple(info_list))
字符串编码方法encode_strings
def encode_strings(self, request_url, custom_headers, further_directives, proxy, upload_fields = None, upload_files = None): """ Turns all given strings into bytes to be processable by PycURL. """ custom_headers = [] if custom_headers is None else custom_headers upload_fields = {} if upload_fields is None else upload_fields upload_files = {} if upload_files is None else upload_files further_directives = {} if further_directives is None else further_directives proxy = {} if proxy is None else proxy request_url = bytes(request_url, 'UTF-8') for i in range(len(custom_headers)): if isinstance(custom_headers[i], str): custom_headers[i] = bytes(custom_headers[i], 'UTF-8') for k, v in further_directives.items(): if isinstance(v, str): further_directives[k] = bytes(further_directives[k], 'UTF-8') for k, v in proxy.items(): if isinstance(v, str): proxy[k] = bytes(proxy[k], 'UTF-8') for k, v in upload_fields.items(): if isinstance(v, str): upload_fields[k] = bytes(upload_fields[k], 'UTF-8') for k, v in upload_files.items(): if isinstance(v, dict): for key, value in v.items(): if isinstance(value, str): upload_files[k][key] = bytes(upload_files[k][key], 'UTF-8') if len(upload_fields) == 0 and len(upload_files) == 0: return request_url, custom_headers, further_directives, proxy else: return request_url, custom_headers, further_directives, proxy, upload_fields, upload_files
PycURL抛出的error 26对应libcurl的CURLE_READ_ERROR错误,本质是底层libcurl无法读取指定路径的文件。
问题出在文件路径的编码逻辑上:Windows系统原生使用UTF-16(宽字符)编码处理文件路径,当前版本PycURL绑定的libcurl在Windows平台不会自动将传入的UTF-8字节路径转换为系统兼容的宽字符路径,直接传入UTF-8编码的带非ASCII字符(如德语变音ö)的路径时,系统无法定位到对应文件,最终触发读文件错误。
之前将所有字符串统一转UTF-8字节的逻辑,只解决了PycURL接口本身不接受Unicode字符串的参数校验问题,没有解决Windows平台文件路径的编码兼容问题。
优先选择手动读取文件内容传入的方案,绕开libcurl的路径解析逻辑,兼容性最好:
- 修改
encode_strings方法,跳过对upload_files中filepath字段的编码转换,保留原生Python字符串类型,交给Python内置的文件IO处理。Python 3在Windows下会自动调用系统宽字符API打开文件,完美支持带任意Unicode字符的路径。 - 改造
build_file_upload方法的filepath分支,不直接传路径给libcurl,而是手动打开文件读取内容,通过FORM_BUFFER系列参数传入:
import os if i[0] == 'filepath': # 用Python原生方法打开文件,自动适配系统路径编码 with open(i[1], 'rb') as f: file_content = f.read() # 提取文件名用于表单提交 upload_filename = os.path.basename(i[1]).encode('utf-8') info_list.append(curl.FORM_BUFFER) info_list.append(upload_filename) info_list.append(curl.FORM_BUFFERPTR) info_list.append(file_content) info_list.append(curl.FORM_BUFFERLENGTH) info_list.append(len(file_content))
如果不想手动读取文件内容,也可以将文件路径的编码从UTF-8改为Windows系统本地ANSI编码(可通过sys.getfilesystemencoding()获取),但该方案存在兼容缺陷,若路径包含超出当前系统ANSI编码范围的字符仍会失败,不推荐使用。
内容的提问来源于stack exchange,提问作者Felix Rösch

