Python:能否用os.path.join()拼接URL?代码TypeError报错求解决
解决你的Python爬虫报错与URL拼接疑问
一、先搞定TypeError报错
你遇到的TypeError: join() argument must be str or bytes, not 'int'原因很明确:
- 你的
page变量是整数类型,而os.path.join()要求所有传入的路径组件必须是字符串或字节。在代码第9行的os.path.join(storing_folder, chapter, page)里,page是int类型,把它转成字符串str(page)就能解决这个问题。 - 另外,你第10行的
imageFile.write()是空调用,应该写入请求到的图片内容res.content,不然会导致文件写入失败。
修正后的这部分代码应该是这样:
imageFile = open(os.path.join(storing_folder, chapter, str(page)), 'wb') imageFile.write(res.content) imageFile.close()
二、关于URL拼接:别用os.path.join()!
os.path.join()是为本地文件系统路径设计的工具,它会根据操作系统自动调整路径分隔符(比如Windows用\,macOS/Linux用/),但URL的分隔符固定是/,用它拼接URL很容易出问题(比如在Windows环境下生成带\的非法URL)。
更优的URL拼接方式推荐这两种:
1. 用f-string直接格式化(最直观)
比如构造章节和页面的URL:
chapter_num = ch_numb_regex.search(chapter).group() page_url = f"{starting_url}/{chapter_num}/{page}" res = requests.get(page_url)
2. 用标准库的urllib.parse.urljoin(更严谨)
这个工具能正确处理URL的相对路径和绝对路径,避免格式错误:
from urllib.parse import urljoin chapter_num = ch_numb_regex.search(chapter).group() base_chapter_url = urljoin(starting_url, chapter_num) page_url = urljoin(base_chapter_url, str(page)) res = requests.get(page_url)
额外小建议
- 可以用
with语句来管理文件写入,这样不用手动调用close(),更安全且代码更简洁:
with open(os.path.join(storing_folder, chapter, str(page)), 'wb') as imageFile: imageFile.write(res.content)
- 循环里的
while True建议加个终止条件,比如当返回的页面找不到图片时就退出,避免出现无限循环的情况。
内容的提问来源于stack exchange,提问作者ulibrau
相关产品推荐
相关产品推荐

