Python 2.7实现文件中长URL转短URL报错,如何修复?
解决Python 2.7中长URL转短URL的报错问题
让我来帮你搞定这个问题!你想要把myfile.txt里带一堆参数的长URL简化成只到.html的短URL,但你的代码出了几个问题,咱们一步步来梳理。
原myfile.txt内容
26-04-2018 | Publication 2018, 88936 , https://search.publications.com/pgm-2018-88936.html?search=%3fzkt%3dextended%26pst%3dPublication%26vrt%3d%26zkd%3dInTitle%26dpr%3dAll%26spd%3d20180529%26epd%3d20180529%26sdt%3dDatePublication%26pubId%3d%26pnr%3d1%26rpp%3d10&resultInx=0&sorttype=1&sortorder=4 19-04-2018 | Publication 2018, 8168 , https://search.publications.com/pgm-2018-8168.html?search=%3fzkt%3dextended%26pst%3dPublication%26vrt%3d%26zkd%3dInTitle%26dpr%3dAll%26spd%3d20180529%26epd%3d20180529%26sdt%3dDatePublication%26pubId%3d%26pnr%3d1%26rpp%3d10&resultInx=1&sorttype=1&sortorder=4 26-03-2018 | Publication 2018, 611724 , https://search.publications.com/pgm-2018-611724.html?search=%3fzkt%3dextended%26pst%3dPublication%26vrt%3d%26zkd%3dInTitle%26dpr%3dAll%26spd%3d20180529%26epd%3d20180529%26sdt%3dDatePublication%26pubId%3d%26pnr%3d1%26rpp%3d10&resultInx=2&sorttype=1&sortorder=4 01-02-2017 | Publication 2017, 1452026 , https://search.publications.com/pgm-2017-1452026.html?search=%3fzkt%3dextended%26pst%3dPublication%26vrt%3d%26zkd%3dInTitle%26dpr%3dAll%26spd%3d20180529%26epd%3d20180529%26sdt%3dDatePublication%26pubId%3d%26pnr%3d1%26rpp%3d10&resultInx=3&sorttype=1&sortorder=4
你的错误代码及报错
错误代码
import re with open('myfile.txt', 'r+') as myfile: data = myfile.read() url = re.findall(r'[^https.+?]', data) urlshort = re.findall(r'[^https.+html?]', data) for url in data: myfile.write(url.replace(url, urlshort, data)) myfile.close()
报错信息
Traceback (most recent call last): File "/pyscripts/data.py", line 9, in <module> myfile.write(url.replace(url, urlshort, data)) TypeError: an integer is required
问题拆解
咱们来看看代码里的几个关键错误:
- 正则表达式完全偏离需求:你写的
[^https.+?]是匹配不是h、t、t、p、s、.、+、?的单个字符,根本不是在找URL;urlshort的正则也一样错了。 - 循环逻辑错误:
for url in data是遍历文件内容的每一个字符,不是遍历找到的URL,完全跑偏了方向。 - replace方法用错:字符串的
replace语法是str.replace(old, new[, count]),第三个参数是替换次数(整数),你传了data这个字符串,这就是报错的直接原因。 - 文件读写模式问题:用
r+模式读完文件后,指针在文件末尾,直接写会追加内容,而且没清空原文件,导致内容混乱。
修改后的正确代码
import re # 打开文件,读取内容后写入修改后的结果 with open('myfile.txt', 'r+') as myfile: data = myfile.read() # 定义替换函数:把长URL截断到.html的位置 def shorten_url(match_obj): full_url = match_obj.group(0) # 按?分割,取前面的部分就是短URL return full_url.split('?')[0] # 正则匹配所有带参数的长URL,替换成短URL # 正则解释:匹配以https开头,直到空格前的完整长URL(包含.html?和后续参数) modified_data = re.sub(r'https://[^ ]+\.html\?[^ ]+', shorten_url, data) # 将文件指针移到开头,写入新内容,截断多余的旧内容 myfile.seek(0) myfile.write(modified_data) myfile.truncate()
代码说明
- 正则匹配:
r'https://[^ ]+\.html\?[^ ]+'精准匹配所有带参数的长URL:https://匹配URL开头[^ ]+匹配到空格前的内容(因为你的文件里URL和其他内容用空格分隔)\.html\?匹配到.html?的位置,确保只处理带参数的URL
- 替换逻辑:用
re.sub找到所有匹配的长URL,调用shorten_url函数把URL按?分割,取前面的短URL部分。 - 文件操作:读完内容后用
seek(0)把指针移到文件开头,写入修改后的内容,再用truncate()截断原文件中多余的内容,避免残留旧内容。
运行这段代码后,myfile.txt里的长URL就会变成简洁的https://search.publications.com/pgm-2018-88936.html这类短URL啦。
内容的提问来源于stack exchange,提问作者green
相关产品推荐
相关产品推荐

