Python正则表达式问题:网页爬取后无法提取目标路径内容
问题分析与解决方案
首先看你遇到的问题:re.findall返回空列表,核心原因是正则表达式里多了一个不必要的空格!
你的正则是:
re.findall("^\.\. (\/img\/gifts\/img.*\.jpg)",x)
注意^\.\.后面的空格——但实际你的x值是../img/gifts/img1.jpg这种格式,../和img之间根本没有空格,这就导致正则完全匹配不到内容,自然返回空列表了。
另外还有个小细节:你用了转义的\/,在Python的原始字符串(r"")里其实不需要转义,写起来更清爽。
修正方案
这里给你两种简单的解决思路:
思路1:修正正则表达式
去掉多余的空格,同时改用原始字符串避免转义麻烦:
import re from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen("http://pythonscraping.com/pages/page3.html") soup = BeautifulSoup(html,'lxml') images = soup.findAll("img", {"src":re.compile("\.\.\/img\/gifts\/img.*\.jpg") }) for image in images: x = image['src'] print(x) # 修正后的正则:去掉空格,用原始字符串 mage = re.findall(r"\.\./(img/gifts/img.*\.jpg)", x) print(mage)
运行后就能得到['img/gifts/img1.jpg']这样的结果,提取出了去掉../后的路径。
思路2:直接字符串替换(更简单)
如果只是要去掉开头的../,其实不需要用正则,直接用字符串的replace方法更高效:
import re from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen("http://pythonscraping.com/pages/page3.html") soup = BeautifulSoup(html,'lxml') images = soup.findAll("img", {"src":re.compile("\.\.\/img\/gifts\/img.*\.jpg") }) for image in images: x = image['src'] print(x) # 直接替换开头的../ cleaned_path = x.replace("../", "") print(cleaned_path)
这种方法代码更简洁,也能精准达到你想要的效果。
内容的提问来源于stack exchange,提问作者Beginner4258
相关产品推荐
相关产品推荐

