You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式问题:网页爬取后无法提取目标路径内容

问题分析与解决方案

首先看你遇到的问题:re.findall返回空列表,核心原因是正则表达式里多了一个不必要的空格!

你的正则是:

re.findall("^\.\. (\/img\/gifts\/img.*\.jpg)",x)

注意^\.\.后面的空格——但实际你的x值是../img/gifts/img1.jpg这种格式,../和img之间根本没有空格,这就导致正则完全匹配不到内容,自然返回空列表了。

另外还有个小细节:你用了转义的\/,在Python的原始字符串(r"")里其实不需要转义,写起来更清爽。

修正方案

这里给你两种简单的解决思路:

思路1:修正正则表达式

去掉多余的空格,同时改用原始字符串避免转义麻烦:

import re
from urllib.request import urlopen
from bs4 import BeautifulSoup

html = urlopen("http://pythonscraping.com/pages/page3.html")
soup = BeautifulSoup(html,'lxml')
images = soup.findAll("img", {"src":re.compile("\.\.\/img\/gifts\/img.*\.jpg") })
for image in images:
    x = image['src']
    print(x)
    # 修正后的正则:去掉空格,用原始字符串
    mage = re.findall(r"\.\./(img/gifts/img.*\.jpg)", x)
    print(mage)

运行后就能得到['img/gifts/img1.jpg']这样的结果,提取出了去掉../后的路径。

思路2:直接字符串替换(更简单)

如果只是要去掉开头的../,其实不需要用正则,直接用字符串的replace方法更高效:

import re
from urllib.request import urlopen
from bs4 import BeautifulSoup

html = urlopen("http://pythonscraping.com/pages/page3.html")
soup = BeautifulSoup(html,'lxml')
images = soup.findAll("img", {"src":re.compile("\.\.\/img\/gifts\/img.*\.jpg") })
for image in images:
    x = image['src']
    print(x)
    # 直接替换开头的../
    cleaned_path = x.replace("../", "")
    print(cleaned_path)

这种方法代码更简洁,也能精准达到你想要的效果。

内容的提问来源于stack exchange,提问作者Beginner4258

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:10:11