You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests爬取网页后正则无法匹配目标内容的问题求助

问题原因及解决办法

原因

  1. 网页返回的HTML是压缩后的单行文本,没有换行符。如果你的正则表达式依赖换行符(比如用^/$匹配行首行尾、\n分割内容),就会匹配失败。
  2. 正则规则可能未适配HTML压缩后的结构,比如忽略了内容中的空格、标签属性顺序变化,或者未启用合适的匹配模式。

解决办法

方法1:调整正则匹配模式

使用re.DOTALL(让.匹配包括换行在内的所有字符,兼容单行压缩内容),同时优化正则规则,避免依赖换行符。示例:

import requests
import re

data = requests.get('https://ru.runetki3.com/?page=1')
# 示例:匹配所有<a>标签内的文本
pattern = re.compile(r'<a.*?>(.*?)</a>', re.DOTALL)
matches = pattern.findall(data.text)
print(matches)

方法2:用HTML解析库替代正则(更可靠)

正则处理HTML容易出错,推荐用BeautifulSoup解析,无需关心换行问题:

import requests
from bs4 import BeautifulSoup

data = requests.get('https://ru.runetki3.com/?page=1')
soup = BeautifulSoup(data.text, 'html.parser')
# 示例:获取页面所有标题元素(根据实际结构调整)
titles = soup.find_all('h2')
for title in titles:
    print(title.get_text())

方法3:格式化单行HTML(可选)

如果需要可视化查看或简化正则编写,可以先格式化单行HTML:

import requests
from bs4 import BeautifulSoup

data = requests.get('https://ru.runetki3.com/?page=1')
soup = BeautifulSoup(data.text, 'html.parser')
formatted_html = soup.prettify()
print(formatted_html)
# 基于格式化后的HTML编写正则

内容的提问来源于stack exchange,提问作者Nascriptah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 06:30:47