You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas .any()函数失效:判断日志含指定串返回False

Pandas .any()结合str.contains()返回False的问题排查

问题场景

通过Selenium获取页面性能日志并转换为Pandas DataFrame后,使用str.contains()结合.any()判断是否存在包含特定字符串的name字段,即使日志中明确存在符合条件的条目,代码仍返回False。

原代码片段

获取网络日志代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options
import time
import pandas as pd

chrome_options = Options()
chrome_options.add_argument("--headless")

driver = webdriver.Chrome(r'chromedriver.exe',options=chrome_options)
driver.get("https://website.com")
wait = WebDriverWait(driver, 5)

element = wait.until(EC.presence_of_element_located((By.XPATH, '/html/body/section/div/div[1]/div[2]/div[1]/button[1]')))
element.click()
time.sleep(5)

network_logs = driver.execute_script("return window.performance.getEntries()")

匹配判断代码

df = pd.DataFrame(network_logs)
if df['name'].str.contains('collect?v=2&tid=G-').any():
    print(True)
else:
    print(False)

日志示例条目

[{'activationStart': 0,
'connectEnd': 2846.5999999996275,
...(省略部分字段),
'name': 'https://data.website.com/g/collect?v=2&tid=G-WPY57YJNRN&...',
...}]

问题原因

pandas.Series.str.contains()默认采用正则表达式匹配模式,而你要匹配的字符串collect?v=2&tid=G-中包含正则特殊字符?——它在正则语法里表示"匹配前面的字符0次或1次",这导致实际匹配规则和你预期的字面量匹配完全不符,自然无法命中目标字符串。

解决方案

有两种直接可行的修复方式:

方式1:关闭正则匹配,使用字面量匹配

设置regex=False参数,让str.contains()直接按字符串字面量进行匹配:

df = pd.DataFrame(network_logs)
if df['name'].str.contains('collect?v=2&tid=G-', regex=False).any():
    print(True)
else:
    print(False)

方式2:转义正则特殊字符

使用re.escape()对字符串中的正则特殊字符进行转义,让正则引擎按字面量处理这些字符:

import re
import pandas as pd

df = pd.DataFrame(network_logs)
pattern = re.escape('collect?v=2&tid=G-')
if df['name'].str.contains(pattern).any():
    print(True)
else:
    print(False)

额外优化

如果存在大小写不一致的可能,可以添加case=False参数忽略大小写:

df['name'].str.contains('collect?v=2&tid=G-', regex=False, case=False).any()

额外检查点

  1. 用df['name'].info()确认name列没有空值或非字符串类型数据
  2. 可以先打印df['name'].str.contains('collect?v=2&tid=G-', regex=False)的结果,查看具体哪些行匹配成功

内容的提问来源于stack exchange,提问作者Aymen Eddaoudi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 07:20:38