You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Cheerio爬虫使用Regex Match出现TypeError: Cannot read properties of null问题

问题描述

我需要从以下HTML中提取括号内的内容:

<dl class="ooa-1o0axny ev7e6t84">
  <dd class="ooa-16w655c ev7e6t83">
    <p class="ooa-gmxnzj">Cekcyn (Kujawsko-pomorskie)</p>
  </dd>
  <dd class="ooa-16w655c ev7e6t83">
    <p class="ooa-gmxnzj">Some other text</p>
  </dd>
</dl>

我尝试用正则表达式提取,代码如下:

$(item[x]).find('.ooa-gmxnzj:first').text().match(/(?<=\().*(?=\))/)[0]

这段代码能成功获取到括号内的字符串,但同时会抛出错误:

$(item[x]).find('.ooa-gmxnzj:first').text().match(/(?<=\().*(?=\))/)[0]                                                                                                                    
                                                                    ^

TypeError: Cannot read properties of null (reading '0')

我确认匹配成功时,match返回的是包含目标元素的数组:

[
  'Kujawsko-pomorskie',
  index: 8,
  input: 'Cekcyn (Kujawsko-pomorskie)',
  groups: undefined
]

我使用的是axios和cheerio,想知道为什么会出现这个null属性读取错误,而用slice操作就没有这个问题。

原因与解决办法

错误原因

match方法在找不到匹配内容时会返回null,而不是数组。你的代码直接去读取[0]属性,当遇到不带括号的文本(比如例子里的Some other text)时,就会触发这个错误。

而slice是字符串的原生方法,不管字符串里有没有括号,调用slice都不会返回null,最多返回空字符串或者原字符串的一部分,所以不会报错。

修复方案

你需要先判断match的返回值是否为null,再去读取数组元素:

const targetText = $(item[x]).find('.ooa-gmxnzj:first').text();
const matchRes = targetText.match(/(?<=\().*(?=\))/);
const result = matchRes ? matchRes[0] : ''; // 无匹配时返回空字符串,可自定义默认值

也可以用可选链操作符简化代码,避免直接访问null的属性:

$(item[x]).find('.ooa-gmxnzj:first').text().match(/(?<=\().*(?=\))/)?.[0] || ''

可选链?.会在左侧值为null或undefined时停止执行,避免报错;|| ''用来给无匹配的情况设置默认值。

内容的提问来源于stack exchange,提问作者Tomek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 19:17:26