You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Erlang中re:run/2与lists:sublist/3索引基数不一致问题问询

Why does re:run/2 use 0-based indices but lists:sublist/3 uses 1-based in Erlang?

Great question—this inconsistency definitely catches a lot of Erlang developers off guard, even if the quick fix (adding 1 to the start index) is straightforward. Let’s break down why this mismatch exists, rooted in Erlang’s design history and the distinct origins of these two functions:

1. Different module heritages

  • The re module is built on POSIX regular expression standards, which universally use 0-based indexing for match positions. This aligns with almost every other regex implementation (Perl, Python, JavaScript, etc.), so Erlang’s regex engine followed this industry norm to avoid confusing developers coming from other regex-heavy environments.
  • The lists module is part of Erlang’s foundational core library, which draws heavy inspiration from older functional languages like ML and Lisp. These languages traditionally use 1-based indexing for lists, as they frame lists as sequences where the first element is logically "position 1" rather than an offset from a starting point.

2. Semantic differences in what indices represent

  • When re:run/2 returns a tuple like {Start, Length}, the Start value is an offset from the beginning of the string (0 means the very first character). This makes sense for regex engines, which scan byte/character streams as linear sequences starting at an initial offset.
  • lists:sublist/3 takes (List, Start, Length) where Start refers to the position of the first element to include in the sublist. Since Erlang lists are linked lists (not arrays), accessing the Nth element requires traversing N nodes from the head—so starting at 1 aligns with how developers intuitively count elements in a list (head = 1, next = 2, etc.).

3. Backward compatibility constraints

Erlang has been around since the late 1980s, and core modules like lists have maintained their 1-based indexing for decades. Changing this now would break an enormous amount of existing production code. Similarly, the re module (added later in Erlang/OTP R7B in 2001) adopted standard regex indexing to avoid deviating from what most developers already knew about regexes.

It’s a perfect example of how "least surprise" is subjective: if you’re coming from a regex background, 0-based feels natural, but if you’re deeply familiar with Erlang’s list operations, 1-based is what you expect. Once you’re aware of the mismatch, adjusting the index (adding 1 to the re:run start position) becomes second nature.

Here’s your original code example formatted clearly:

172> Line = "8#123abc#".
"8#123abc#"
173> re:run(Line,"^(\\d+)#(.+#$)").
{match,[{0,9},{0,1},{2,7}]}
174> lists:sublist(Line,0,1).
** exception error: no function clause matching lists:nthtail(-1,"8#123abc#") (lists.erl, line 180)
     in function  lists:sublist/3 (lists.erl, line 345)
175> lists:sublist(Line,1,1).
"8"

内容的提问来源于stack exchange,提问作者alexakarpov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:26:52