关于Elixir推导式中二进制生成器语法的疑问
Great question! Let's break down exactly what's happening here—this is a piece of Elixir syntax that's specific to comprehensions, so it's totally normal to be confused when it behaves differently outside that context.
Why <<c <- " hello world ">> works in comprehensions (but not alone)
First off: this is a binary generator, a special syntax that only exists within Elixir's for comprehensions. It's not a standalone binary expression, which is why you hit a compile error when trying to run it on its own.
Let's unpack its behavior:
- Elixir strings are UTF-8 encoded binaries, so
" hello world "is just a sequence of bytes representing characters. - The
<<c <- binary>>syntax in a comprehension tells Elixir to iterate over the binary (your string) one full UTF-8 character at a time, binding each complete character to the variablec. - Compare this to a regular generator like
c <- " hello world ": that would iterate over the string one byte at a time. This works fine for ASCII characters, but breaks for multi-byte UTF-8 characters (likeéorñ), splitting them into individual bytes instead of keeping them as single characters.
Why standalone <<c <- " hello world ">> throws an error
When you write <<c <- " hello world ">> outside a comprehension, Elixir tries to parse it as a standard binary literal. But binary literals don't support the <- operator—there's no built-in <-/2 function for binary expressions. This is the same reason you can't run x <- [1,2,3] on its own: that syntax is exclusive to comprehensions as a way to define iteration logic.
Is this syntax valid outside comprehensions?
Nope! This binary generator syntax is strictly limited to the generator part of for comprehensions. Anywhere else, Elixir will reject it as invalid syntax, just like you saw in your iex session.
Let's walk through your example code
Your comprehension:
for <<c <- " hello world ">>, c != ?\s, into: "", do: <<c>>
Does three key things:
- Uses the binary generator to loop over each UTF-8 character in
" hello world ". - Filters out any characters that are spaces (
c != ?\s—?\sis Elixir's shorthand for the ASCII space character code). - Collects the remaining characters into a new string (via
into: "") by converting each character back to its binary representation with<<c>>.
The end result is the string "helloworld", with all spaces removed.
As a quick side note: if you used a regular generator here (c <- " hello world "), it would still work for ASCII spaces, but it would fail on multi-byte characters. The binary generator is the safe, correct way to iterate over UTF-8 strings character-by-character in comprehensions.
内容的提问来源于stack exchange,提问作者intentionally-left-nil

