如何在过程宏中检查字符串是否为合法非关键字变量名?
Great question—handling user input for identifiers in proc macros can be tricky, especially since proc_macro2::Ident has that annoying quirk of letting keywords pass while panicking on invalid formats. Here are the most reliable ways to do this without rolling your own error-prone regex:
1. Use the syn Crate (Simplest & Most Robust)
The syn crate is the go-to for parsing Rust syntax in proc macros, and it includes built-in validation for identifiers that rejects both invalid formats and keywords while returning detailed, user-friendly errors.
Here's how to use it:
use syn::{parse_str, Ident, Error}; fn validate_ident(input: &str) -> Result<Ident, Error> { parse_str(input) } // Example usage in your proc macro: match validate_ident(user_input) { Ok(ident) => { // Use the valid identifier } Err(err) => { // Return a compile error with the friendly message return Err(err.to_compile_error()); } }
When you run this:
- If the input is a keyword like
fn, you'll get an error:expected identifier, found keywordfn`` - If the input is invalid (like
123varorvar!), you'll get context-specific messages likeexpected identifier, found literal123`` - Valid identifiers (including Unicode ones like
caféorλ) pass through without issues.
This is the best option because syn handles all the edge cases of Rust's identifier rules (including Unicode support) out of the box, so you don't have to worry about missing anything.
2. Use rustc_lexer + proc_macro2 (No Heavy Dependencies)
If you want to avoid adding syn as a dependency, you can use rustc_lexer (the same lexer Rust itself uses) to manually check for valid identifier syntax and keywords.
use proc_macro2::{Ident, Span}; use rustc_lexer::{is_keyword, is_ident_start, is_ident_continue}; fn validate_ident(input: &str) -> Result<Ident, &'static str> { // Check for empty string if input.is_empty() { return Err("Identifier cannot be empty"); } // Check if it's a Rust keyword if is_keyword(input) { return Err("Cannot use a Rust keyword as a variable name"); } // Validate identifier syntax character by character let mut chars = input.chars(); let first_char = chars.next().unwrap(); if !is_ident_start(first_char) { return Err("Identifier must start with a letter or underscore"); } for c in chars { if !is_ident_continue(c) { return Err("Identifier can only contain letters, digits, or underscores (or valid Unicode identifier characters)"); } } // If all checks pass, create the Ident Ok(Ident::new(input, Span::call_site())) }
This approach uses Rust's own lexing logic to ensure your checks match exactly what the compiler accepts. The is_ident_start and is_ident_continue functions correctly handle Unicode characters that are allowed in Rust identifiers (like accented letters or mathematical symbols), which regex would struggle to cover accurately.
Why Avoid Regex?
You're right to be wary of regex here—Rust's identifier rules are more complex than just [a-zA-Z_][a-zA-Z0-9_]* because they include Unicode categories (like Letter, Mark, Number(Decimal), etc.). Writing a regex that covers all valid Unicode identifiers is error-prone and hard to maintain, whereas the crates above use the same logic as the Rust compiler itself.
内容的提问来源于stack exchange,提问作者Henry Gomersall

