如何用pyparsing解析支持多行续行的键属性定义?
Hey there! Let's break down how to parse this key-value syntax with pyparsing—since you're new to the library, I'll walk you through each step clearly so you understand how it all fits together.
Your input looks like this (with optional line continuations using \):
some_key a b c d other_key /some/path /some/other/path.pl \ '$SOME_ENV_VAR$' \ '$OTHER_ENV_VAR$'
Here's what we need to handle:
- Keys: Identifiers like
some_keyorother_key(made up of letters, numbers, and underscores) - Values: Everything after a key until the next key starts (includes spaces, file paths, environment variables like
$SOME_ENV_VAR$, and content split across lines with\) - Line Continuations: A trailing
\means the next line's content is part of the current value (we need to merge these lines into a single string)
Let's build the parser piece by piece.
1. Import pyparsing and Define Core Elements
First, import the library and set up the basics:
import pyparsing as pp
Handle Line Continuations
We need to detect when a line ends with \ and merge the next line's content into the current value. This rule will suppress the \ and the following newline/whitespace:
# Match a line ending with \, followed by newline and any whitespace line_continuation = pp.LineEnd().suppress() + pp.White().suppress() escaped_newline = pp.Literal("\\").suppress() + line_continuation
Define Key Syntax
Keys are standard identifiers (letters, numbers, underscores—pyparsing has a built-in rule for this):
key = pp.Identifier()
Define Value Syntax
Values can include any characters except the start of a new key (which is a letter or underscore). We'll use Combine to merge content split by line continuations into a single string:
value_content = pp.Combine( pp.OneOrMore( # Match any character that isn't the start of a new key, or a line continuation pp.CharsNotIn(pp.alphas + "_", exact=1) | escaped_newline ), joinString="", adjacent=False )
2. Build the Key-Value Pair Parser
Now we'll group keys and their corresponding values together, so we can easily access them later:
key_value_pair = pp.Group( key + pp.Group(value_content).setName("value") )
3. Define the Full Parser
Our input is just one or more of these key-value pairs:
parser = pp.OneOrMore(key_value_pair)
Let's run the parser on your sample input (including a line continuation to test that feature):
# Test input with a line continuation for other_key's value test_input = """some_key a b c d other_key /some/path /some/other/path.pl \\ '$SOME_ENV_VAR$' \\ '$OTHER_ENV_VAR$'""" # Parse the input result = parser.parseString(test_input) # Print the parsed results print("Parsed Output:") for pair in result: print(f"Key: {pair[0]}") print(f"Value: {pair[1].strip()}") # Strip leading/trailing whitespace from the value print("---")
Expected Output
Parsed Output: Key: some_key Value: a b c d --- Key: other_key Value: /some/path /some/other/path.pl '$SOME_ENV_VAR$' '$OTHER_ENV_VAR$' ---
- If your keys need to include other characters (like hyphens), you can redefine
keyusingpp.Word(pp.alphanums + "-_")instead ofpp.Identifier(). - If your values can include characters that are now blocked (like letters at the start of a word that's not a key), you can adjust the
CharsNotInrule—just make sure it doesn't accidentally match the start of a new key.
内容的提问来源于stack exchange,提问作者rookie099

