OpenRewrite ChangeKey/ChangeValue用KeyPath性能差?求替代方案
Problem Description
When running the following OpenRewrite YAML recipes on a large number of files, execution time increases drastically, becoming a performance bottleneck:
Slow ChangeKey Recipes
- org.openrewrite.yaml.ChangeKey: oldKeyPath: $.['resilience4j.retry'].configs.*[?(@.maxRetryAttempts)] newKey: maxAttempts - org.openrewrite.yaml.ChangeKey: oldKeyPath: $.resilience4j.retry.configs.*[?(@.maxRetryAttempts)] newKey: maxAttempts - org.openrewrite.yaml.ChangeKey: oldKeyPath: $.['resilience4j.retry'].configs.*[?(@.max-retry-attempts)] newKey: max-attempts - org.openrewrite.yaml.ChangeKey: oldKeyPath: $.resilience4j.retry.configs.*[?(@.max-retry-attempts)] newKey: max-attempts
Slow ChangeValue Recipe
- org.openrewrite.yaml.ChangeValue: keyPath: $.spec.template.spec.containers.env[?(@.name == 'SERVER_TOMCAT_INTERNAL_PROXIES')].name value: SERVER_TOMCAT_REMOTEIP_INTERNAL_PROXIES
In contrast, recipes that don't use KeyPath (like the example below) perform well:
- org.openrewrite.java.spring.ChangeSpringPropertyKey: oldPropertyKey: "spring.data.cassandra.provider-ref" newPropertyKey: "spring.cassandra.provider-ref"
Optimization Attempt
To address this, a custom precondition recipe was created to filter files that contain the target string before applying the slow KeyPath recipes:
Custom Precondition Code
@Value @EqualsAndHashCode(callSuper = false) public class FindFilesContaining extends Recipe { private static final Logger log = LoggerFactory.getLogger(FindFilesContaining.class); transient SourcesFiles results = new SourcesFiles(this); @Option(displayName = "String to contain", description = "A String that needs to be contained in the file", required = true, example = "this is the string") @Nullable String searchString; // ... (omitted code) @Override public TreeVisitor<?, ExecutionContext> getVisitor() { return new TreeVisitor<Tree, ExecutionContext>() { @Override public @Nullable Tree visit(@Nullable Tree tree, ExecutionContext ctx) { if (tree instanceof SourceFile) { SourceFile sourceFile = (SourceFile) tree; Path sourcePath = sourceFile.getSourcePath(); try { String content = new String(Files.readAllBytes(sourcePath)); assert searchString != null; if (content.contains(searchString)) { log.warn("Found file: " + sourcePath); results.insertRow(ctx, new SourcesFiles.Row(sourcePath.toString(), tree.getClass().getSimpleName())); return SearchResult.found(sourceFile); } } catch (IOException e) { log.error("Error reading file: " + sourcePath, e); } } return tree; } }; } }
Optimized Recipe Configuration
preconditions: - org.me.FindFilesContaining: searchString: "max-retry-attempts" recipeList: - org.openrewrite.yaml.ChangeKey: oldKeyPath: $.['resilience4j.retry'].configs.*[?(@.max-retry-attempts)] newKey: max-attempts - org.openrewrite.yaml.ChangeKey: oldKeyPath: $.resilience4j.retry.configs.*[?(@.max-retry-attempts)] newKey: max-attempts
This approach drastically reduced runtime, but concerns exist about handling extremely large files with the current implementation (reading the entire file into memory).
Answers
1. Are KeyPath-based recipes inherently less scalable?
Yes. The performance overhead stems from two core factors:
- Full YAML Parsing: Every file must be parsed into an abstract syntax tree (AST) to evaluate the KeyPath, which is resource-intensive when processing hundreds or thousands of files.
- Complex JSONPath Evaluation: Wildcards (
*) and filter expressions ([?()]) require traversing the entire YAML structure for each file, adding significant per-file processing time.
Recipes like ChangeSpringPropertyKey avoid this overhead by targeting specific property keys directly without full AST traversal or complex path evaluation, leading to better scalability.
2. Is using preconditions/filters the right optimization approach?
Absolutely. Preconditions let you skip applying expensive recipes to files that don't need modification, which is the most impactful way to cut total runtime for large codebases. This aligns with OpenRewrite's design principles for efficient recipe execution.
3. Feedback on the custom precondition implementation
Your approach is effective but has limitations with large files:
- Memory Overhead: Reading the entire file into a
Stringcan trigger out-of-memory errors for gigabyte-scale files. - Redundant IO: OpenRewrite already reads file content during AST parsing; re-reading via
Files.readAllBytesadds unnecessary disk access.
Improvements to consider:
- Stream file content: Read the file line-by-line instead of loading it all into memory to check for the target string. This drastically reduces memory usage for large files:
try (Stream<String> lines = Files.lines(sourcePath)) { if (lines.anyMatch(line -> line.contains(searchString))) { log.warn("Found file: " + sourcePath); results.insertRow(ctx, new SourcesFiles.Row(sourcePath.toString(), tree.getClass().getSimpleName())); return SearchResult.found(sourceFile); } } - Use OpenRewrite's cached content: Access the file content via
sourceFile.printAll()instead of re-reading from disk. OpenRewrite caches parsed content, making this more efficient:String content = sourceFile.printAll(); if (content.contains(searchString)) { // Mark file for processing } - Leverage built-in preconditions: Use OpenRewrite's existing
org.openrewrite.search.FindTextprecondition instead of writing a custom one. This avoids reinventing the wheel and benefits from optimized, maintained code.
4. Additional Optimization Tips
- Combine similar recipes: Simplify your four
ChangeKeyrecipes into two by using KeyPaths that handle both quoted and unquoted keys, reducing redundant processing. - Narrow KeyPath specificity: If possible, replace wildcards (
*) with exact config names to limit traversal scope. - Enable parallel execution: Configure the Maven Rewrite plugin to run in parallel (check plugin docs for settings) to distribute workload across multiple threads.
内容的提问来源于stack exchange,提问作者JGleason

