Linux CentOS 7下dplyr 0.7.5 select函数失效及旧代码兼容问题求助
Hey there, I totally get how frustrating it is to slog through all those dependency updates just to get sparklyr working, only to have your trusted old dplyr code break left and right. Let’s break down the most common issues you’re facing and how to fix them:
1. Fix the Broken select Function
The biggest shift between dplyr 0.5.0 and 0.7.x is the introduction of tidy evaluation (tidy eval), which reworked how functions like select handle variable references. Here’s how to tackle this:
- Namespace conflicts first: If you have other packages loaded (like MASS, which has its own
selectfunction), explicitly call dplyr’s version withdplyr::select(df, your_columns)instead of justselect(). This stops other packages from "masking" dplyr’s function. - Update custom
selectfunctions: If you had helper functions like this in your old code:
Rewrite it to use tidy eval syntax (required in dplyr 0.7+):old_select <- function(data, col) { select(data, col) }
If you work with column names as strings, usenew_select <- function(data, col) { select(data, !!enquo(col)) }select_at()instead:string_select <- function(data, col_string) { select_at(data, col_string) } - Basic
selectcalls: Simple calls likeselect(df, col1, col2)should still work, but if they’re failing, double-check for typos or masked functions as noted above.
2. Resolve General Code Crashes
Beyond select, dplyr 0.7.x deprecated or changed behavior for several core functions. Here’s what to check:
- Replace deprecated functions: Old functions like
summarise_each()were replaced withsummarise_at()/summarise_if()—swap these out for the newer equivalents. - Adjust variable scoping: If your
mutateorfiltercalls are crashing, make sure you’re referencing variables correctly. For complex expressions, you might need to use!!(bang-bang) to "unquote" variables in line with tidy eval rules. - Check for conflicts: Run
dplyr::conflicts()to see if any other packages are overriding dplyr functions. Resolve these by either unloading conflicting packages or using explicitdplyr::prefixes for affected functions.
3. Prevent Future Dependency Headaches
Since you had to wade through so many dependency updates to install dplyr 0.7.5 on CentOS, consider using a dependency manager like packrat to lock your package versions. This way, you won’t accidentally upgrade packages later and break your code again. It also makes reinstalling exact versions a breeze if you ever need to set up your environment again.
Bonus: Partial Rollback Option (If Needed)
If rewriting all your old code feels overwhelming right now, check which dplyr versions are compatible with your sparklyr installation. You can install a specific version with:
devtools::install_version("dplyr", version = "0.7.0", repos = "http://cran.us.r-project.org")
Just make sure the version you pick is new enough to support sparklyr’s parquet/Spark DataFrame functionality.
内容的提问来源于stack exchange,提问作者Ke Cheng

