将R语言dplyr代码转换为Python dfply代码报错求助
Let's break down what's going wrong with your code and fix it step by step. Your error AttributeError: 'DataFrame' object has no attribute 'invoiceprob' happens because you're trying to access the invoiceprob column from the original test DataFrame, but this column is only created later in the summarize step of your pipeline. Plus, there are a couple of other syntax issues that don't align with your original R dplyr logic.
What's Wrong with Your Original Code?
- Referencing the original DataFrame instead of pipeline data: When you use
test.InvoiceDocNumberortest.itemprob, you're pulling directly from the originaltestDataFrame, not the data being passed through the pipeline. This causes two big problems:max(test.itemprob)calculates the maximum value of the entireitemprobcolumn across all rows, not the maximum per group (which is what your R code does).test.invoiceprobdoesn't exist—this column is created in thesummarizestep, so the original DataFrame has no idea about it.
- Incorrect ranking approach: Using
rankdata(test.invoiceprob)again references the original DataFrame (which lacks the column) and ignores the pipeline's current state.
Corrected dfply Code
Here's the fixed version that matches your original R dplyr logic exactly:
from dfply import * # Option 1: Using pandas' built-in rank method (most similar to R's rank) tmp = (test >> group_by(X.InvoiceDocNumber) >> summarize(invoiceprob=X.itemprob.max()) >> mutate(invoicerank=X.invoiceprob.rank(ascending=False)))
Or if you prefer using rankdata from scipy (note the negative sign to replicate descending ranking):
from dfply import * from scipy.stats import rankdata tmp = (test >> group_by(X.InvoiceDocNumber) >> summarize(invoiceprob=X.itemprob.max()) >> mutate(invoicerank=rankdata(-X.invoiceprob)))
Key Explanations
- The
Xsymbol: This is dfply's way of referring to the current DataFrame being passed through the pipeline. It's equivalent to the implicit data reference in R's dplyr. UsingX.InvoiceDocNumber,X.itemprob, andX.invoiceprobensures you're always working with the data at each step of the pipeline. - Group-wise max:
X.itemprob.max()calculates the maximumitemprobfor each group defined byInvoiceDocNumber, just likemax(itemprob)in your R code. - Descending ranking:
X.invoiceprob.rank(ascending=False)replicates R'srank(desc(invoiceprob))—it ranks values from highest to lowest.- If using
rankdata, adding a negative sign (-X.invoiceprob) converts descending ranking into ascending (sincerankdatadefaults to ascending order).
Quick Comparison to Your Original R Code
| R dplyr Code | dfply Python Code |
|---|---|
group_by(InvoiceDocNumber) | group_by(X.InvoiceDocNumber) |
summarise(invoiceprob=max(itemprob)) | summarize(invoiceprob=X.itemprob.max()) |
mutate(invoicerank=rank(desc(invoiceprob))) | mutate(invoicerank=X.invoiceprob.rank(ascending=False)) |
This fix will resolve the AttributeError and ensure your Python code behaves exactly like your original R code.
内容的提问来源于stack exchange,提问作者vinay karagod

