cluster包pam函数执行起始逻辑及代码缺失对象问题咨询
Hey there, let's break down your questions about cluster::pam() step by step.
Troubleshooting the "Missing Object" Error in
cluster::pam() From the truncated code snippet you shared (ending at trace.lev...), here are the most likely causes for the "missing object" error:
- Missing mandatory parameters:
pam()requires two core arguments to run:x(your input data matrix or distance object) andk(the number of clusters). If either is omitted when calling the function, it will throw a missing object error. - Invalid input reference: Double-check that the
xobject you're passing actually exists in your current R environment. Typos in variable names or accidentally deleted objects are common triggers here. - Invalid
medoidsspecification: If you used themedoidsparameter, ensure the indices or objects you referenced are valid and exist in your dataset. A non-existent medoid reference can also lead to this error.
Starting Execution Logic of
cluster::pam() The pam() function follows these key steps right at the start of execution:
- Parameter Validation & Initial Setup
- First, it verifies that mandatory parameters (
x,k) are present and valid:kmust be a positive integer less than the number of samples, andxmust be either a numeric matrix/data frame or a validdistobject. - It resolves the
dissparameter: ifxis adistobject, it uses that precomputed distance matrix directly; otherwise, it calculates distances using the specifiedmetric(default is Euclidean). - If
stand = TRUE, it standardizes the input data matrix (scales variables to mean 0 and variance 1) to eliminate bias from differing variable scales.
- First, it verifies that mandatory parameters (
- Initial Medoid Selection
- If you don't specify
medoids, the function uses the BUILD algorithm to pick initial medoids: it starts with a random sample, then iteratively selects the sample farthest from all already chosen medoids untilkmedoids are selected. - If you provide
medoids, it uses those directly as the initial cluster centers.
- If you don't specify
- Initial Cluster Assignment
- For every sample in the dataset, it calculates the distance to each initial medoid, then assigns the sample to the cluster corresponding to the closest medoid.
- Optional Optimization (SWAP Step)
- If
do.swap = TRUE(the default), it runs the SWAP algorithm: it tests swapping each non-medoid sample with each medoid, calculates the total clustering cost (sum of distances from each sample to its cluster medoid) after each swap, and keeps swaps that reduce the cost—this continues until no more cost-reducing swaps are possible.
- If
内容的提问来源于stack exchange,提问作者theBotelho
相关产品推荐
相关产品推荐

