PySpark ML中用MinHash LSH实现Jaccard相似度时遇导入错误
Hey there, that import error usually pops up because of one of two common issues: either you're using the wrong library, or your import statement doesn't match the library's structure. Let's walk through how to fix this step by step.
Step 1: Make sure you're using the right library
The most popular Python library for MinHash LSH (especially for shopping basket data like yours) is datasketch. If you haven't installed it yet, run this command in your terminal:
pip install datasketch
Step 2: Use the correct import statement
In the latest versions of datasketch, MinHashLSH lives directly under the datasketch module. Your import should look like this:
from datasketch import MinHash, MinHashLSH
Avoid trying to import it from submodules (like datasketch.lsh), as that's deprecated in newer releases.
Step 3: Upgrade if you have an older version
If you already have datasketch installed but still get the error, your version might be out of date. Upgrade it with:
pip install --upgrade datasketch
Quick Example with Your Shopping Basket Data
Here's a tiny snippet to test things out with your Cust_ID/Item_id dataset structure:
from datasketch import MinHash, MinHashLSH import pandas as pd # Sample shopping basket data data = pd.DataFrame({ 'Cust_ID': ['C1', 'C1', 'C2', 'C2', 'C3'], 'Item_id': ['I1', 'I2', 'I2', 'I3', 'I1'] }) # Create MinHash objects for each customer lsh = MinHashLSH(threshold=0.5, num_perm=128) customer_hashes = {} for cust_id, group in data.groupby('Cust_ID'): mh = MinHash(num_perm=128) for item in group['Item_id']: mh.update(item.encode('utf8')) customer_hashes[cust_id] = mh lsh.insert(cust_id, mh) # Query similar customers for C1 result = lsh.query(customer_hashes['C1']) print(f"Customers similar to C1: {result}")
This should work without the import error once you've set up the library correctly.
内容的提问来源于stack exchange,提问作者Sai Kiran Kodukula

