如何将Python中余弦相似度匹配的打印输出转为Pandas DataFrame
解决方案
你可以通过以下步骤将输出转换为Pandas DataFrame:
- 导入
pandas库 - 初始化一个空列表用于存储所有匹配结果
- 替换原有的打印逻辑,将每个query、匹配的corpus文本和相似度分数以字典形式存入列表
- 最后将列表转换为DataFrame
修改后的完整代码如下:
import pandas as pd from sentence_transformers import SentenceTransformer, util import torch embedder = SentenceTransformer('all-MiniLM-L6-v2') corpus = ['About us. About Us · Our Coffees · Starbucks Stories & News · Starbucks® Ready to Drink · Foodservice Coffee · Customer Service · Tax Strategy 2022 · Careers.', 'Costa is the Nation Favourite coffee shop and the largest and fastest growing coffee shop chain in the UK.', 'Leading UK speciality coffee roaster with a focus on sustainability. B Corp certified. Become a wholesale partner or buy coffee beans online today.', 'Kick-start your morning with our amazing range of speciality coffee and equipment. World-class coffee, direct from the farmer, delivered free every time.', 'Coffee Direct - Freshly roasted coffee beans delivered to your door. Origin coffee, coffee blends and flavoured coffee for bean-to-cup', 'Whether you prefer whole coffee beans or freshly ground coffee, Whittard of Chelsea selection of light, medium and dark roast luxury coffees has something', 'Coffee beans are the seeds of a fruit called a coffee cherry. Coffee cherries grow on coffee trees from a genus of plants called Coffea.', 'On these coffee plants, bunches of cherries grow and inside these you will find two coffee beans, Arabica and Robusta coffee.', ] corpus_embeddings = embedder.encode(corpus, convert_to_tensor=True) # Query sentences: queries = ['coffee', 'coffee near me', 'coffee bean', 'coffee house', 'coffee jelly','coffee order nyt crossword clue','coffee quotes', 'coffee shops near me'] # 初始化存储结果的列表 results = [] # Find the closest 5 sentences of the corpus for each query sentence based on cosine similarity top_k = min(5, len(corpus)) for query in queries: query_embedding = embedder.encode(query, convert_to_tensor=True) cos_scores = util.cos_sim(query_embedding, corpus_embeddings)[0] top_results = torch.topk(cos_scores, k=top_k) # 将结果存入列表 for score, idx in zip(top_results[0], top_results[1]): results.append({ "query": query, "corpus": corpus[idx], "score": round(score.item(), 4) }) # 转换为DataFrame similarity_df = pd.DataFrame(results) # 可以打印查看结果,或者保存为CSV print(similarity_df.head(10)) # similarity_df.to_csv("coffee_similarity_results.csv", index=False)
关键修改说明:
- 使用
results列表收集所有匹配项,每个元素是包含query、corpus、score的字典 - 通过
score.item()将PyTorch张量转换为Python浮点数,并用round()保留4位小数,和原打印格式一致 - 最后用
pd.DataFrame()将列表转换为结构化的表格数据
生成的DataFrame示例(部分):
| query | corpus | score |
|---|---|---|
| coffee | Coffee Direct - Freshly roasted coffee beans delivered to your door. Origin coffee, coffee blends and flavoured coffee for bean-to-cup | 0.6477 |
| coffee | Whether you prefer whole coffee beans or freshly ground coffee, Whittard of Chelsea selection of light, medium and dark roast luxury coffees has something | 0.5873 |
| coffee | Kick-start your morning with our amazing range of speciality coffee and equipment. World-class coffee, direct from the farmer, delivered free every time. | 0.5739 |
内容的提问来源于stack exchange,提问作者Simone De Palma
相关产品推荐
相关产品推荐

