You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Databricks PySpark中使用Azure认知文本分析时遭遇「不支持混合字符串与字典/对象文档输入」类型错误的解决求助

Fix "Mixing string and dictionary/object document input unsupported" Error in Azure Text Analytics

Let's break down what's causing this error and how to fix it quickly.

Your Scenario Recap

You're trying to extract key phrases from a text file stored in DBFS using Azure Text Analytics, but hitting a type error: Mixing string and dictionary/object document input unsupported. Here's your original code flow:

  1. Reading the file with binary mode:
with open("/dbfs/mnt/lake/RAW/export/dummy.txt", "rb") as fd:
    documents = fd.read()
  1. Calling the Text Analytics API:
response = text_analytics_client.extract_key_phrases(documents, language="en")

Root Cause

The extract_key_phrases method expects an iterable of strings (like a list of documents), not a single raw bytes object or a single string. When you pass the raw documents variable (which is bytes from rb mode) or even a single string, the SDK misinterprets the input format, leading to the type mismatch error.

Fixed Solution

Here's the corrected code with explanations:

  1. Read the file as text (not binary)
    Use text mode (r) instead of binary (rb) to get a proper string, and strip any extra whitespace to avoid unintended issues:

    with open("/dbfs/mnt/lake/RAW/export/dummy.txt", "r", encoding="utf-8") as fd:
        document_text = fd.read().strip()
    
  2. Wrap the single document in a list
    The API is built to handle multiple documents at once, so even for one file, you need to pass it as a list to match the expected input format:

    from azure.core.credentials import AzureKeyCredential
    from azure.ai.textanalytics import TextAnalyticsClient
    
    credential = AzureKeyCredential("xxxxxxxxxxxxxxxxxxxx")
    endpoint= "https://xxxxxxxx.cognitiveservices.azure.com/"
    text_analytics_client = TextAnalyticsClient(endpoint, credential)
    
    # Pass the text as a list (even for a single document)
    response = text_analytics_client.extract_key_phrases([document_text], language="en")
    result = [doc for doc in response if not doc.is_error]
    for doc in result:
        print(doc.key_phrases)
    

Expected Output

Running this corrected code will give you the desired result:

['King County', 'United States', 'Redmond', 'city', 'Washington', 'Seattle']

内容的提问来源于stack exchange,提问作者Patterson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 22:02:44