更新qdap情感分析词典:组合词极性修正需求
Great question—this is a common gotcha with rule-based sentiment analysis tools like qdap::polarity, where multi-word phrases get split into individual terms and their sentiments cancel out. Here's how you can add custom multi-word negative phrases to the dictionary to get accurate results:
Step 1: Access and Modify the Default Polarity Dictionary
First, we'll make a copy of qdap's built-in polarity dictionary, then add your custom phrases as single entries marked as negative.
# Load required packages library(qdap) library(qdapDictionaries) # Create a copy of the default polarity dictionary custom_polarity <- polarity_df # Add your multi-word negative phrase (use lowercase to match qdap's default text normalization) custom_polarity <- rbind( custom_polarity, data.frame( term = "pretty bad", polarity = -1, stringsAsFactors = FALSE ) ) # Add more phrases if needed—just repeat the data.frame row for each one # Example: # custom_polarity <- rbind( # custom_polarity, # data.frame(term = "terribly wrong", polarity = -1, stringsAsFactors = FALSE), # data.frame(term = "super disappointing", polarity = -1, stringsAsFactors = FALSE) # )
Step 2: Run Polarity Analysis with the Custom Dictionary
When you call the polarity function, pass your modified dictionary using the custom.dict parameter. This tells qdap to prioritize your custom multi-word phrases over individual term matches.
# Test with your target phrase test_text <- "That restaurant experience was pretty bad" # Run polarity analysis with the custom dictionary sentiment_result <- polarity(test_text, custom.dict = custom_polarity) # View the results—you'll see the correct negative polarity score print(sentiment_result)
Key Notes
- Case Consistency: qdap automatically converts text to lowercase during analysis, so make sure your custom phrases are in lowercase too (e.g., "pretty bad" not "Pretty Bad").
- Phrase Priority: The
polarityfunction will match multi-word phrases first before splitting into individual terms, so your custom entries will override conflicting single-word sentiment scores. - Batch Additions: For multiple phrases, just add more rows to the
data.framewhen binding to the default dictionary.
This approach ensures phrases like "Pretty Bad" are treated as a single negative unit, so their sentiment doesn't get canceled out by conflicting individual terms.
内容的提问来源于stack exchange,提问作者Rana Usman

