You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用stringr从字符变量中提取姓氏?技术求助

Hey there! Let's figure out why your code was throwing an error and fix it properly.

Your original code failed because str_match_all(Min3$Name, Name) tried to pair a 25-element vector (Min3$Name) with a 63-element list (Name). R can't handle this length mismatch, hence the error. Here's how to solve it with stringr:

Solution: Extract Matching Surnames to a New Column with stringr

Step 1: Build a Unified Regex Pattern

First, we'll turn all surnames in the Name list into a single regex pattern, using | to separate each option. This lets the regex engine match any surname from the list in Min3$Name:

library(stringr)

# Convert the Name list into a regex alternative pattern
# Add str_escape() inside if your surnames have special characters like . or ()
name_pattern <- str_c(Name, collapse = "|")

Step 2: Extract Matches to the New Column

Use str_extract instead of str_match_all—it returns a character vector with the same length as Min3$Name, which fits perfectly as a data frame column:

Min3$Name2 <- str_extract(Min3$Name, name_pattern)

Handling Edge Cases

  • If each entry in Min3$Name has only one matching surname, the code above works perfectly. For entries with multiple matches (e.g., a filename containing multiple surnames), use str_extract_all and combine results:
    # Combine multiple matches into a comma-separated string
    Min3$Name2 <- str_extract_all(Min3$Name, name_pattern) %>% 
      sapply(str_c, collapse = ", ")
    
  • If Min3$Name contains pure surnames (no extra filename characters), you don't even need regex—just use %in% to filter:
    Min3$Name2 <- ifelse(Min3$Name %in% Name, Min3$Name, NA)
    

That should get you the new column you need!

内容的提问来源于stack exchange,提问作者NColl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:39:33