You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中从产品文本提取年龄信息并生成新列?

Got it, using stringr::str_extract is exactly the right call here! Let's walk through how to pull that age number out cleanly.

First, you'll want a regular expression that targets the digits immediately following age:. A positive lookbehind is perfect for this—it lets us check for age: without including it in our final extracted result.

Here's a complete, reproducible example with sample data matching your use case:

# Load the stringr package (install first if you haven't: install.packages("stringr"))
library(stringr)

# Sample data frame with your product info column
product_data <- data.frame(
  product_details = c(
    "Technical Details Manufacturer recommended age:14 years and up Manufacturer reference176-1308 Scale1::160 Track Width/GaugeNo Additional Information ....",
    "Kids Toy age:3 years and up Some other details here", # Test a different age
    "No age info here" # Test a row without age data
  )
)

# Extract the age number and create a new column
product_data$recommended_age <- str_extract(product_data$product_details, "(?<=age:)\\d+")

# Optional: Convert to numeric if you need to use the age for calculations/analysis
product_data$recommended_age <- as.numeric(product_data$recommended_age)

# View the result
product_data

Quick regex breakdown:

  • (?<=age:): This is a positive lookbehind assertion. It tells R "only match what comes next if it's directly preceded by age:".
  • \\d+: Matches one or more consecutive digits (so it works for ages like 14, 3, 21, etc.)

If a row doesn't have an age: entry, this will return NA—which is ideal for handling missing values consistently.

内容的提问来源于stack exchange,提问作者Nastya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:39:49