finding if boolean is ever true by groups in R

I want a simple way to create a new variable determining whether a boolean is ever true in R data frame. Here is and example: Suppose in the dataset I have 2 variables (among other variables which are not relevant) 'a' and 'b' and 'a' determines a group, while 'b' is a boolean with values TRUE (1) or FALSE (0). I want to create a variable 'c', which is also a boolean being 1 for all entries in groups where 'b' is at least once 'TRUE', and 0 for all entries in groups in which 'b' is never TRUE. From entries like below:

I want to get variable 'c' like below:

a   b   c
-----------
1   1   1 
2   0   0
1   0   1
1   0   1
1   1   1
2   0   0
2   0   0
3   0   1
3   1   1
3   0   1
-----------

I know how to do it in Stata, but I haven't done similar things in R yet, and it is difficult to find information on that on the internet. In fact I am doing that only in order to later remove all the observations for which 'c' is 0, so any other suggestions would be fine as well. The application of that relates to multinomial logit estimation, where the alternatives that are never-chosen need to be removed from the dataset before estimation.

标签： r boolean mlogit

4条回答

你好瞎i

2楼-- · 2019-02-28 00:16

if X is your data frame

library(dplyr)
X <- X %>%
  group_by(a) %>%
  mutate(c = any(b == 1))

0人赞添加讨论(0) 举报

beautiful°

3楼-- · 2019-02-28 00:21

A base R option would be

 df1$c <- with(df1, ave(b, a, FUN=any))

 library(sqldf)
 sqldf('select * from df1
      left join(select a, b,
         (sum(b))>0 as c
         from df1 
         group by a)
         using(a)')

0人赞添加讨论(0) 举报

Juvenile、少年°

4楼-- · 2019-02-28 00:31

Simple data.table approach

require(data.table)
data <- data.table(data)
data[, c := any(b), by = a]

Even though logical and numeric (0-1) columns behave identically for all intents and purposes, if you'd like a numeric result you can simply wrap the call to any with as.numeric.

0人赞添加讨论(0) 举报

唯我独甜

5楼-- · 2019-02-28 00:32

An answer with base R, assuming a and b are in dataframe x

c value is a 1-to-1 mapping with a, and I create a mapping here

cmap <- ifelse(sapply(split(x, x$a), function(x) sum(x[, "b"])) > 0, 1, 0)

Then just add in the mapped value into the data frame

x$c <- cmap[x$a]

Final output

edited to change call to split.

0人赞添加讨论(0) 举报

finding if boolean is ever true by groups in R

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间